Integrate the Uncensored AI Model API
Get your uncensored AI model API key and integrate the chat completions endpoint in minutes. Change your base_url and key to switch to our drop-in OpenAI compatible API for immediate, refusal-free responses.
Authentication and Base URL
Our API is designed as a drop-in replacement for existing OpenAI clients. You only need to change two variables: the base_url and your api_key. The base URL for all requests is https://api.llmmodelsapi.com/v1. Authentication is handled via the Authorization header using a Bearer token. Ensure you keep your API key secure, as it grants access to your prepaid credit. You can regenerate your key at any time from your dashboard, which immediately revokes the old one.
First Request
Send a standard chat completion request to test connectivity. Use the model ID uncensored to access our purpose-built large language model. This model is tuned to answer without content refusals for lawful adult use, making it ideal for agents that need consistent outputs. The request body follows the standard OpenAI chat completions format.
curl https://api.llmmodelsapi.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
Verify the response includes a valid completion. If you receive a 200 status code, your integration is working correctly. This ai model api endpoint supports both single-turn and multi-turn conversations within the context window.
Python SDK Integration
Using the official OpenAI Python SDK is the fastest way to integrate. Point the client to our base URL and provide your API key. The SDK handles serialization and connection pooling automatically. This approach works with any OpenAI-compatible client library that respects the base_url override.
from openai import OpenAI
client = OpenAI(base_url="https://api.llmmodelsapi.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
The Python SDK supports streaming and synchronous calls. For most agent workflows, synchronous calls are simpler, but streaming is available if you need to display tokens as they generate. The model ID remains uncensored regardless of the SDK version.
Node SDK Integration
For JavaScript and TypeScript environments, the Node SDK follows the same pattern. Initialize the client with the correct base URL and credentials. This ensures your existing codebase requires minimal changes to switch providers. The uncensored model ID is compatible with all standard chat completion parameters.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.llmmodelsapi.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
The Node SDK also supports async/await patterns, making it easy to integrate into Express routes or serverless functions. Keep your API key in environment variables to prevent accidental exposure. This openai compatible api design means you can swap back to OpenAI or switch to other providers by changing only the configuration object.
Streaming Responses
Enable streaming by setting stream: true in your request. The API returns a Server-Sent Events (SSE) stream, allowing you to process tokens as they are generated. This reduces perceived latency for end-users and is essential for real-time agent interactions. The stream format is identical to the OpenAI SSE specification.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Each chunk contains partial delta data. You must accumulate these deltas to reconstruct the full response. Streaming does not affect pricing; you are charged for the total input and output tokens regardless of streaming mode. This feature is part of the robust chat completions api design we provide.
Limits, Errors, and Context
Our infrastructure enforces strict limits to ensure reliability. You are limited to 300 requests per minute per key, with a maximum request body size of 8 MB. The context window is 64,000 tokens, shared between prompt and completion. If you exceed the rate limit, you will receive a 429 error. An invalid key returns 401, and insufficient credit returns 402. Unlike many llm api provider options, we do not hide these limits; they are transparent and consistent. No training is performed on your data, and privacy is maintained by not using prompts for model improvement.
API specifications
One table with every limit, feature and price that applies to your key.
| Feature | Support |
|---|---|
| Compatibility | OpenAI-compatible: any OpenAI SDK or client works — change the base URL and the key |
| Base URL | https://api.llmmodelsapi.com/v1 |
| Endpoints | POST /v1/chat/completions · GET /v1/models |
| Authentication | Bearer token in the Authorization header |
| Model ID | uncensored |
| Tools / tool calls | Supported: tools + tool_choice, tool_calls in the reply (streamed too), tool results as role: tool messages |
| Structured output | response_format: {"type": "json_object"} |
| Sampling parameters | temperature, top_p, stop, seed, presence_penalty, frequency_penalty |
| Streaming | Yes — server-sent events; the last chunk carries token usage |
| Max output | up to 16,000 tokens per request (default 2,048) |
| Max context | 64,000 tokens (prompt + completion together) |
| Headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Request size | 8 MB request body |
| Requests per minute | 300 requests per minute per key |
| Concurrency | 8 requests at the same time per key |
| Credit expiry | no monthly fee; paid credit does not expire |
| Free trial | $0.50 of credit valid 7 days, no card needed |
| How you pay | prepaid credit, charged by real token usage; errors and refusals are free |
| Token prices | input $0.25 / 1M tokens, output $1.00 / 1M tokens |
| Volume bonus | +5% from $50, +10% from $100 |
| Top-up | crypto: USDT on TRON or USDC on Base, $10–$500, any whole sum |
| Content | adult content allowed; sexual content involving minors is refused |
| Account | Google or e-mail and password |
| Key management | one active key per account; a new key replaces the old one |
Errors and what to do
Errors come back as JSON with a stable type; failed and refused requests are not billed.
| Status | Type | Reason |
|---|---|---|
400 | bad_request | malformed request or too long for the context window |
401 | missing_key · invalid_key · key_revoked | check the Authorization header or use your current key |
402 | no_credit | balance is empty — top up, requests resume at once |
403 | content_blocked | refused by the content policy |
404 | not_found | only /v1/chat/completions and /v1/models exist |
413 | request_too_large | request body larger than 8 MB |
429 | rate_limited · concurrency | slow down: rate or parallel limit reached |
503 | upstream_busy | model busy — retry in a few seconds |
Questions and answers
Is this the same model as GPT-4 or Claude?
No. Our model is an open-weight model tuned specifically for uncensored responses. It is not GPT, Claude, Gemini, or any other vendor's model. It is a distinct model run on our own GPU servers.
How does pricing work?
We use a pay-as-you-go prepaid credit system with no monthly fees. Input tokens cost $0.25 per 1M, and output tokens cost $1.00 per 1M. Credits never expire, and you can top up with crypto (USDT or USDC) starting at $10.
What is the trial credit?
Every new account receives $0.50 in trial credit valid for 7 days. No credit card or phone number is required to start. This allows you to test the <code>uncensored ai model</code> before committing to a purchase.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.