Authentication & Base URL
Our API is designed to be a drop-in replacement for existing OpenAI clients. You only need to change two things in your configuration: the base URL and the API key. The base URL for this service is https://api.uncensoredgpt.top/v1. You can generate your API key immediately after signing up via Google or email on the dashboard. There is no phone number required, and the key is displayed once upon creation.
Ensure you store your key securely. If you lose it, you can generate a new one from the dashboard, which will invalidate the previous key. This simple setup allows you to connect to our uncensored ai api using standard libraries without custom adapters.
Chat Completions Endpoint
The core functionality is served through the standard POST /v1/chat/completions endpoint. This endpoint accepts text input and returns text output, supporting both synchronous and asynchronous requests. You do not need to worry about model routing or aggregation; we serve a single, high-performance uncensored large language model tuned for unrestricted content.
When making a request, you must specify the model ID as uncensored. This model is open-weight and hosted on our own servers, distinct from GPT, Claude, or other vendor models. It is designed to answer without content refusals for lawful adult use, making it ideal for NSFW or controversial topics.
To get started, you can test the endpoint with a simple cURL command. Replace YOUR_API_KEY with your actual key.
curl https://api.uncensoredgpt.top/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
Streaming Responses (SSE)
For real-time applications, the API supports Server-Sent Events (SSE). Streaming allows you to receive tokens as they are generated, providing a better user experience for chat interfaces. When you set stream: true in your request, the API returns a stream of chunks rather than a single complete response.
Each chunk contains partial text, and the final chunk includes the token usage statistics for billing purposes. This feature is particularly useful for applications requiring immediate feedback. You can handle the stream in your preferred language using the appropriate SDK.
Here is an example of how to initiate a streaming request using cURL:
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Function Calling & Tools
The API supports function calling, allowing your application to interact with external tools and services. You can define a list of functions in the tools parameter, and the model will respond with a function call if appropriate. This is useful for building agents or applications that need to perform specific actions.
You can control the model's behavior using tool_choice, which can be set to auto, none, or a specific function ID. The model will return the function name and arguments in the response, which you can then execute in your application.
Here is a Python example demonstrating how to use the function calling feature with the OpenAI-compatible SDK:
from openai import OpenAI
client = OpenAI(base_url="https://api.uncensoredgpt.top/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
JSON Mode Configuration
If you need structured output, you can enable JSON mode by setting the response_format parameter to {"type": "json_object"}. This instructs the model to generate output that conforms to valid JSON syntax. This is particularly useful for applications that require parsing the response programmatically.
When using JSON mode, the model will avoid adding markdown fences or extra text around the JSON object. This ensures that your parser can directly consume the response without additional cleanup steps. It is a reliable way to get structured data from an LLM.
Here is a Node.js example showing how to configure the client for JSON mode:
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.uncensoredgpt.top/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
Parameters: Temperature & Top P
You can control the randomness and creativity of the model's responses using parameters like temperature and top_p. The temperature parameter adjusts the sampling temperature, with lower values making the output more deterministic and higher values making it more creative. The top_p parameter controls nucleus sampling, limiting the model to consider only the top p probability mass of tokens.
Other supported parameters include stop sequences, seed for reproducibility, and penalties for presence_penalty and frequency_penalty. These parameters allow you to fine-tune the model's behavior to suit your specific use case.
Understanding these parameters is crucial for optimizing the quality of your generated text. Experiment with different values to find the right balance for your application.
Capabilities and limits
If your tool speaks the OpenAI API, these are the details that matter.
| Item | Value |
|---|---|
| API format | OpenAI Chat Completions schema; official openai SDKs work unchanged |
| Methods | POST /v1/chat/completions · GET /v1/models |
| Base URL | https://api.uncensoredgpt.top/v1 |
| Authentication | Bearer token in the Authorization header |
| Model ID | uncensored |
| Streaming | Supported (stream: true), usage included at the end |
| Sampling parameters | temperature, top_p, stop, seed, presence_penalty, frequency_penalty |
| Function calling | Yes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool |
| Structured output | JSON object mode via response_format json_object |
| Max output | up to the rest of the 100,000-token window; max_tokens optional (no separate cap) |
| Max context | 100,000 tokens (prompt + completion together) |
| Request size | 8 MB request body |
| Requests per minute | 300 requests per minute per key |
| Concurrency | 8 requests at the same time per key |
| Headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Token prices | $0.25 per 1M input tokens · $1.00 per 1M output tokens |
| Trial credit | $0.50 of credit valid 7 days, no card needed · Trial key: 2 parallel requests, 60 req/min; full limits (8 and 300) after first top-up |
| Billing | prepaid credit, charged by real token usage; errors and refusals are free |
| Top-up | crypto: USDT on TRON or USDC on Base, $10–$500, any whole sum |
| Volume bonus | +5% from $50, +10% from $100 |
| Subscription | no monthly fee; paid credit does not expire |
| Key management | one active key per account; a new key replaces the old one |
| Content policy | adult content allowed; sexual content involving minors is refused |
| Account | sign in with Google or with e-mail + password |
Errors and what to do
The type field is stable, the message is for humans. Errors cost nothing.
| Status | Type | Reason |
|---|---|---|
400 | bad_request | invalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend |
401 | missing_key · invalid_key · key_revoked | no key, wrong key, or a key replaced by a newer one |
402 | no_credit | balance is empty — top up, requests resume at once |
403 | content_blocked | sexual content involving minors — refused, not billed |
404 | not_found | only /v1/chat/completions and /v1/models exist |
413 | request_too_large | request body larger than 8 MB |
429 | rate_limited · concurrency | slow down: rate or parallel limit reached |
503 | upstream_busy | model busy — retry in a few seconds |
Questions and answers
What happens if I exceed the rate limit?
If you exceed the limit of 300 requests per minute or 8 concurrent requests, you will receive a 429 Too Many Requests error. Your API key is limited to 8 simultaneous connections, so ensure your application handles concurrent requests appropriately to avoid hitting this limit.
Why do I get a 402 error?
A 402 error indicates that your prepaid credit has been exhausted. Since we use a prepaid token model, you must top up your account to continue using the API. Errors and refusals do not consume credit, so you only pay for successful token usage.
What is the context window limit?
The model supports a context window of 100,000 tokens, which includes both the prompt and the completion. The maximum output per request is 32,000 tokens, or 2,048 tokens if you do not specify a max_tokens value. This allows for long-form content generation and complex reasoning tasks.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.