Get API key

Uncensored Chat Completions API Documentation

This documentation guides you through integrating our uncensored LLM via the standard OpenAI-compatible chat completions endpoint. You can start generating text immediately by configuring your client with our base URL and a freshly generated API key.

On this page
  1. Authentication & Setup
  2. Chat Completions Endpoint
  3. Python SDK Integration
  4. Streaming Responses
  5. Rate Limits & Quotas

Authentication & Setup

Access requires a valid API key associated with your account. You generate this key on the Get API Key page after signing up with an email and password. The key is displayed immediately upon creation. To switch from a vendor like OpenAI, update your client configuration to point to our base URL: https://api.unfilteredaimodels.com/v1 and supply your unique API key. We support the standard Authorization header format.

Chat Completions Endpoint

The primary interface is the POST /v1/chat/completions endpoint. It accepts standard message objects and returns text responses. The model ID to specify is uncensored. This endpoint handles raw text generation without the typical corporate guardrails, allowing for adult, controversial, or creative content within legal bounds. You can send your first request using the following example:

curl https://api.unfilteredaimodels.com/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

This request demonstrates sending a simple system prompt and a user query to generate a response.

Python SDK Integration

Use the official openai Python package. Initialize the client with our base URL and your API key. Then call chat.completions.create with the model ID uncensored. This approach requires zero code changes if you are already using the OpenAI SDK for other services. The following snippet shows how to initialize the client and send a request:

from openai import OpenAI

client = OpenAI(base_url="https://api.unfilteredaimodels.com/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

The SDK handles JSON serialization and response parsing automatically.

Streaming Responses

Set the stream parameter to true to receive server-sent events (SSE). This allows you to process tokens as they are generated, reducing perceived latency for long responses. The API returns a stream of chunks, each containing partial deltas of the response. Use the following example to implement streaming:

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Streaming is ideal for chat interfaces where users expect to see text appearing in real-time.

Rate Limits & Quotas

Your API key is limited to 300 requests per minute. If you exceed this, you receive a 429 Too Many Requests error. You can regenerate your key at any time from your dashboard, which revokes the old key and provides a fresh one with a new rate limit counter. Other errors include 401 for invalid keys and 402 if your prepaid credit is exhausted. The context window supports up to 64,000 tokens for prompt and completion combined. Request bodies are limited to 8 MB.

Node.js

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.unfilteredaimodels.com/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Questions and answers

Does regenerating my API key reset my rate limit?

Yes. Regenerating your key revokes the old one and issues a new one with a fresh rate limit counter. The 300 requests per minute limit applies to the specific key in use.

What content is restricted?

The model does not refuse most adult, fictional, or controversial topics. However, a hard limit always applies: sexual content involving minors is blocked.

Is the model fine-tuned or a standard GPT?

It is an open-weight model run on our own GPU servers, tuned for uncensored output. It is not GPT, Claude, or any other vendor's model.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.