Uncensored LLM API Documentation

Get started with the uncensored LLM API in minutes. This quickstart covers authentication, basic requests, streaming, and key constraints for building your NSFW chat application.

Base URL and Authentication

The API is hosted at https://api.nsfwchat.top/v1. It follows the OpenAI-compatible protocol, so you can use existing SDKs by updating the base URL and providing your API key. Generate your key on the Get API key page via Google or email. The key grants full access to the chat endpoint. Keep it secure, as it controls your prepaid credit. You can have only one active key per account; generating a new one replaces the previous key immediately.

  • Base URL: https://api.nsfwchat.top/v1
  • Auth Header: Authorization: Bearer YOUR_API_KEY

Verify your connection by listing available models. This confirms your key is valid and the service is responsive.

First Request

Send a simple text completion to the /v1/chat/completions endpoint. Use the model ID uncensored. This open-weight model is tuned to answer without content refusals for lawful adult use. You can specify parameters like temperature or stop sequences to control the output.

Here is a basic cURL example to get a response:

curl https://api.nsfwchat.top/v1/chat/completions \
  -H "Authorization: Bearer $API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "uncensored",
    "messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
  }'

The response includes the generated text in choices[0].message.content. Errors like invalid keys return 401, and insufficient credit returns 402. Mistakes in your prompt do not charge credit if the request fails before processing.

Python SDK Integration

Use the official openai Python package. Configure the client to point to our base URL. This allows you to leverage familiar methods like chat.completions.create(). The SDK handles JSON serialization and retry logic automatically.

Example setup:

from openai import OpenAI

client = OpenAI(base_url="https://api.nsfwchat.top/v1", api_key="YOUR_KEY")

resp = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)

Set the model to uncensored. You can pass messages as a list of dictionaries with roles like user or assistant. This structure supports multi-turn conversations. The API handles context retention within the 64k token window. No additional setup is required for JSON mode or function calling.

Node SDK Usage

The Node.js SDK works identically to the Python version. Initialize the client with your API key and the correct base URL. This ensures compatibility with any OpenAI-compatible client library.

Example initialization:

import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://api.nsfwchat.top/v1", apiKey: process.env.API_KEY });

const resp = await client.chat.completions.create({
  model: "uncensored",
  messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);

Call chat.completions.create() with the uncensored model. You can stream responses or get full JSON objects. The API supports standard parameters like top_p and seed for reproducibility. Error handling should check for 401 (auth) and 402 (billing) statuses. These errors are explicit and do not consume credit.

Streaming Responses via SSE

Enable streaming by setting stream: true. The API returns Server-Sent Events (SSE). Each chunk contains partial token data. The final chunk includes the full token usage statistics for billing.

Streaming example:

stream = client.chat.completions.create(
    model="uncensored",
    messages=[{"role": "user", "content": "Tell the story in second person."}],
    stream=True,
)
for chunk in stream:
    if chunk.choices and chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="", flush=True)

Process each chunk as it arrives. This reduces perceived latency for your users. Token counts are only reported in the last chunk. If you need accurate usage metrics, parse the final event. Streaming is ideal for chat interfaces where real-time text display is preferred.

Limits, Errors, and Context

The context window is 64,000 tokens total. Max output is 16,000 tokens, or 2,048 if max_tokens is unset. Rate limits are 300 requests per minute and 8 concurrent requests per key. Request bodies must stay under 8 MB. Errors return standard HTTP codes: 401 for invalid keys, 402 for no credit, and 429 for rate limits. These errors are free. Content restrictions apply only to minor sexual content. All other lawful adult content is allowed. Use these constraints to design your application's retry and buffering logic.

Questions and answers

How much does the API cost?

Pricing is $0.25 per 1M input tokens and $1.00 per 1M output tokens. Credit is prepaid and never expires. Errors and refusals do not charge credit.

Can I use credit cards?

No. Top-ups are crypto only: USDT (TRC20) or USDC (Base). Minimum top-up is $10. No card, PayPal, or bank transfer is accepted.

Is there a free trial?

Yes. New accounts get $0.50 trial credit valid for 7 days. No card is needed. One trial per person.

Your key is one form away

Create an account, copy the key, change the base URL. That is the whole setup.

Get API key