gpt5apikey.com/Quickstart for GPT 5 users
Quickstart for GPT 5 users
Get started with the uncensored LLM API by configuring your base URL and API key. Use the standard OpenAI-compatible endpoints to send text, stream responses, or call functions with zero configuration overhead.
Base URLhttps://api.gpt5apikey.com/v1
Authentication: Using Your API Key
To begin, you need an API key from your account dashboard. The gpt 5 api uses standard bearer token authentication. Include your key in the Authorization header as Bearer YOUR_API_KEY. This key grants access to the single uncensored model via the base URL https://api.gpt5apikey.com/v1. If you lose your key, you can regenerate it at any time, which immediately invalidates the previous one. There is no need to link a credit card for the initial trial credit, which provides $0.50 valid for 7 days.
Endpoint: POST /v1/chat/completions
Send your prompt to the chat completions endpoint. The API accepts standard OpenAI JSON payloads. Ensure you set the model field to uncensored. The system expects a messages array containing your conversation history. This endpoint supports both synchronous and asynchronous requests, making it easy to integrate into existing workflows without changing your client library.
curl https://api.gpt5apikey.com/v1/chat/completions \
-H "Authorization: Bearer $API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "uncensored",
"messages": [{"role": "user", "content": "Write a blunt product review of a cheap VPN."}]
}'
The response returns a JSON object containing the generated text, token usage statistics, and finish reasons. This structure mirrors the standard OpenAI response format, allowing you to use any compatible SDK or HTTP client without custom parsing logic.
Python SDK Integration
Using the official Python SDK is straightforward. Set the base_url to our endpoint and provide your API key. The library handles serialization and error handling automatically. You can pass your entire conversation history as a list of message dictionaries. The API processes the input and returns the model's completion.
from openai import OpenAI
client = OpenAI(base_url="https://api.gpt5apikey.com/v1", api_key="YOUR_KEY")
resp = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Summarise this thread without softening it."}],
)
print(resp.choices[0].message.content)
This approach works identically for simple prompts and multi-turn conversations. The SDK manages the HTTP connection pool efficiently, ensuring minimal latency for subsequent requests. You can also pass additional parameters like temperature or max_tokens to control the output behavior.
Node SDK Integration
For JavaScript developers, the Node SDK follows the same pattern. Initialize the client with your base URL and key. The API supports standard tool definitions, allowing you to define functions that the model can call. This is useful for generating structured data or triggering external actions based on user input.
import OpenAI from "openai";
const client = new OpenAI({ baseURL: "https://api.gpt5apikey.com/v1", apiKey: process.env.API_KEY });
const resp = await client.chat.completions.create({
model: "uncensored",
messages: [{ role: "user", content: "Draft a villain monologue for my game." }],
});
console.log(resp.choices[0].message.content);
The Node environment handles asynchronous operations naturally. You can chain requests or run them in parallel depending on your application's needs. The SDK provides type definitions if you are using TypeScript, ensuring better developer experience during development.
Support for Streaming Responses
Enable streaming by setting stream: true in your request. The API returns a Server-Sent Events (SSE) stream. Each chunk contains a partial response, allowing you to display text to the user as it is generated. This improves perceived latency, especially for long responses. The streaming response includes token usage data in the final chunk.
stream = client.chat.completions.create(
model="uncensored",
messages=[{"role": "user", "content": "Tell the story in second person."}],
stream=True,
)
for chunk in stream:
if chunk.choices and chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="", flush=True)
Handle the stream by reading chunks until the response ends. Error handling is crucial; if the connection drops, you may need to retry or handle the partial output. The streaming format is compatible with most modern HTTP clients and SDKs that support SSE.
Rate Limits, Errors, and Context
Each API key is limited to 300 requests per minute. If you exceed this, you will receive a 429 status code. The request body must not exceed 8 MB. Authentication errors return a 401 status if the key is invalid or expired. If your prepaid credit is exhausted, you receive a 402 status. The model supports a context window of 100,000 tokens, including both input and output. Plan your prompt size accordingly to avoid truncation. The model does not refuse lawful adult or controversial topics, providing a true uncensored experience.
What the API supports
Everything the endpoint can and cannot do, in one place — check it before you top up.
| Item | Value |
|---|---|
| API format | OpenAI Chat Completions schema; official openai SDKs work unchanged |
| Model | uncensored |
| Authentication | Authorization: Bearer YOUR_KEY |
| Base URL | https://api.gpt5apikey.com/v1 |
| Endpoints | POST /v1/chat/completions · GET /v1/models |
| JSON mode | response_format: {"type": "json_object"} |
| Streaming | Supported (stream: true), usage included at the end |
| Tools / tool calls | Yes — tools, tool_choice; replies carry tool_calls, also when streaming; send results back as role: tool |
| Context window | 100,000 tokens, input and output combined |
| Sampling parameters | temperature, top_p, stop, seed and the two penalties are passed through |
| Max output | 16,000 tokens max; 2,048 if max_tokens is not set |
| Max body | up to 8 MB per request |
| Concurrency | up to 8 in parallel per key |
| Response headers | X-Request-Id, X-Balance-USD, X-RateLimit-Limit-Requests, X-RateLimit-Limit-Concurrency |
| Rate limit | 300/min per key |
| Payment | USDT (TRC20) or USDC (Base), any whole amount from $10 to $500 |
| Token prices | input $0.25 / 1M tokens, output $1.00 / 1M tokens |
| Subscription | no monthly fee; paid credit does not expire |
| Free trial | $0.50 for 7 days, no card |
| Volume bonus | +5% from $50, +10% from $100 |
| Billing | pay as you go from prepaid credit; nothing is charged for failed or refused requests |
| Content | uncensored for adults; the only hard rule: no sexual content involving minors |
| Sign-in | sign in with Google or with e-mail + password |
| Key management | one active key per account; a new key replaces the old one |
Error codes
Every error is JSON with a type you can switch on. You are never charged for an error.
| HTTP | Type | What to do |
|---|---|---|
400 | bad_request | invalid JSON, empty messages, bad parameter, or prompt + max_tokens over the window — fix and resend |
401 | missing_key · invalid_key · key_revoked | no key, wrong key, or a key replaced by a newer one |
402 | no_credit | out of credit; add credit and retry |
403 | content_blocked | refused by the content policy |
404 | not_found | unknown endpoint |
413 | request_too_large | request body larger than 8 MB |
429 | rate_limited · concurrency | slow down: rate or parallel limit reached |
503 | upstream_busy | model busy — retry in a few seconds |
Questions and answers
What is the model ID for the uncensored model?
The model ID is "uncensored". You must specify this in your requests to use our tuned open-weight model.
Does the API support function calling?
Yes, the endpoint supports tool and function calling. You can define functions in your request, and the model will return structured data to call them.
What happens if my prepaid credit runs out?
You will receive a 402 error. You can top up your account with a minimum of $10 via crypto (USDT or USDC). Credits never expire.
Your key is one form away
Create an account, copy the key, change the base URL. That is the whole setup.