Developer docs

DocsEverything an application needs to call the gateway. Admin screens are documented in the portal itself.
Authentication

Every request carries a gateway key as a bearer token. Keys start with rt-, belong to one tenant, and are created in the portal under API Keys. The gateway stores only a hash: a lost key is replaced, not recovered.

Authorization: Bearer rt-xxxxxxxx-................................

Each key has a requests-per-minute and tokens-per-minute limit. Over either, the answer is 429 with a Retry-After header.

Endpoints

Base URL: https://ai.redtone-dev.com/v1. The surface is OpenAI's; the request and response bodies pass through unchanged apart from model, which the gateway rewrites from the alias to the backend's own model name.

POST /v1/chat/completionsChat. Streaming and non-streaming.
POST /v1/completionsLegacy text completion, when the backend supports it.
POST /v1/embeddingsEmbeddings.
GET /v1/modelsThe aliases your key may call, in OpenAI's list shape.
POST /v1/*Any other OpenAI-shaped path with a model field is forwarded to that alias's backend, logged but not token-metered.

Every response carries X-Request-Id, X-RT-Backend and X-RT-Location, so a support question can be matched to a placement-log row.

Data classification

Every tenant has a classification: public internal confidential restricted. A key may be created with a higher one. Each backend declares which classifications it accepts; a request only ever reaches a backend that accepts its class, and an internal-only backend is where confidential and restricted traffic goes.

A single request can raise its class with a header. It can never lower it.

curl https://ai.redtone-dev.com/v1/chat/completions \
  -H "Authorization: Bearer $REDTONE_API_KEY" \
  -H "X-Data-Classification: restricted" \
  -H "Content-Type: application/json" \
  -d '{"model": "redtone/chat-default", "messages": [{"role":"user","content":"..."}]}'

Fail-closed: if no backend for the alias accepts the class, the call is refused with 403 policy_denied and logged as such. Nothing is sent anywhere.

Streaming

Set "stream": true. Server-sent events are relayed as they arrive. The gateway asks the backend to include a final usage chunk (stream_options.include_usage) so streamed calls are metered exactly like the rest; you will see it as the last chunk before [DONE].

Errors

OpenAI's error envelope, with a stable code.

401invalid_api_keyMissing, unknown, revoked or expired key.
403policy_deniedNo backend for this alias accepts the request's classification.
404model_not_foundThe alias does not exist or is not visible to your tenant.
429rate_limitedOver the key's rpm or tpm limit. Retry after the header says.
502upstream_errorThe backend answered with an error; its body is passed through.
503no_healthy_backendEvery candidate backend is down. Retry with backoff.
{
  "error": {
    "message": "no backend for alias redtone/chat-default accepts classification restricted",
    "type": "policy_denied",
    "code": "policy_denied"
  }
}
Migrating from LiteLLM

Two settings change. Your code does not. LiteLLM stays where it is, behind the gateway as a backend; the gateway adds the tenant, the key, the classification check and the metering in front of it.

# before — LiteLLM directly
OPENAI_BASE_URL=https://llm.redtone.com/v1
OPENAI_API_KEY=sk-...

# after — through the gateway
OPENAI_BASE_URL=https://ai.redtone-dev.com/v1
OPENAI_API_KEY=rt-...
# model names can stay: an alias may equal the upstream name, e.g. openai/gpt-5
Code samples
curl
curl https://ai.redtone-dev.com/v1/chat/completions \
  -H "Authorization: Bearer $REDTONE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "redtone/chat-default",
    "messages": [{"role": "user", "content": "Summarise this in one line: ..."}]
  }'
Python · openai
from openai import OpenAI

client = OpenAI(
    base_url="https://ai.redtone-dev.com/v1",
    api_key="rt-...",            # a gateway key, not a provider key
)

stream = client.chat.completions.create(
    model="redtone/chat-default",
    messages=[{"role": "user", "content": "Hello"}],
    stream=True,
)
for chunk in stream:
    print(chunk.choices[0].delta.content or "", end="")
Node · openai
import OpenAI from "openai";

const client = new OpenAI({ baseURL: "https://ai.redtone-dev.com/v1", apiKey: process.env.REDTONE_API_KEY });

const r = await client.chat.completions.create({
  model: "redtone/chat-default",
  messages: [{ role: "user", content: "Hello" }],
});
console.log(r.choices[0].message.content);
Embeddings
curl https://ai.redtone-dev.com/v1/embeddings \
  -H "Authorization: Bearer $REDTONE_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{"model": "redtone/embed-default", "input": "The quick brown fox"}'