Integrate once. Govern centrally. Route anywhere.
REDtone AI Gateway is the enterprise AI governance layer between your applications and every model you use — cloud providers, LiteLLM, or a GPU in your own rack. Bring your own AI backend. One OpenAI-compatible endpoint decides who may call, whether a request may leave the network, what it cost, and keeps the evidence.

How it works
Northbound, an OpenAI-compatible API your applications already speak. Southbound, any provider, gateway or runtime. In the middle, the controls REDtone owns.

The product is the control planes
Backends change every quarter. These do not. Each one is a screen in the portal and a row in a table you can audit.

Identity & Access
Virtual API keys per tenant and application, hashed at rest and shown once. Per-key request and token rate limits. REDtone ID sign-in for the portal.

Governance & Routing
Every tenant and key carries a data classification. Restricted requests reach internal backends only. Fail-closed: no permitted backend, no call.

FinOps & Quota
Token usage read from the backend, cost from a price table per backend and model, monthly budgets with 80% and 100% alerts. Usage by tenant, key, model and day.

Observability & Audit
A placement log for every request: which backend, which region, and why. That log is the residency evidence. Prometheus metrics and an admin audit trail.
Bring your own backend, or use REDtone's
Provider-agnostic and gateway-agnostic. Connect your own with a name, a base URL and a key, or start on the models REDtone already runs. Either way, the same policies apply.

LiteLLM
Your existing LiteLLM proxy becomes one backend. It keeps the provider keys and the protocol translation; the gateway keeps the governance.
How to connectOpenAI-compatible
vLLM, Ollama, SGLang, a vendor gateway, a partner's API — anything that answers /v1/chat/completions plugs in unchanged.
How to connectPrivate & local models
An on-prem inference server marked as an internal location. Confidential and restricted traffic routes here and nowhere else.
How to connectREDtone backend
No provider accounts or keys to manage. Use the models REDtone already connects, billed through your tenant and governed like the rest.
See available modelsAvailable models
Leading AI brands behind one endpoint, governed by REDtone. Call them by name, or by a REDtone alias.
OpenAI GPT-6
Frontier chat, from a low-cost everyday model to the most capable one, plus embeddings.
gpt-6-luna · gpt-6.1-sol · gpt-6-astra · +1 more
Google Gemini 3.8
Fast multimodal chat and natural-sounding speech.
gemini-3.8-flash · gemini-3.8-flash-tts
DeepSeek V4
Strong reasoning and coding at a low price.
deepseek-v4-flash · deepseek-v4-pro
Qwen 3.5
Alibaba's open-weight chat models, strong in English and Chinese.
qwen3.5 · qwen3.5-27b
Anthropic Claude
Careful, capable assistants for writing, analysis and code.
claude-opus-5-5 · claude-sonnet-5-5 · claude-haiku-4-5
Change two lines. Keep your code.
Every OpenAI SDK, LangChain and agent framework works as before. Migrating from LiteLLM? The model names can stay.
- 1
Get a tenant
Your REDtone contact creates the tenant and invites the first owner. Members, keys and budgets are self-service after that. Contact REDtone
- 2
Create an API key
Sign in with REDtone ID and create a key for each application. It is shown once and can be revoked any time. Sign in to get a key
- 3
Point your code at the gateway
Set the base URL and swap the key. Call a model by its own name or a REDtone alias. More samples
from openai import OpenAIclient = OpenAI(− api_key="sk-...",+ base_url="https://ai.redtone-dev.com/v1",+ api_key="rt-...",)reply = client.chat.completions.create(model="redtone/chat-default",messages=[{"role": "user", "content": "Hello from REDtone"}],)