Integrate once. Govern centrally. Route anywhere.

REDtone AI Gateway is the enterprise AI governance layer between your applications and every model you use — cloud providers, LiteLLM, or a GPU in your own rack. Bring your own AI backend. One OpenAI-compatible endpoint decides who may call, whether a request may leave the network, what it cost, and keeps the evidence.

Applications connect through the REDtone gateway to cloud, on-premise and GPU backends; one path is blocked by policy

How it works

Northbound, an OpenAI-compatible API your applications already speak. Southbound, any provider, gateway or runtime. In the middle, the controls REDtone owns.

REDtone AI Gateway: applications on top, the gateway's control planes in the middle, AI providers and infrastructure below; all traffic passes through the gateway
Northbound. OpenAI SDKs, LangChain, n8n, chatbots and copilots — with a gateway key instead of a provider key.
Gateway. Tenant, project, user, policy, budget, placement, audit and compliance in one Go core.
Southbound. Commercial providers, LiteLLM, vLLM and Ollama, private models on-prem, and specialised services.

The product is the control planes

Backends change every quarter. These do not. Each one is a screen in the portal and a row in a table you can audit.

Key card, key and padlock

Identity & Access

Virtual API keys per tenant and application, hashed at rest and shown once. Per-key request and token rate limits. REDtone ID sign-in for the portal.

A shield routes traffic to an internal data centre and blocks the path to the cloud

Governance & Routing

Every tenant and key carries a data classification. Restricted requests reach internal backends only. Fail-closed: no permitted backend, no call.

Coin stacks beside a budget gauge

FinOps & Quota

Token usage read from the backend, cost from a price table per backend and model, monthly budgets with 80% and 100% alerts. Usage by tenant, key, model and day.

Dashboard, log entries and a magnifying glass

Observability & Audit

A placement log for every request: which backend, which region, and why. That log is the residency evidence. Prometheus metrics and an admin audit trail.

Bring your own backend, or use REDtone's

Provider-agnostic and gateway-agnostic. Connect your own with a name, a base URL and a key, or start on the models REDtone already runs. Either way, the same policies apply.

One gateway hub plugged into a server rack, a cloud and a GPU server

LiteLLM

Your existing LiteLLM proxy becomes one backend. It keeps the provider keys and the protocol translation; the gateway keeps the governance.

How to connect

OpenAI-compatible

vLLM, Ollama, SGLang, a vendor gateway, a partner's API — anything that answers /v1/chat/completions plugs in unchanged.

How to connect

Private & local models

An on-prem inference server marked as an internal location. Confidential and restricted traffic routes here and nowhere else.

How to connect

REDtone backend

No provider accounts or keys to manage. Use the models REDtone already connects, billed through your tenant and governed like the rest.

See available models

Available models

Leading AI brands behind one endpoint, governed by REDtone. Call them by name, or by a REDtone alias.

OpenAI GPT-6

Frontier chat, from a low-cost everyday model to the most capable one, plus embeddings.

gpt-6-luna · gpt-6.1-sol · gpt-6-astra · +1 more

Google Gemini 3.8

Fast multimodal chat and natural-sounding speech.

gemini-3.8-flash · gemini-3.8-flash-tts

DeepSeek V4

Strong reasoning and coding at a low price.

deepseek-v4-flash · deepseek-v4-pro

Qwen 3.5

Alibaba's open-weight chat models, strong in English and Chinese.

qwen3.5 · qwen3.5-27b

Anthropic Claude

Careful, capable assistants for writing, analysis and code.

claude-opus-5-5 · claude-sonnet-5-5 · claude-haiku-4-5

Zhipu GLM

Bilingual chat and agent models from Zhipu AI.

Change two lines. Keep your code.

Every OpenAI SDK, LangChain and agent framework works as before. Migrating from LiteLLM? The model names can stay.

  1. 1

    Get a tenant

    Your REDtone contact creates the tenant and invites the first owner. Members, keys and budgets are self-service after that. Contact REDtone

  2. 2

    Create an API key

    Sign in with REDtone ID and create a key for each application. It is shown once and can be revoked any time. Sign in to get a key

  3. 3

    Point your code at the gateway

    Set the base URL and swap the key. Call a model by its own name or a REDtone alias. More samples

app.py · OpenAI SDK
from openai import OpenAI
client = OpenAI(
− api_key="sk-...",
+ base_url="https://ai.redtone-dev.com/v1",
+ api_key="rt-...",
)
reply = client.chat.completions.create(
model="redtone/chat-default",
messages=[{"role": "user", "content": "Hello from REDtone"}],
)