AI Gateway
One API, any provider. Use the OpenAI SDK you already know — 2kw.ai routes to the right backend.
Overview
The AI Gateway exposes OpenAI-compatible endpoints (/v1/chat/completions, /v1/audio/transcriptions, /v1/models) and routes requests to whatever provider you've configured. Swap between OpenAI, Azure OpenAI, Anthropic, xAI, Mistral, or a self-hosted Ollama server without touching your code.
Why use it?
- No vendor lock-in — switch providers by changing the model string, not your codebase
- Standard API — any OpenAI-compatible SDK, tool, or framework just works
- Streaming — full SSE support for real-time responses
- Centralized credentials — manage API keys and routing per-organization
Models
Platform Models
2kw.ai comes with pre-configured models available on every tier. Use them by name — no provider prefix required.
Checking available models
Use the GET /v1/models endpoint to see all models currently available to your organization, including both platform and BYOK models.
Bring Your Own Key (BYOK)
On the Team, Business and Enterprise plans, you can connect your own provider accounts and use the provider/model format:
openai/gpt-5.1
azure-openai/gpt-4
anthropic/claude-sonnet-4-5-20250929
xai/grok-4
mistral/mistral-large-latest
ollama/llama3
The prefix tells the gateway where to route. The model name after the slash is passed directly to the provider. A BYOK request is billed through your own provider account, and its model charge does not count against your plan's usage.
BYOK availability
Bring Your Own Key requires the Team, Business or Enterprise plan. Configure your providers under Providers in the sidebar (section Build); see Providers & BYOK for provider types, test connection and the /v1/providers API.
Chat Completions
POST /v1/chat/completions
The main endpoint. Drop-in replacement for https://api.openai.com/v1/chat/completions.
Request
curl -X POST https://api.2kw.ai/v1/chat/completions \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk_your_api_key" \
-d '{
"model": "gpt-5.1",
"messages": [
{"role": "system", "content": "You are a helpful assistant."},
{"role": "user", "content": "What is Spring Boot?"}
]
}'
Request Parameters
| Field | Type | Required | Description |
|---|---|---|---|
model | string | Yes | Platform model name (e.g., gpt-5.1) or provider/model for BYOK |
messages | array | Yes | Chat messages (role + content). content is a string, or an array of text and image_url parts — see Images in messages |
stream | boolean | No | Enable SSE streaming (default: false) |
temperature | number | No | Sampling temperature (0-2) |
max_tokens | number | No | Max tokens in response |
top_p | number | No | Nucleus sampling (0-1) |
frequency_penalty | number | No | Frequency penalty (-2 to 2) |
presence_penalty | number | No | Presence penalty (-2 to 2) |
stop | array | No | Stop sequences |
response_format | object | No | Output format — e.g. {"type": "json_object"} for JSON mode |
tools | array | No | Tool definitions for function calling |
tool_choice | string/object | No | Control tool selection: "auto", "none", or a specific tool |
max_completion_tokens | number | No | Max completion tokens (newer alternative to max_tokens) |
parallel_tool_calls | boolean | No | Accepted for OpenAI compatibility but not forwarded to the provider on this endpoint; the provider's own default applies |
prompt_cache_key | string | No | Up to 256 characters. Groups requests that share a long prompt prefix so the provider's prompt cache can serve them together. On the built-in models the provider receives a hash scoped to your organization, never the value itself. Longer than 256 characters is a 400 |
stream_options | object | No | Streaming options. {"include_usage": true} appends a final chunk with the call's token usage — see Streaming |
Response
{
"id": "chatcmpl-123",
"object": "chat.completion",
"model": "gpt-5.1",
"choices": [{
"index": 0,
"message": {
"role": "assistant",
"content": "Spring Boot is a Java-based framework..."
},
"finish_reason": "stop"
}],
"usage": {
"prompt_tokens": 25,
"completion_tokens": 150,
"total_tokens": 175,
"prompt_tokens_details": { "cached_tokens": 0 }
}
}
usage.prompt_tokens_details.cached_tokens is how many of the prompt_tokens the provider served from its prompt cache. Cached input is billed at the cache-read price. The object is left out when the share is not known, so a missing prompt_tokens_details does not mean zero. On the built-in models the field is reported only while the organization-scoped prompt_cache_key is sent to the provider; calls on your own provider (BYOK) always report it when the provider does.
Images in Messages
A message's content can be an array of parts instead of a string. Text parts and image parts mix freely, in the same shape the OpenAI API uses:
{
"model": "gpt-5.1",
"messages": [
{
"role": "user",
"content": [
{ "type": "text", "text": "What does this label say, and is the part within tolerance?" },
{ "type": "image_url", "image_url": { "url": "https://example.com/label.jpg" } }
]
}
]
}
image_url.url takes an https URL or a data: URL with base64 content; the provider you route to decides which image types and sizes it accepts. Pick a model that supports vision, or the provider rejects the request. Calls that carry an image are recorded under the VISION surface in analytics, so image traffic is visible next to text traffic in the analytics dashboard.
Responses are always plain text; the model does not return image parts.
Streaming
Set "stream": true and you'll get Server-Sent Events:
Request
from openai import OpenAI
client = OpenAI(
api_key="sk_your_api_key",
base_url="https://api.2kw.ai/v1"
)
stream = client.chat.completions.create(
model="gpt-5.1",
messages=[{"role": "user", "content": "Write a haiku about code"}],
stream=True
)
for chunk in stream:
if chunk.choices[0].delta.content:
print(chunk.choices[0].delta.content, end="")
SSE format
Streaming responses follow the standard OpenAI SSE format — data: {...} lines terminated with data: [DONE].
A stream carries no token counts by default. Send "stream_options": {"include_usage": true} and the last chunk before data: [DONE] has an empty choices array and a usage object in the shape of the non-streaming response, prompt_tokens_details.cached_tokens included. Every other chunk omits usage. When the provider reports no token counts for the stream, that chunk is not sent, so do not wait for it.
List Models
GET /v1/models
Lists all models available to your organization — both platform models and BYOK providers you've configured:
curl https://api.2kw.ai/v1/models \
-H "Authorization: Bearer sk_your_api_key"
Responses API
POST /v1/responses
The gateway also serves the OpenAI Responses API. With model: "provider/model" it is a direct call with the same routing as chat completions; with model: "agent/{name}" it runs a stored agent with server-side tools, knowledge search, citations and approvals. The request and response shapes, conversations and the continuation protocol are documented on the Agents page. Set "stream": true for Server-Sent Events in the Open Responses format; the same page lists the events a stream carries.
Embeddings
POST /v1/embeddings
Drop-in replacement for https://api.openai.com/v1/embeddings. There is no CLI or MCP command for it; call it from an OpenAI SDK, LangChain or any OpenAI-compatible client.
Request
curl -X POST https://api.2kw.ai/v1/embeddings \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk_your_api_key" \
-d '{
"model": "text-embedding-3-small",
"input": ["First text", "Second text"]
}'
The response is the OpenAI shape: object: "list", one data[] entry per input in input order with its index and embedding, the model you sent, and a usage with prompt_tokens and total_tokens.
| Field | Description |
|---|---|
model | Required. A bare name runs on the built-in models; provider/model runs on your own provider (see below). |
input | Required. A string, or an array of strings. |
dimensions | Optional. The width of each vector, for models that can be shortened. |
encoding_format | Optional. float (default), or base64: each vector as little-endian float32, base64-encoded. The OpenAI SDKs decode it for you. |
user | Optional. An identifier of your end user, recorded on the trace. |
Two routes
model | Runs on | Cost |
|---|---|---|
text-embedding-3-small, text-embedding-3-large (bare name) | The platform's built-in embedding models | Charged against your plan's usage: €0.0825 per million input tokens on text-embedding-3-small, €0.53625 on text-embedding-3-large (see the rate card). Each call is charged for the tokens it used. |
openai/text-embedding-3-small, azure-openai/<deployment> (provider/model) | Your own provider | Not charged by 2kw.ai; your provider bills you. Needs the Team, Business or Enterprise plan, like any other BYOK call. Counts toward your concurrent-call limit. |
Only the openai and azure-openai providers serve embeddings; any other prefix answers 400 with code model_not_supported. The model in the response is the one you sent.
Where the data goes follows the route: the built-in models run on the platform's Azure OpenAI deployments in the EU Data Zone, and provider/model sends your input to the provider you configured.
Limits
- At most 2,048 inputs per request.
- Every input is a non-empty string.
- At most 1 MiB of UTF-8 in total across the inputs.
- OpenAI's own token limits per input are the provider's, and a request over them is answered with the provider's
400.
Arrays of token ids are not supported. LangChain's OpenAIEmbeddings sends token ids by default and gets a 400 until you set check_embedding_ctx_length=False (see LangChain).
Errors
Errors use the gateway's OpenAI-style envelope {"error": {"message", "type", "param", "code"}}.
| Case | Status | type / code |
|---|---|---|
Accept header that rules out JSON | 406 | invalid_request_error / not_acceptable |
Malformed body, input of another type, token-id array, unknown encoding_format | 400 | invalid_request_error |
| More than 2,048 inputs, an empty input, more than 1 MiB | 400 | invalid_request_error / too_many_inputs, empty_input, input_too_large |
A bare name that is not a built-in embedding model, or a malformed provider/model | 400 | invalid_request_error |
| A provider that serves no embeddings | 400 | invalid_request_error / model_not_supported |
| No provider configured for the prefix | 404 | invalid_request_error |
| Viewer role | 403 | invalid_request_error / forbidden |
| Usage used up with no extra usage available, or a plan without BYOK | 402 | billing_error |
| Too many concurrent calls | 429, with Retry-After | rate_limit_error |
The provider refused the request (over-long input, bad dimensions) | 400 | upstream_error |
| The provider failed | 502 | upstream_error |
Two answers come from the security layer in front of every /v1 endpoint and are not in this envelope: a request with no credentials at all gets a bare 401, and a rejected surface assertion gets a flat {"error": "<code>", "message": "..."}.
Framework Integrations
The gateway works with anything that speaks OpenAI — just point base_url at your 2kw.ai instance.
LangChain
from langchain_openai import ChatOpenAI
llm = ChatOpenAI(
api_key="sk_your_api_key",
base_url="https://api.2kw.ai/v1",
model_name="gpt-5.1"
)
response = llm.invoke("What is the capital of France?")
print(response.content)
For embeddings, turn off LangChain's token-id chunking, which the gateway does not accept:
from langchain_openai import OpenAIEmbeddings
embeddings = OpenAIEmbeddings(
api_key="sk_your_api_key",
base_url="https://api.2kw.ai/v1",
model="text-embedding-3-small",
check_embedding_ctx_length=False
)
vectors = embeddings.embed_documents(["First text", "Second text"])
Works with any OpenAI-compatible tool
LiteLLM, LlamaIndex, Haystack, Semantic Kernel — if it has a base_url setting, it works with 2kw.ai.
Supported BYOK Providers
Connect your own accounts from any of these providers (Team, Business and Enterprise plans):
| Provider | Prefix | Example Models |
|---|---|---|
| OpenAI | openai | gpt-5.1, gpt-5-mini, gpt-4.1 |
| Azure OpenAI | azure-openai | Your Azure deployment names |
| Anthropic | anthropic | claude-sonnet-4-5, claude-opus-4-5, claude-haiku-4-5 |
| xAI | xai | grok-4 |
| Mistral | mistral | mistral-large-latest, mistral-small-latest |
| Ollama | ollama | llama3, mistral, codellama (models pulled on your own server) |
The name after the prefix is the provider's own model id, passed through unchanged. Platform model names are not provider ids: the platform's gpt-5.1-mini, for example, is OpenAI's gpt-5-mini, so a BYOK call names it openai/gpt-5-mini.
A few request features depend on the provider:
- Mistral takes
tool_choiceas"auto","none"or"required". A specific named function is not passed on, so the model chooses among the tools itself. - Images by URL must use
httpsfor every provider; a plainhttpimage URL is refused with400. OpenAI, Azure OpenAI, Anthropic, xAI and Mistral fetch anhttpsimage themselves. - Ollama accepts images only inline, as base64
data:URLs. Anhttpsimage URL sent to an Ollama model is refused with400.
Google Vertex AI is not supported. The vertex-ai prefix is recognized, and appears in the list of valid prefixes an unknown prefix returns, but no call routed to it can succeed.
Configure providers under Providers in the sidebar (section Build) or through the API, see Providers & BYOK.
Provider Configuration
POST /v1/providers
Connects a provider account to your organization. The Providers page in the console makes the same call. Providers & BYOK documents every provider endpoint in full.
curl -X POST https://api.2kw.ai/v1/providers \
-H "Content-Type: application/json" \
-H "Authorization: Bearer sk_your_api_key" \
-d '{
"name": "Production Azure",
"provider": "AZURE_OPENAI",
"apiKey": "your-azure-openai-key",
"config": { "endpoint": "https://my-resource.openai.azure.com" }
}'
| Field | Type | Required | Description |
|---|---|---|---|
name | string | Yes | Display name, up to 255 characters |
provider | string | Yes | OPENAI, AZURE_OPENAI, ANTHROPIC, XAI, MISTRAL or OLLAMA |
apiKey | string | Yes | The provider's API key, a top-level field and never part of config. Stored encrypted and never returned. Required for every provider except OLLAMA, which needs no key: leave it out |
config | object | Yes | Provider-specific settings from the table below. Send {} when none apply |
The API also lists VERTEX_AI as a provider type, but a Vertex AI provider cannot be configured or serve requests.
Config keys
| Provider | Key | Required | Default | Notes |
|---|---|---|---|---|
| OpenAI | baseUrl | No | https://api.openai.com/v1 | For an OpenAI-compatible endpoint. Starts with http:// or https:// |
| OpenAI | organizationId | No | Your OpenAI organization ID | |
| OpenAI | useResponsesApi | No | false | true sends every request through OpenAI's Responses API. See below |
| Azure OpenAI | endpoint | Yes | Your resource URL, such as https://my-resource.openai.azure.com. Must start with https:// | |
| Azure OpenAI | apiVersion | No | 2024-08-01-preview | Used only to list your deployments |
| Anthropic | baseUrl | No | https://api.anthropic.com | Starts with http:// or https:// |
| Anthropic | version | No | 2023-06-01 | Sent as the anthropic-version header |
| xAI | baseUrl | No | https://api.x.ai/v1 | Starts with http:// or https:// |
| Mistral | baseUrl | No | https://api.mistral.ai | Starts with http:// or https:// |
| Ollama | baseUrl | Yes | Your Ollama server, such as http://ollama.internal:11434. Starts with http:// or https:// |
Azure OpenAI has no deployment setting: the deployment name is the model name, as in azure-openai/my-gpt-4o-deployment.
useResponsesApi. You rarely need it. A request that carries function tools for a reasoning model is sent through the Responses API automatically, because chat completions refuses that combination, and so is a model that only exists on the Responses API. Setting the flag forces the Responses API for every request on this provider, tools or not. Streaming works on both. The flag takes a boolean or the string "true" or "false"; the console calls it Use the Responses API.
Managing providers
| Method and path | Purpose |
|---|---|
GET /v1/providers | List your providers. apiKey is never included |
GET /v1/providers/{id} | Read one provider |
PATCH /v1/providers/{id} | Change name, apiKey or config. Every field is optional; the provider type cannot change |
DELETE /v1/providers/{id} | Remove a provider |
POST /v1/providers/test | Check provider, apiKey and config before saving. Answers success, a message and the models found |
GET /v1/providers/models | List the models available across all your providers |
The CLI equivalents are under backbone providers on the CLI page.