AI Gateway

One API, any provider. Use the OpenAI SDK you already know — 2kw.ai routes to the right backend.

Overview

The AI Gateway exposes OpenAI-compatible endpoints (/v1/chat/completions, /v1/audio/transcriptions, /v1/models) and routes requests to whatever provider you've configured. Swap between OpenAI, Azure OpenAI, Anthropic, xAI, Mistral, or a self-hosted Ollama server without touching your code.

Why use it?

  • No vendor lock-in — switch providers by changing the model string, not your codebase
  • Standard API — any OpenAI-compatible SDK, tool, or framework just works
  • Streaming — full SSE support for real-time responses
  • Centralized credentials — manage API keys and routing per-organization

Models

Platform Models

2kw.ai comes with pre-configured models available on every tier. Use them by name — no provider prefix required.

Bring Your Own Key (BYOK)

On the Team, Business and Enterprise plans, you can connect your own provider accounts and use the provider/model format:

openai/gpt-5.1
azure-openai/gpt-4
anthropic/claude-sonnet-4-5-20250929
xai/grok-4
mistral/mistral-large-latest
ollama/llama3

The prefix tells the gateway where to route. The model name after the slash is passed directly to the provider. A BYOK request is billed through your own provider account, and its model charge does not count against your plan's usage.

Chat Completions

POST /v1/chat/completions

The main endpoint. Drop-in replacement for https://api.openai.com/v1/chat/completions.

Request

curl -X POST https://api.2kw.ai/v1/chat/completions \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer sk_your_api_key" \
  -d '{
    "model": "gpt-5.1",
    "messages": [
      {"role": "system", "content": "You are a helpful assistant."},
      {"role": "user", "content": "What is Spring Boot?"}
    ]
  }'

Request Parameters

FieldTypeRequiredDescription
modelstringYesPlatform model name (e.g., gpt-5.1) or provider/model for BYOK
messagesarrayYesChat messages (role + content). content is a string, or an array of text and image_url parts — see Images in messages
streambooleanNoEnable SSE streaming (default: false)
temperaturenumberNoSampling temperature (0-2)
max_tokensnumberNoMax tokens in response
top_pnumberNoNucleus sampling (0-1)
frequency_penaltynumberNoFrequency penalty (-2 to 2)
presence_penaltynumberNoPresence penalty (-2 to 2)
stoparrayNoStop sequences
response_formatobjectNoOutput format — e.g. {"type": "json_object"} for JSON mode
toolsarrayNoTool definitions for function calling
tool_choicestring/objectNoControl tool selection: "auto", "none", or a specific tool
max_completion_tokensnumberNoMax completion tokens (newer alternative to max_tokens)
parallel_tool_callsbooleanNoAccepted for OpenAI compatibility but not forwarded to the provider on this endpoint; the provider's own default applies
prompt_cache_keystringNoUp to 256 characters. Groups requests that share a long prompt prefix so the provider's prompt cache can serve them together. On the built-in models the provider receives a hash scoped to your organization, never the value itself. Longer than 256 characters is a 400
stream_optionsobjectNoStreaming options. {"include_usage": true} appends a final chunk with the call's token usage — see Streaming

Response

{
  "id": "chatcmpl-123",
  "object": "chat.completion",
  "model": "gpt-5.1",
  "choices": [{
    "index": 0,
    "message": {
      "role": "assistant",
      "content": "Spring Boot is a Java-based framework..."
    },
    "finish_reason": "stop"
  }],
  "usage": {
    "prompt_tokens": 25,
    "completion_tokens": 150,
    "total_tokens": 175,
    "prompt_tokens_details": { "cached_tokens": 0 }
  }
}

usage.prompt_tokens_details.cached_tokens is how many of the prompt_tokens the provider served from its prompt cache. Cached input is billed at the cache-read price. The object is left out when the share is not known, so a missing prompt_tokens_details does not mean zero. On the built-in models the field is reported only while the organization-scoped prompt_cache_key is sent to the provider; calls on your own provider (BYOK) always report it when the provider does.

Images in Messages

A message's content can be an array of parts instead of a string. Text parts and image parts mix freely, in the same shape the OpenAI API uses:

{
  "model": "gpt-5.1",
  "messages": [
    {
      "role": "user",
      "content": [
        { "type": "text", "text": "What does this label say, and is the part within tolerance?" },
        { "type": "image_url", "image_url": { "url": "https://example.com/label.jpg" } }
      ]
    }
  ]
}

image_url.url takes an https URL or a data: URL with base64 content; the provider you route to decides which image types and sizes it accepts. Pick a model that supports vision, or the provider rejects the request. Calls that carry an image are recorded under the VISION surface in analytics, so image traffic is visible next to text traffic in the analytics dashboard.

Responses are always plain text; the model does not return image parts.

Streaming

Set "stream": true and you'll get Server-Sent Events:

Request

from openai import OpenAI

client = OpenAI(
    api_key="sk_your_api_key",
    base_url="https://api.2kw.ai/v1"
)

stream = client.chat.completions.create(
    model="gpt-5.1",
    messages=[{"role": "user", "content": "Write a haiku about code"}],
    stream=True
)

for chunk in stream:
    if chunk.choices[0].delta.content:
        print(chunk.choices[0].delta.content, end="")

A stream carries no token counts by default. Send "stream_options": {"include_usage": true} and the last chunk before data: [DONE] has an empty choices array and a usage object in the shape of the non-streaming response, prompt_tokens_details.cached_tokens included. Every other chunk omits usage. When the provider reports no token counts for the stream, that chunk is not sent, so do not wait for it.

List Models

GET /v1/models

Lists all models available to your organization — both platform models and BYOK providers you've configured:

curl https://api.2kw.ai/v1/models \
  -H "Authorization: Bearer sk_your_api_key"

Responses API

POST /v1/responses

The gateway also serves the OpenAI Responses API. With model: "provider/model" it is a direct call with the same routing as chat completions; with model: "agent/{name}" it runs a stored agent with server-side tools, knowledge search, citations and approvals. The request and response shapes, conversations and the continuation protocol are documented on the Agents page. Set "stream": true for Server-Sent Events in the Open Responses format; the same page lists the events a stream carries.

Embeddings

POST /v1/embeddings

Drop-in replacement for https://api.openai.com/v1/embeddings. There is no CLI or MCP command for it; call it from an OpenAI SDK, LangChain or any OpenAI-compatible client.

Request

curl -X POST https://api.2kw.ai/v1/embeddings \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer sk_your_api_key" \
  -d '{
    "model": "text-embedding-3-small",
    "input": ["First text", "Second text"]
  }'

The response is the OpenAI shape: object: "list", one data[] entry per input in input order with its index and embedding, the model you sent, and a usage with prompt_tokens and total_tokens.

FieldDescription
modelRequired. A bare name runs on the built-in models; provider/model runs on your own provider (see below).
inputRequired. A string, or an array of strings.
dimensionsOptional. The width of each vector, for models that can be shortened.
encoding_formatOptional. float (default), or base64: each vector as little-endian float32, base64-encoded. The OpenAI SDKs decode it for you.
userOptional. An identifier of your end user, recorded on the trace.

Two routes

modelRuns onCost
text-embedding-3-small, text-embedding-3-large (bare name)The platform's built-in embedding modelsCharged against your plan's usage: €0.0825 per million input tokens on text-embedding-3-small, €0.53625 on text-embedding-3-large (see the rate card). Each call is charged for the tokens it used.
openai/text-embedding-3-small, azure-openai/<deployment> (provider/model)Your own providerNot charged by 2kw.ai; your provider bills you. Needs the Team, Business or Enterprise plan, like any other BYOK call. Counts toward your concurrent-call limit.

Only the openai and azure-openai providers serve embeddings; any other prefix answers 400 with code model_not_supported. The model in the response is the one you sent.

Where the data goes follows the route: the built-in models run on the platform's Azure OpenAI deployments in the EU Data Zone, and provider/model sends your input to the provider you configured.

Limits

  • At most 2,048 inputs per request.
  • Every input is a non-empty string.
  • At most 1 MiB of UTF-8 in total across the inputs.
  • OpenAI's own token limits per input are the provider's, and a request over them is answered with the provider's 400.

Arrays of token ids are not supported. LangChain's OpenAIEmbeddings sends token ids by default and gets a 400 until you set check_embedding_ctx_length=False (see LangChain).

Errors

Errors use the gateway's OpenAI-style envelope {"error": {"message", "type", "param", "code"}}.

CaseStatustype / code
Accept header that rules out JSON406invalid_request_error / not_acceptable
Malformed body, input of another type, token-id array, unknown encoding_format400invalid_request_error
More than 2,048 inputs, an empty input, more than 1 MiB400invalid_request_error / too_many_inputs, empty_input, input_too_large
A bare name that is not a built-in embedding model, or a malformed provider/model400invalid_request_error
A provider that serves no embeddings400invalid_request_error / model_not_supported
No provider configured for the prefix404invalid_request_error
Viewer role403invalid_request_error / forbidden
Usage used up with no extra usage available, or a plan without BYOK402billing_error
Too many concurrent calls429, with Retry-Afterrate_limit_error
The provider refused the request (over-long input, bad dimensions)400upstream_error
The provider failed502upstream_error

Two answers come from the security layer in front of every /v1 endpoint and are not in this envelope: a request with no credentials at all gets a bare 401, and a rejected surface assertion gets a flat {"error": "<code>", "message": "..."}.

Framework Integrations

The gateway works with anything that speaks OpenAI — just point base_url at your 2kw.ai instance.

LangChain

from langchain_openai import ChatOpenAI

llm = ChatOpenAI(
    api_key="sk_your_api_key",
    base_url="https://api.2kw.ai/v1",
    model_name="gpt-5.1"
)

response = llm.invoke("What is the capital of France?")
print(response.content)

For embeddings, turn off LangChain's token-id chunking, which the gateway does not accept:

from langchain_openai import OpenAIEmbeddings

embeddings = OpenAIEmbeddings(
    api_key="sk_your_api_key",
    base_url="https://api.2kw.ai/v1",
    model="text-embedding-3-small",
    check_embedding_ctx_length=False
)

vectors = embeddings.embed_documents(["First text", "Second text"])

Supported BYOK Providers

Connect your own accounts from any of these providers (Team, Business and Enterprise plans):

ProviderPrefixExample Models
OpenAIopenaigpt-5.1, gpt-5-mini, gpt-4.1
Azure OpenAIazure-openaiYour Azure deployment names
Anthropicanthropicclaude-sonnet-4-5, claude-opus-4-5, claude-haiku-4-5
xAIxaigrok-4
Mistralmistralmistral-large-latest, mistral-small-latest
Ollamaollamallama3, mistral, codellama (models pulled on your own server)

The name after the prefix is the provider's own model id, passed through unchanged. Platform model names are not provider ids: the platform's gpt-5.1-mini, for example, is OpenAI's gpt-5-mini, so a BYOK call names it openai/gpt-5-mini.

A few request features depend on the provider:

  • Mistral takes tool_choice as "auto", "none" or "required". A specific named function is not passed on, so the model chooses among the tools itself.
  • Images by URL must use https for every provider; a plain http image URL is refused with 400. OpenAI, Azure OpenAI, Anthropic, xAI and Mistral fetch an https image themselves.
  • Ollama accepts images only inline, as base64 data: URLs. An https image URL sent to an Ollama model is refused with 400.

Google Vertex AI is not supported. The vertex-ai prefix is recognized, and appears in the list of valid prefixes an unknown prefix returns, but no call routed to it can succeed.

Configure providers under Providers in the sidebar (section Build) or through the API, see Providers & BYOK.

Provider Configuration

POST /v1/providers

Connects a provider account to your organization. The Providers page in the console makes the same call. Providers & BYOK documents every provider endpoint in full.

curl -X POST https://api.2kw.ai/v1/providers \
  -H "Content-Type: application/json" \
  -H "Authorization: Bearer sk_your_api_key" \
  -d '{
    "name": "Production Azure",
    "provider": "AZURE_OPENAI",
    "apiKey": "your-azure-openai-key",
    "config": { "endpoint": "https://my-resource.openai.azure.com" }
  }'
FieldTypeRequiredDescription
namestringYesDisplay name, up to 255 characters
providerstringYesOPENAI, AZURE_OPENAI, ANTHROPIC, XAI, MISTRAL or OLLAMA
apiKeystringYesThe provider's API key, a top-level field and never part of config. Stored encrypted and never returned. Required for every provider except OLLAMA, which needs no key: leave it out
configobjectYesProvider-specific settings from the table below. Send {} when none apply

The API also lists VERTEX_AI as a provider type, but a Vertex AI provider cannot be configured or serve requests.

Config keys

ProviderKeyRequiredDefaultNotes
OpenAIbaseUrlNohttps://api.openai.com/v1For an OpenAI-compatible endpoint. Starts with http:// or https://
OpenAIorganizationIdNoYour OpenAI organization ID
OpenAIuseResponsesApiNofalsetrue sends every request through OpenAI's Responses API. See below
Azure OpenAIendpointYesYour resource URL, such as https://my-resource.openai.azure.com. Must start with https://
Azure OpenAIapiVersionNo2024-08-01-previewUsed only to list your deployments
AnthropicbaseUrlNohttps://api.anthropic.comStarts with http:// or https://
AnthropicversionNo2023-06-01Sent as the anthropic-version header
xAIbaseUrlNohttps://api.x.ai/v1Starts with http:// or https://
MistralbaseUrlNohttps://api.mistral.aiStarts with http:// or https://
OllamabaseUrlYesYour Ollama server, such as http://ollama.internal:11434. Starts with http:// or https://

Azure OpenAI has no deployment setting: the deployment name is the model name, as in azure-openai/my-gpt-4o-deployment.

useResponsesApi. You rarely need it. A request that carries function tools for a reasoning model is sent through the Responses API automatically, because chat completions refuses that combination, and so is a model that only exists on the Responses API. Setting the flag forces the Responses API for every request on this provider, tools or not. Streaming works on both. The flag takes a boolean or the string "true" or "false"; the console calls it Use the Responses API.

Managing providers

Method and pathPurpose
GET /v1/providersList your providers. apiKey is never included
GET /v1/providers/{id}Read one provider
PATCH /v1/providers/{id}Change name, apiKey or config. Every field is optional; the provider type cannot change
DELETE /v1/providers/{id}Remove a provider
POST /v1/providers/testCheck provider, apiKey and config before saving. Answers success, a message and the models found
GET /v1/providers/modelsList the models available across all your providers

The CLI equivalents are under backbone providers on the CLI page.

Was this page helpful?