Endpoints
Base URL for everything:
https://api.kral.ai/v1
Authentication is a Bearer token (Authorization: Bearer sk-kral-...) on every request. All endpoints follow the OpenAI wire format unless noted otherwise.
Core
| Endpoint | Method | Purpose |
|---|---|---|
/models |
GET | List the models your plan can use, with capabilities and parameters |
/chat/completions |
POST | Chat with any model, streaming or not |
/embeddings |
POST | Embedding vectors, billed on input tokens |
/images/generations |
POST | Image generation, billed per image |
/moderations |
POST | Content moderation, free of charge |
Audio
| Endpoint | Method | Purpose |
|---|---|---|
/audio/speech |
POST | Text-to-speech, billed per character |
/audio/transcriptions |
POST | Speech-to-text (multipart upload, up to 25 MB), billed per audio second |
/audio/translations |
POST | Speech-to-text with translation to English |
Native protocols
You are not locked to the OpenAI format. Two provider-native protocols are exposed directly, with the same key and the same billing:
- Anthropic Messages:
POST /messagesaccepts the native Anthropic request shape. Themodelfield decides routing. - Gemini:
POST https://api.kral.ai/v1beta/models/{model}:generateContent, plus:streamGenerateContentand:embedContent, accept Google's native shape.
Existing code written against the Anthropic or Google SDKs only needs the base URL and key swapped.
Assistants family (OpenAI models)
For OpenAI's stateful APIs, the gateway passes requests through with your account's gates applied: /assistants, /threads, /files, /vector_stores, /batches, and /responses. These reach OpenAI models only, since other providers have no equivalent API.
The models endpoint
GET /v1/models returns what your plan can access. Beyond the OpenAI-standard fields, each entry carries capabilities (for example image_generation) and, for media models, a media_params_schema describing the parameters that model accepts, so a client can render the right controls per model.
Account and key introspection
| Endpoint | Method | Purpose |
|---|---|---|
/key |
GET | Status of the key you are calling with: label, limits, spend (day, month, lifetime), expiry, model allowlist |
/credits |
GET | Account balance (USD equivalent plus per-currency breakdown) and lifetime usage |
Both work with any valid key, including one that has hit a spend limit, so your tooling can find out why requests are being rejected.
Key management (provisioning keys)
Platforms that issue keys automatically, one per customer or tenant, can manage keys programmatically. Create a provisioning key in the dashboard; it authenticates the endpoints below but can never call models itself.
| Endpoint | Method | Purpose |
|---|---|---|
/keys |
GET | List your keys |
/keys |
POST | Create a key. The plain-text key is in the response exactly once |
/keys/{id} |
GET | One key with limits and spend |
/keys/{id} |
PATCH | Update fields you send: name, description, monthly_limit_usd, daily_limit_usd, total_limit_usd, expires_at, allowed_models, rate_limit_rpm, disabled |
/keys/{id} |
DELETE | Delete permanently |
POST /keys accepts the same fields as PATCH (except disabled). Fields you omit in a PATCH keep their value; sending null clears one. Provisioning keys themselves can only be managed in the dashboard.
Request size
Chat and embedding requests accept large payloads (up to 50 MB) so document-heavy RAG contexts fit. Audio uploads cap at 25 MB.
Calling your agents
The endpoints above give you raw model access. To call an agent you configured in the app, complete with its instructions, knowledge, and tools, use the separate Agents API.