CoralBricks Inference API
OpenAI-compatible inference for chat, interactive agents, and long-running autonomous agents. One API key, three surfaces, open models.
Coral Bricks runs open models against the OpenAI wire format. Point your
existing OpenAI SDK at the Coral gateway and you get three progressively
more powerful surfaces:
- Chat Completions — turn-based chat. Client owns the conversation.
- Responses — stateful, interactive agents. Server carries reasoning
and thread state across turns.
- Responses in background mode — autonomous agents that run for
minutes to tens of minutes, headless. Fire-and-poll, resumable
streams, no HTTP hold.
Same auth. Same model ids. Same SDK. Choose the surface that matches the
shape of your workload.
Access is self-serve: sign up, create a key at
/api-keys, and pay per token from a prepaid credit balance —
no waitlist or allowlist. Models and prices are at /models.
Building with a coding agent? This page is served as raw markdown at
https://www.coralbricks.ai/docs.md.
A machine-readable site index lives at /llms.txt, and the
complete docs in one file at /llms-full.txt.
Using OpenCode, Cursor, Cline, or Claude Code? One config stanza (or
one click) routes their agent loops through Coral — see the setup guides
for OpenCode, Cursor,
Cline, Claude Code, and
Kilo Code.
Coding agent setup guides
Point your coding agent at Coral with one config change:
- OpenCode — open-source, multi-provider terminal agent
- Cline — VS Code extension for autonomous coding
- Claude Code — Anthropic's terminal coding agent
- Kilo Code — open-source VS Code agent
- Pi — minimal, extensible terminal coding harness
Which API should I use?
| Your workload |
Use |
Why |
| Turn-based chat, single-shot generation, you already have OpenAI Chat code |
Chat Completions |
Stateless per request. Drop-in for any Chat Completions client. |
| Multi-step interactive agent, tool use, reasoning that should carry across turns, a human is waiting |
Responses |
Server-side state via previous_response_id. Reasoning preserved between turns — measurably better on reasoning models. |
| Long-running headless agent, runs in CI / cron / a cloud box / a coding agent, no human waiting on the socket |
Responses (background) |
Fire-and-poll. Resumable streams. Survives caller death, network drops, and multi-minute runs. |
All three share the same base URL (https://inference.coralbricks.ai/v1),
the same API key, and the same model ids.
Authentication
Every request needs:
Authorization: Bearer <CORAL_API_KEY>
- Mint and rotate keys at /api-keys. Keys look like
cb_… (older ak_… keys remain valid).
- Keep keys server-side. Never embed in a browser bundle or a mobile
app — anyone with the key can spend against your account.
- Newly-minted keys may take up to ~30 seconds to be honoured.
Any account can create keys and call the API; spend is drawn from the
account's prepaid credit balance (see Credits).
Credits
Check the calling key's remaining prepaid balance:
curl -sS https://inference.coralbricks.ai/v1/quota \
-H "Authorization: Bearer $CORAL_API_KEY"
{
"object": "credits",
"currency": "USD",
"balance": 34.14,
"balance_cents": 3420,
"pending_charge_cents": 6
}
balance is the effective remaining amount in USD — the prepaid balance
minus any un-flushed pending spend (balance_cents - pending_charge_cents).
The endpoint has no credit gate, so it works even at a zero balance. Spend is
also rejected up front: a request that would run against a non-positive
balance returns 429 insufficient_credits.
Zero balance:
{
"object": "credits",
"currency": "USD",
"balance": 0,
"balance_cents": 0,
"pending_charge_cents": 0
}
Pending spend:
{
"object": "credits",
"currency": "USD",
"balance": 0.4,
"balance_cents": 100,
"pending_charge_cents": 60
}
Models
The production model catalog is loaded live from the same DynamoDB control
plane that enables models in the gateway. Its machine-readable form is
available at
/api/public/models.
| Model | Context | Precision | Input / 1M | Cache write / 1M | Cached read / 1M | Output / 1M | Status |
GLM 5.3 glm-5.3-fast |
1M |
NVFP4 |
$1.12 |
$1.68 |
$0.00 |
$4.40 |
Available |
GLM 5.3 Flash glm-5.3-flash-fast |
1M |
NVFP4 |
$0.15 |
$0.23 |
$0.00 |
$0.50 |
Available |
DeepSeek V4.1 Flash deepseek-v4.1-flash-fast |
1M |
Native MXFP4 |
$0.30 |
$0.09 |
$0.00 |
$1.20 |
Available |
Pass the model id verbatim as the model field on any request. Older
-fp4 model IDs still work: they reach the same model at the same price.
The exact set enabled for your key is also available programmatically:
curl -sS https://inference.coralbricks.ai/v1/models \
-H "Authorization: Bearer $CORAL_API_KEY"
The same model ids work across Chat Completions, Responses, and background
mode. Capability fields such as supports_image_input on each /v1/models
row are authoritative; unsupported content returns 400 unsupported_content_type rather than being silently dropped.
Measured serving stats
Each row also carries what our own serving measured over the last 30 minutes
and the last day — the same two windows OpenRouter's endpoints API uses.
| Field |
Meaning |
decode_speed_last_30m / _1d |
Output tokens per second while a response streams — the rate one caller's answer arrives at, not fleet capacity. |
latency_last_30m / _1d |
Time to first token, median, in seconds. |
cache_hit_rate_last_30m / _1d |
Share of prompt tokens served from cache, percent. |
These are measured on requests we served on our own GPUs, so they describe
our serving rather than a vendor's. A model that was too quiet to measure in a
window simply carries no field for it — never a zero, which would claim a
measurement we do not have.
Cached input tokens are free — a prefix that is already resident is billed
at $0, on every model and both APIs, with no cache-management flags to set.
Agent loops that re-send a growing conversation each turn benefit most: the
repeated prefix costs nothing after the first turn.
Cached reads and physical writes are reported per response. Cache writes are the novel
input for that request: prompt_tokens - cached_tokens. The separate billable counter is
non-zero whenever retention is in effect. Those tokens bill at the cache-write rate instead
of the input rate, on every model and every route; opting out bills them as plain input.
A new token is billed once either way, never at both rates.
The charge sits on the write because memory capacity is what limits an inference stack: a
cache write is the point where tokens take up capacity, and a read reuses capacity the write
already paid for. In an agent loop this keeps the cost of context in step with the context.
Each turn re-reads everything before it, so read tokens grow with the square of the number
of turns (50 turns that each add 2,000 tokens write 100,000 tokens and re-read 2.45 million),
while each token is written once.
Retention controls
prompt_cache_retention has only two states:
- Empty — omit the field or send
null to use the platform's default retention.
- Off — send
"off" to opt out of retained-cache billing.
OpenAI's values "in_memory" and "24h" are also accepted, so a client that sends them
works unchanged. They behave like an empty field and do not set a duration. Any other
value, including an empty string, returns 400 invalid_prompt_cache. To request a
specific retention duration, leave prompt_cache_retention empty and use a positive,
whole-minute TTL in prompt_cache_options.ttl:
{
"prompt_cache_options": { "ttl": "30m" }
}
To opt out:
{
"prompt_cache_retention": "off"
}
Opting out does not turn off engine prefix caching. Physical writes and cache hits can still
occur and are still reported; it only keeps novel input in the plain-input billing bucket.
"usage": {
"prompt_tokens": 20767,
"prompt_tokens_details": {
"cached_tokens": 20032,
"cache_write_tokens": 735,
"billable_cache_write_tokens": 0
},
"completion_tokens": 246,
"cost": 0.001098,
"cost_details": {
"currency": "USD",
"input": 0.00099,
"cached_input": 0,
"cache_write": 0,
"output": 0.000108
}
}
Every response's usage also carries the request's cost in USD (usage.cost)
and a per-bucket usage.cost_details breakdown — a Coral extension that equals
what the request is billed. Present on both APIs, unary and streamed (on the
final usage-bearing chunk). cached_input is $0 (cached reads are free) and
cache_write is non-zero only when a paid retention request was delivered.
Prompt cache key
Caching needs no key. A conversation that re-sends its history is served from
the cache on its next turn whether or not the request names one.
prompt_cache_key is an optional string on /v1/chat/completions and
/v1/responses that says which cache a request belongs to:
{
"model": "glm-5.3-fast",
"prompt_cache_key": "session-7f3a9c",
"messages": [{ "role": "user", "content": "..." }]
}
- Same key, same prefix: reused. Requests that carry one key find each
other's cached prefix. Without a key, two requests that share a long prefix
but aren't turns of one conversation don't always find it.
- Different keys never share. The same prompt under another key, or with
no key, is written to the cache again.
- Use one stable ID per session or conversation, and keep it for the whole
conversation. A new key starts cold: its first request writes the full
prompt, system prompt and tools included.
- Share one key across conversations that share a prefix, such as forks of
one parent response or workers reading the same document. See
Cross-session prompt cache reuse.
promptCacheKey is accepted as an alias. Send either one. If a request
carries both, they must match.
- A key is at most 512 characters. A longer key, a key that isn't a string, or
two aliases that differ return
400 with the code invalid_prompt_cache.
/v1/messages has no such field. Caching there is automatic, and a
prompt_cache_key in the body is ignored.
Whether a coding agent sends a key is up to the agent:
| Agent |
Sends a key |
| Codex CLI |
Always |
| OpenCode |
With setCacheKey: true; always in OpenCode 2 |
| Kilo Code |
With setCacheKey: true |
| Pi |
On the Responses API; with supportsPromptCacheKey: true in oh-my-pi |
| Cline |
No |
| GitHub Copilot |
Not with the Chat Completions setup |
| Cursor |
Not established |
| Claude Code |
Not applicable: it uses /v1/messages |
1. Chat Completions
Turn-based chat. Stateless per request — you send the full message
history on every call, and the server returns one assistant turn. The
foundational OpenAI-compatible chat endpoint, unchanged from what your
existing SDK expects.
Use when: you already have working chat.completions code, you're
generating a single response, or you don't need the server to remember
anything about the conversation.
POST /v1/chat/completions
curl -sS -X POST https://inference.coralbricks.ai/v1/chat/completions \
-H "Authorization: Bearer $CORAL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5.3-fast",
"messages": [
{"role": "system", "content": "You are a careful research assistant."},
{"role": "user", "content": "What are the key metrics to watch for a Series A SaaS company?"}
]
}'
Streaming is supported the OpenAI way — pass stream: true and consume
the SSE. Tool calling (tools / tool_choice) works the same as
upstream. Background mode is not available on this endpoint — use
/v1/responses for that.
from openai import OpenAI
client = OpenAI(base_url="https://inference.coralbricks.ai/v1", api_key="cb_...")
resp = client.chat.completions.create(
model="glm-5.3-fast",
messages=[
{"role": "system", "content": "You are a careful research assistant."},
{"role": "user", "content": "Explain KV attention in one paragraph."},
],
)
print(resp.choices[0].message.content)
2. Responses — interactive agents
Stateful. The server holds reasoning and conversation state between
turns, so subsequent calls can reference an earlier response by id
instead of resending the full history. This preserves the model's
reasoning context across turns, which measurably improves multi-step
tool-using agents.
Use when: you're building an interactive agent — multi-step tool
calling, iterative reasoning, threaded conversation — and a human (or
an upstream agent) is waiting on the reply. Timescale: seconds.
POST /v1/responses
curl -sS -X POST https://inference.coralbricks.ai/v1/responses \
-H "Authorization: Bearer $CORAL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5.3-fast",
"input": "Sketch a plan to refactor the auth module."
}'
Follow-up turns pass previous_response_id to continue the thread. The
server carries state; you don't have to replay the transcript:
first = client.responses.create(
model="glm-5.3-fast",
input="Sketch a plan to refactor the auth module.",
)
# Same thread, no history replay needed.
next_turn = client.responses.create(
model="glm-5.3-fast",
input="OK, now do step 1 in detail.",
previous_response_id=first.id,
)
print(next_turn.output_text)
| Field |
Type |
Notes |
model |
string |
Required. One of the ids from GET /v1/models. |
input |
string or array |
Required. Same shape as OpenAI Responses. |
instructions |
string |
Optional system / developer instructions. |
previous_response_id |
string |
Continues an existing thread. Sticky-routed back to the same replica for locality. |
stream |
boolean |
When true, SSE stream of Responses events. |
| ...other OpenAI fields |
|
Forwarded verbatim. |
GET /v1/responses/{id}
Retrieve any prior response by id — useful for auditing, resuming a
UI after a page reload, or re-hydrating a thread from persistent
storage. Also serves stream resumption (see background mode).
curl -sS https://inference.coralbricks.ai/v1/responses/$RESPONSE_ID \
-H "Authorization: Bearer $CORAL_API_KEY"
DELETE /v1/responses/{id}
Delete a stored response.
Long generations and client timeouts
Streaming keeps bytes moving, but your SDK or HTTP client still owns a
total-request timeout. Many OpenAI-compatible clients default to roughly
300 seconds; a large max_tokens value, a long agent turn, or a large
prompt can legitimately take longer than that even while tokens stream
normally.
For interactive calls:
- Keep
max_tokens aligned with what the user will actually read.
- Set your client's timeout based on the expected whole turn, not just
time-to-first-token.
- If a call is intentionally long, use
/v1/responses with
background: true and poll or resume the stream rather than holding the
original request open.
If the client disconnects while Coral is still generating, usage records
show HTTP 499 / client_disconnected. The abandon_verdict
distinguishes a normal caller hang-up from a Coral deadline miss.
3. Responses in background mode — autonomous agents
Same endpoint as above, plus one flag: background: true. The call
returns immediately with a response id and status: "queued". The
model then runs on Coral's cluster — not on your socket — for as long
as the work takes. You poll (or resume the stream) on your own cadence.
Use when: the caller is not a human at a chat window. Long-running
research loops, coding agents, batch pipelines, CI jobs, cron tasks,
overnight runs. Timescale: minutes to tens of minutes. Nobody is
waiting on the next token.
Why the shape matters for autonomous agents:
- No HTTP hold. The socket closes when
create returns. No proxy
timeouts, no keepalive plumbing, no reconnect logic.
- Survives caller death. If your agent process crashes or your CI
runner is preempted, the response keeps running server-side. A fresh
process retrieves by id and gets the completed result.
- Resumable streams. If you were streaming and the socket dropped,
a
GET /v1/responses/{id}?stream=true from any process picks up
where you left off.
- Location-independent. Kick off from your laptop, poll from a
Lambda, render in a Slack bot. One id, three consumers.
POST /v1/responses (with background: true)
curl -sS -X POST https://inference.coralbricks.ai/v1/responses \
-H "Authorization: Bearer $CORAL_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "glm-5.3-fast",
"input": "Plan a 5-step refactor for the attached repo, then execute step 1.",
"background": true
}'
{
"id": "resp_01HXY…",
"object": "response",
"status": "queued",
"model": "glm-5.3-fast"
}
GET /v1/responses/{id} (polling)
The status moves through queued → in_progress → completed | failed | cancelled. Poll on any cadence you like.
import time
queued = client.responses.create(
model="glm-5.3-fast",
input="Run the full research loop. Take your time.",
background=True,
)
while True:
current = client.responses.retrieve(queued.id)
if current.status in ("completed", "failed", "cancelled"):
break
time.sleep(5)
print(current.output_text)
POST /v1/responses/{id}/cancel
Cooperative cancel of an in-progress background response. Queued
responses flip to cancelled synchronously.
curl -sS -X POST https://inference.coralbricks.ai/v1/responses/$RESPONSE_ID/cancel \
-H "Authorization: Bearer $CORAL_API_KEY"
Stream + background
You can combine stream: true with background: true. The initial
call returns an SSE stream that you can consume; if the socket drops or
you want a second consumer, GET /v1/responses/{id}?stream=true picks
up from where the last event left off. Useful when the agent loop and
the renderer live on different machines.
Streaming
Every surface supports OpenAI-shape SSE streaming via stream: true.
Consume it exactly as you would from OpenAI's SDK. See the background
section above for stream resumption via GET.
The balance may take up to ~30 seconds to reflect a top-up or completed
inference charge because API-key account data is briefly cached.
Errors
Errors follow the OpenAI error shape:
{"error": {"message": "…", "type": "…", "code": "…"}}
| Status |
code |
Meaning |
400 |
bad_request / model_required |
Malformed body or missing required field. |
400 |
invalid_prompt_cache |
A prompt cache field is invalid, e.g. a prompt_cache_key over 512 characters. The message names the field. |
401 |
missing_api_key / invalid_api_key |
Auth header absent or rejected. Re-mint at /api-keys. |
403 |
access_denied |
The key's account can't use this endpoint. Contact us if you think that's wrong. |
404 |
model_not_accepted |
The model id isn't enabled for your account or doesn't match a known id. |
404 |
response_not_found |
Wrong response_id, or it belongs to another account. |
429 |
rate_limit_exceeded |
Per-API-key rate limit. Slow down and retry with backoff. |
499 |
client_disconnected |
The caller closed the connection. For long generations, use background mode or raise the client's whole-request timeout. |
502 |
upstream_error |
Upstream model server returned an error. Safe to retry. |
503 |
backend_unconfigured |
Transient. Retry. |
504 |
timeout |
Sync request waited too long. Re-issue with background: true. |
Limits
Rate, concurrency, and context-length limits are set per API key and
account. The defaults fit a single interactive agent or a single
background loop comfortably. If you're fanning out (multiple
background agents in parallel, a batch pipeline, a coding agent that
spawns sub-agents), talk to us so we can raise the cap on your key.
Support