API
One key, three families of models: Nestor and Lambert for generation, Versatile for zero-shot classification across text, image, audio and video.
The API is free while in beta. Pricing may change — we will announce it before it does, and keys created now keep working. Usage is currently capped by the same quota as the web chat.
Your keys
Quickstart
curl https://korollr.com/api/v1/chat/completions \
-H "Authorization: Bearer $KOROLLR_API_KEY" \
-H "Content-Type: application/json" \
-d '{
"model": "lemma-0.2",
"messages": [{"role": "user", "content": "Explain entropy in one sentence."}]
}' The shape follows the OpenAI convention so existing SDKs work unchanged. It is not full compatibility: no tools, no n, no logprobs. Add "stream": true for SSE.
Models
GET /api/v1/models returns the live list. The type field tells you which endpoint to call.
| Model | Type | Notes |
|---|---|---|
| nestor-1.1 | chat | Small and fast · effort low → xhigh · web search |
| lemma-0.2 | chat | Everyday conversation |
| theorem-0.2 | chat | Most capable, slower |
| versatile-low | classify | 5.5 ms · BTZSC avg 0.588 |
| versatile-high | classify | 10.6 ms · 0.608 · default |
| versatile-max | classify | ~24 ms · 0.634 |
| versatile-mm | classify | Image / audio / video |
Latencies are single-item p50 on a 12-core desktop CPU, no GPU. BTZSC averages are measured on the complete test sets (12 480 examples), not a sample. Full methodology and confidence intervals are in the engine’s SPECS.md.
Limits & errors
- 401 — missing, invalid or revoked key.
- 429 — quota exhausted; it resets on a rolling 5-hour window.
- 503 — Versatile weights are not installed on this deployment. Chat models are unaffected.
- Up to 5 active keys per account, and 256 inputs per classification request.
- A classification input is truncated at 512 tokens — a hard limit of the encoder, not a setting.