API

One key, three families of models: Nestor and Lambert for generation, Versatile for zero-shot classification across text, image, audio and video.

Free

The API is free while in beta. Pricing may change — we will announce it before it does, and keys created now keep working. Usage is currently capped by the same quota as the web chat.

Your keys

Quickstart

curl https://korollr.com/api/v1/chat/completions \
  -H "Authorization: Bearer $KOROLLR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "lemma-0.2",
    "messages": [{"role": "user", "content": "Explain entropy in one sentence."}]
  }'

The shape follows the OpenAI convention so existing SDKs work unchanged. It is not full compatibility: no tools, no n, no logprobs. Add "stream": true for SSE.

Models

GET /api/v1/models returns the live list. The type field tells you which endpoint to call.

ModelTypeNotes
nestor-1.1chatSmall and fast · effort low → xhigh · web search
lemma-0.2chatEveryday conversation
theorem-0.2chatMost capable, slower
versatile-lowclassify5.5 ms · BTZSC avg 0.588
versatile-highclassify10.6 ms · 0.608 · default
versatile-maxclassify~24 ms · 0.634
versatile-mmclassifyImage / audio / video

Latencies are single-item p50 on a 12-core desktop CPU, no GPU. BTZSC averages are measured on the complete test sets (12 480 examples), not a sample. Full methodology and confidence intervals are in the engine’s SPECS.md.

Limits & errors

  • 401 — missing, invalid or revoked key.
  • 429 — quota exhausted; it resets on a rolling 5-hour window.
  • 503 — Versatile weights are not installed on this deployment. Chat models are unaffected.
  • Up to 5 active keys per account, and 256 inputs per classification request.
  • A classification input is truncated at 512 tokens — a hard limit of the encoder, not a setting.