All research Announcement

Introducing Versatile 1.1

September 28, 2026

We're introducing Versatile 1.1, a zero-shot classification engine from Korollr. You give it a piece of content and a set of labels, and it tells you which one fits - no training, no fine-tuning, no examples. It works across text, image, audio, and video, and it answers in three typed shapes instead of free text: Choice (pick one of N options), Score (a graded position on a scale), and Noul (a yes/no probability).

Versatile is available today through Korollr chat and the Korollr API, under the same key as our Lambert models.

Here's what you can expect:

  • Three speed tiers. low, high, and max - pick the tradeoff that fits your latency budget. Moving up a tier costs more time and buys more accuracy; we measure both so you don’t have to guess.

  • Multimodal. Text, image, audio, and video through the same decision layer. On external image and audio benchmarks, Versatile lands within a point of published reference scores for models built specifically for that modality.

  • Long context. We’ve measured Versatile scanning documents past ten million tokens for a small set of matching passages. The architecture has no fixed context ceiling.

  • Small and self-hostable. 129-540 MB depending on the tier. It runs in-process - no external API call, no data leaving your own stack, nothing to audit but your own deployment.

  • Typed answers you can act on directly. A Choice question returns the selected option plus a probability for every option you offered. A Score returns a position on your own scale - it can land between two levels, not just on one. A Noul returns a single probability. Nothing to parse out of generated text.

Benchmarks (BTZSC full test sets, 12,480 examples; CIFAR-10 and ESC-50 with external ground truth):

  • low - 5.5 ms, 0.588 average accuracy

  • high - 10.6 ms, 0.608 average accuracy

  • max - 23.5–34.2 ms, 0.634 average accuracy

  • CIFAR-10 (image) - 0.890, published reference 0.90

  • ESC-50 (audio) - 0.818, published reference 0.82

Per-benchmark, Versatile High vs. published Jev numbers:

  • agnews - Versatile High 0.769, Jev 0.910

  • banking77 - Versatile High 0.612, Jev 0.870

  • emotion - Versatile High 0.522, Jev 0.480

Jev leads the first two by a wide margin. On emotion, Versatile comes out ahead.

What it's for:

  • Content moderation - text, image and audio through the same pipeline

  • Ticket triage - urgency, category and sentiment in one call, without customer data leaving your own stack

  • Fraud pre-screening - score every transaction before escalating the borderline ones

  • Agent routing - decide which sub-agent or tool to call without paying full-LLM cost for a simple branch

  • Sentiment analysis

A structural test, clearly labeled as such: we ran Versatile against the problem shape behind content moderation - finding a rare category inside a long stream of images. Using real photographs with external ground truth, it recovered 14 of 15 planted targets (93%) with a 0.2% false-positive rate across the surrounding images. This isn’t a moderation benchmark - telling a cat from a ship is a much easier problem than telling a violent scene from an action movie - but it shows the retrieval mechanics work: rare-class recall stays high even as the target gets sparser in the stream.

Versatile is free while in beta. Pricing may change - we’ll announce it ahead of time, and existing keys will keep working.