All research Announcement

Introducing Nestor 1.1

September 23, 2026

We're introducing Nestor 1.1, a new small language model from Korollr. At 12.7 MB, it is small enough to run on an ordinary CPU, and it holds its own against models up to eight times its size.

Nestor 1.1 is available today in Korollr chat and through the Korollr API.

Here's what you can expect:

  • Small by design. 19.9M parameters and 12.7 MB. SmolLM2-135M is twenty times larger on disk, and Pythia-160M twenty-four times.

  • Strong where it counts. 80.7% on SciQ, level with SmolLM2-135M, and the best score in its comparison group on reading comprehension (BoolQ, 62.9%) and truthfulness (TruthfulQA, 27.1%).

  • Grounded answers. When the answer isn't in the text it's given, Nestor 1.1 makes something up less often than the Pythia models, and no more often than SmolLM2-135M.

  • Tools. Nestor 1.1 calls tools, and when it does, it invents arguments far less often than the Pythia models.

  • Web search. One switch, and Nestor can look things up on the web.

  • Checked answers. Every answer is checked against the source it came from, and Korollr tells you when it could not be.

  • Four effort levels. One model, one menu: Low, Medium, High and Extra high.

Performance

We compared Nestor 1.1 with three open models between 70 and 162 million parameters. Every model was scored on the same harness, zero-shot.

Nestor 1.1 benchmark results against Pythia-70M, Pythia-160M and SmolLM2-135M

SmolLM2-135M stays ahead on average (52.6% against 47.7%). It is also seven times larger in parameters and twenty times larger on disk. Nestor 1.1 comes second, ahead of Pythia-160M, a model eight times its size.

Multiple-choice benchmarks only say so much about how a model answers, so we also ran head-to-head duels: every model answered the same 80 open instructions, and a much larger model judged each pair of answers, in both orders so that the position of an answer could not decide the outcome. Nestor 1.1 scores 1012 Elo, ahead of Pythia-160M (958) and Pythia-70M (878). SmolLM2-135M leads at 1151.

Honest by default

A small model knows less than a large one. What matters is what it does when it doesn't know. On SQuAD 2.0, where half of the questions we asked have no answer in the passage, Nestor 1.1 has the lowest hallucination rate of the four models (53.7%), level with SmolLM2-135M and well below both Pythia models.

The same holds for tools. When Nestor 1.1 calls a tool, it fills in values it was not given in 25.5% of calls, against 67.1% for Pythia-160M and 99.4% for Pythia-70M. It solves 23% of BFCL v3 simple tasks, where the Pythia models solve 3.5% or less.

Effort levels

Nestor 1.1 is a single model with four effort levels, chosen from the model menu:

  • Low

  • Medium

  • High

  • Extra high

Medium is the default.

Tools and web search

Nestor 1.1 decides for itself when a request needs a tool: a calculation, a word count, the time, or a web search when you turn it on. The tool runs, and Nestor answers from its result. You see every call it makes and the source behind the answer.

Every answer is checked against the sentence it came from before it reaches you. When the check fails, Nestor says so, with the reason, instead of passing a guess off as a fact. A small model will not know everything; it should at least tell you when it does not.

Limits

Nestor 1.1 is a small model. It knows less than large models, it can be wrong with confidence, and it is at its best on short, factual exchanges. Long conversations drop their oldest turns. Check anything important.

Availability

Nestor 1.1 is available now in Korollr chat: pick it from the model menu, choose an effort, and turn web search on or off. Developers can use it through the Korollr API with the model id nestor-1.1:

curl https://korollr.com/api/v1/chat/completions \
  -H "Authorization: Bearer $KOROLLR_API_KEY" \
  -H "Content-Type: application/json" \
  -d '{
    "model": "nestor-1.1",
    "effort": "medium",
    "web_search": true,
    "messages": [{"role": "user", "content": "Who wrote The Little Prince?"}]
  }'

effort accepts low, medium, high and xhigh. Each response carries a nestor field: whether the answer was checked against its source, why, the source itself, and every step Nestor took. The API is free during the beta.