← Blog

Kolibri-1 on the NVIDIA DGX Spark: a first comparison with Qwen3.6

Iristrace
Kolibri-1 on the NVIDIA DGX Spark: a first comparison with Qwen3.6

By Iristrace B.V., with André Kingham, CEO

On 3 October, the Day of German Unity, Aleph Alpha in Heidelberg released Kolibri-1, an open-weight language model built for German and English. We were a day early: its page on Hugging Face was already up on 2 October, before the model itself, and we came across it while going through the previous day’s new repositories. Kolibri is German for hummingbird, the smallest bird there is. Its model card, though, sets the hardware bar high. It says, verbatim: “Model memory footprint: ~78 GB (FP8 weights). Minimum: 2× A100 80 GB, 2× H100 SXM5, 1× H200, 1× B200 or 1× B300. Recommended: 2× H100 SXM5, 2× H200, 1× B200 or 1× B300.” Those are data-centre GPUs, the kind that go into large servers.

By the card’s own figures, and not only by its name, it ought to manage with less: about 78 GB fits inside a DGX Spark’s 128 GB.

We wanted to know two things. Does Kolibri-1 run on a DGX Spark, yes or no? And how good is it at the job our customers ask of an assistant, finding and using the right information in their own documents, compared with Qwen3.6?

Shortly after the official release, then, we set out to answer those two questions ourselves: on one of our NVIDIA DGX Sparks, desk-sized workstations built around NVIDIA’s Grace Blackwell chip, mounted in our server racks, and on the kind of work our customers do, in a direct comparison with Qwen3.6-35B-A3B, both in FP8, on identical machines, through the same pipeline.

We threw AI at the question to gauge the potential: an AI-written test library, AI markers, 1,428 answers. What that shows is promise, and promise is what practice, with our customers’ documents and questions, has to confirm. Two days of hard work – a first encounter with the model.

The short version

  • DGX Spark, yes or no? Yes, and on the first day. Kolibri-1 runs on a single DGX Spark, with Aleph Alpha’s own plugin for the inference server, no changes to the code and no tuning. Its model card’s minimum, verbatim, is “2× A100 80 GB, 2× H100 SXM5, 1× H200, 1× B200 or 1× B300”.
  • Answering from documents: level, on equal terms. With search and without reasoning, the two score the same, 0.55 against 0.55 per answer under the first marker and 0.50 against 0.47 under an independent second marker, at every sampling setting. Kolibri-1 takes more rounds and reads more to get there: typically four model rounds and five tool calls per answer against Qwen3.6’s two and two.
  • Level on average, not on every question. On two of the seventeen questions the models differ far beyond chance, in opposite directions. Kolibri-1 found a customer’s stricter 90-day rule in 11 of 14 answers, Qwen3.6 in 1, because Kolibri-1 kept searching. Qwen3.6 noticed in all 14 that the contract names a grease grade no longer approved; Kolibri-1 noticed it too and still concluded “approved” in 11. Same average, different strengths.
  • With reasoning switched on, Kolibri-1 was slightly ahead in both runs, at a price. 0.65 against 0.58 with search under the first marker, 0.60 against 0.51 under the second: a lead not yet distinguishable from chance. A Kolibri-1 answer then took a median of five minutes against Qwen3.6’s one and a half, under our test load.
  • The clearest difference is behaviour, not accuracy. When a search returns nothing, Kolibri-1’s reasoning mode keeps retrying it with new wordings until a round budget stops it; in a test harness with no fallback answer, that meant no answer at all. It is something to watch, measure and pre-empt with a response strategy, not a reason to rule a model out, and it is the kind of thing a benchmark score does not show.
  • Searching is cheaper, not more accurate. On a library that fits in the prompt, putting everything in answered as well or better for three of four settings. Searching costs up to three times fewer tokens and keeps working when the library outgrows the window.
  • A gap in the library does not produce “I don’t know”. When our search missed the one document that held an answer, all 34 answers to that question were confident and incomplete; not one said the library did not cover it. Coverage is part of the answer, whichever model gives it.
  • Where a model comes from is now a real choice. With Kolibri-1, a capable German model joins the American, Chinese and French ones. For many companies, that is a supply-chain decision as much as a performance one.

Why we ran this test

Our customers’ documents hold their contracts, procedures, quality records and HR rules. When they ask us about AI, the first question is rarely how clever it is. It is where their documents go.

Open-weight models change the answer. When a maker publishes the model itself, you can run it on hardware you control, and no document has to leave for an outside AI service. Under a licence like Apache 2.0, you can also keep running it, whatever its maker decides next.

Most of the capable open models in use today come from the United States and China. Qwen3.6 is one of them. Europe has had fewer, and fewer still built with German in mind. A capable model from Heidelberg matters to organisations that would rather keep their AI supplier, like their data, in Europe.

A European address does not answer anyone’s questions, though. The only way to know whether a model is good enough for our customers’ work is to give it that work. So we did.

Why these two

Both are open models under the Apache 2.0 licence, both can run on hardware you own, and both are built the same way: a large model of which only a small part (about 3 billion parameters) works on each word. That is why both are quick on modest hardware.

We ran both in FP8, which stores each of the model’s numbers in eight bits instead of sixteen and so halves the memory it needs. Both makers publish their models in this format, so each model ran in its maker’s own version, not a third party’s.

Kolibri-1 was built for this job. Its model card names, among the model’s intended uses, “retrieval-augmented generation”, “long-document processing” and “agentic tool calling”, and says it suits “question-answering systems over an organisation’s own material”. That is what we tested.

Qwen3.6 is not a soft target. Alibaba’s Qwen family has become the workhorse of open-weight AI: it is what many teams reach for first when they run a model themselves, and it has been the base for many derivative models around the world. NVIDIA, for one, has fine-tuned several of its Qwen3-Nemotron models from Qwen3 and publishes its own compressed build of the very Qwen3.6 we tested. Even its Nemotron 3 Nano, built on NVIDIA’s own architecture, says on its model card that it “was improved using” Qwen models, among others. This release also reads images and covers a broad range of languages.

Our own experience with the model has been largely good. Seen that way, a newcomer that merely matches it on a company’s documents has cleared a high bar, and that is what we wanted to look at more closely.

Kolibri-1Qwen3.6 (35B-A3B)
Made byAleph Alpha Research GmbH, HeidelbergQwen Team, Alibaba Group’s Tongyi Lab
ReleasedOctober 2026April 2026
Size, total / active per word78 billion / 3.5 billion35 billion / 3 billion
LanguagesGerman and English, officiallyA broad range
Reads imagesNoYes (not compared here)
Longest input262,144 tokens262,144 tokens
Minimum hardware, per its model card”2× A100 80 GB, 2× H100 SXM5, 1× H200, 1× B200 or 1× B300”not stated
Model card on Hugging FaceAleph-Alpha/Kolibri-1Qwen/Qwen3.6-35B-A3B-FP8

Where the test ran: the Iristrace Docs knowledge base

Iristrace Docs turns documents into usable data: an invoice or a delivery note goes in, structured fields come out, matched to the right accounts and codes in systems such as SAP Business One. It connects to the places documents already live, SharePoint, Exchange and Gmail mailboxes among them. The step we are building now is a knowledge base you can ask questions, by chat and also by voice, with a spoken answer, and that answers from a company’s own documents. The test ran on that knowledge base as it stands today. This is not a product review; it is a report on what we learned running two models through it.

Highlighted: used in this test

Iristrace Docs

Web app

AI chat · document upload · access management

Iristrace Docs

Mobile app

Chat · ticket capture and quality processes · voice questions

Iristrace DocsOne platform, in our cloud or in your own infrastructure

Workspaces: Engineering · HR · Quality · Purchasing · Finance · …

Permissions per user and workspace, on every query

Intelligence

  • RAGAnswers from the documents' text, with citations
  • Tool callingSearch, open documents, query the data warehouse
  • AI extractionDocument → structured data (JSON) and text (Markdown); scans and photos need a vision model

Storage

  • Text storeEvery document whole, as Markdown
  • Record storeExtracted data, as JSON
  • Data warehouseTables built automatically from the records
  • Vector storeThe search index, one per workspace

Integrations

  • SharePointWhere the documents live. Fetched through Microsoft Graph; the synchroniser that picks up changes in a library or folder on its own is in development. Who may read what is set by the workspace
  • SAP Business OneFinds the right GL account and the SAP codes for business partner, product and more
  • Exchange and Gmail mailboxes
  • Webhooks
  • Transcriptions
  • Iristrace Checks
  • More, built with our customers

The prompt carries only the context the user is allowed to see

Private LLM

Text and vision: one model that does both, such as Qwen or Mistral, or two models side by side

  • LiteLLMModel router
  • NVIDIA DGX Spark 1Kolibri-1 in this test
  • NVIDIA DGX Spark 2Qwen3.6 in this test; earlier, the vision model that read the scans
Iristrace Docs under the hood.

It works in five steps, and the test went through all five:

  1. Connect. Documents come from where they already live. For this test, a SharePoint library, fetched once through Microsoft Graph. The synchroniser that notices a changed file and fetches only that file is in development.
  2. Read. Every file becomes text: Word, PDF, spreadsheets. Scanned pages are read by a vision model that runs on our own hardware. Each document is kept whole, as Markdown, not only in fragments. That text store is the one the knowledge base is built on.
  3. Index. Documents are split into passages along their headings and indexed for search, each workspace in its own index.
  4. Answer. The assistant searches, opens whole documents when a passage is not enough, and answers with citations. Each citation leads to the passage and to the original file.
  5. Permissions. The workspace sets the perimeter: a person only gets answers from workspaces they are allowed to read, and the database itself enforces this, not only the application on top of it. Because a workspace can be fed from a single folder of a SharePoint site, its audience can be narrowed to the group in charge. We deliberately do not copy SharePoint’s permission lists: one clear perimeter is easier to understand and to check.

The assistant is also told what a workspace contains, so that it can say when a question falls outside it instead of searching on and on.

For this test, every step ran on our own hardware: reading the scans, indexing, searching and answering. The outside services were Microsoft’s SharePoint, where the library was stored as a customer’s would be, and two AI models from other makers: Anthropic’s Claude wrote the fictional documents and the tests and marked every answer, and OpenAI’s GPT-6.1 Sol marked every answer a second time, independently and blind. Both worked only with synthetic material, and for a reason: only with documents whose content and flaws we set ourselves do we know for certain which answer is correct.

The test: a company that does not exist, with problems that do

We did not want to test the models on trivia. We wanted to know whether they help an employee who has a question about the company’s own rules. So we needed a company.

We guided Claude Opus 5.5, at extra-high effort, to write the document library of a fictional manufacturer: a bearing maker for the car industry, part of a larger group, with sites in six countries. Forty-seven documents: policies, procedures, technical specifications, work instructions, approval matrices and a supply agreement, in English, Spanish, Polish, German and Chinese. About 160,000 tokens of text in all.

Then we did what time does to every real document library: we planted 18 deficiencies, each one documented in an answer key. A few examples:

  • Two revisions of the travel policy in the same library, with different hotel limits.
  • Two documents that give the same manager two different approval limits.
  • A procedure that cites a checklist that does not exist.
  • A plant’s work instruction that quietly departs from the group specification.
  • A customer’s requirement that is stricter than our own procedure.
  • A regional HR document overtaken by a change in the law.

One more document was written and kept out of the library on purpose, to see whether a model admits that the library does not cover a topic instead of making something up.

A detail we found telling: the AI that wrote the library had, on its own initiative, added further contradictions that nobody had ordered. A second review found 18 of them, and they were removed again. That kept the focus on the traps we had planned deliberately.

Why go to this trouble? Because a document assistant that answers “the hotel limit is €120” because it found the old revision first is worse than no assistant at all. The question is not whether a model writes well. It is whether it notices.

From text to a real document library

Writing the documents was half the job. The other half was making them behave like real files.

Each document was produced in the format a real company would use: Word for most, PDF for others, spreadsheets for the matrices. Three were turned into scans, slightly rotated, speckled and saved as images with no text behind them, so the only way in is to read the page. Every document carries a control block (number, revision, owner, approver, effective date, next review) in one of two house styles, because the group and its subsidiary format their documents differently.

We then uploaded the library to a SharePoint site set up for demos, one folder per department, and loaded it into Iristrace Docs from there.

Before any model saw a question, a script checked that every fact in the answer key could be found in the converted text, the scans included. When a model misses a fact, the fact was there to be found.

One job, tried three ways

You will not find a table of benchmark scores here. Those exist, and the model makers publish them. We tested the one thing an employee most often asks a document assistant to do: answer a question from the company’s own documents, and get it right when the documents disagree. Extraction and structured data analysis are part of Iristrace Docs as well, but not of this test; more on that at the end.

Answer from your documents: read the second document

This is the job most people mean when they talk about AI on company documents. Before it answers, the assistant searches the library, reads what it finds and cites it. The jargon is “retrieval-augmented generation”, or RAG.

We asked 17 questions an employee might ask, each with an answer key. Some examples:

  • May I accept a €75 gift from a supplier? The group’s code of conduct and the local purchasing policy give different thresholds. A correct answer mentions both.
  • What is the current hotel limit for business travel? An older revision with a lower limit is still in the library.
  • How much notice must we give a particular customer before we change a product? The customer’s requirements are stricter than our general procedure.
  • What happens if a supplier does not return its conflict minerals report? The document that answers this is the one we held back. The right answer is “the library does not cover this”.

To count as correct, an answer needed the key fact and had to point out the conflict planted behind it. The right fact with the conflict missed counts as partial.

Each question was asked in English and in German, under three sampling settings (the model’s fixed setting our chat uses, a low temperature, and the maker’s own recommendation), with reasoning off and on, and one to three times per setting: 1,428 answers in all. Every answer was shuffled, stripped of the model’s name and marked blind against the answer key, twice: by agents running Claude Opus 5.5, and independently by OpenAI’s GPT-6.1 Sol. The two markers agreed on 78% of grades (Cohen’s kappa 0.68). The second was stricter about how complete a correct answer must be; where it says something different, we say so.

With search, the way the product works. Shares of all answers, the first marker first and the second in brackets:

Kolibri-1Qwen3.6
Correct / partial / wrong, reasoning off42% / 26% / 32% (32 / 36 / 32)45% / 21% / 34% (28 / 38 / 34)
Mean score per answer, reasoning off0.55 (0.50)0.55 (0.47)
Seconds per answer, median, reasoning off, under test load5532
Correct / partial / wrong, reasoning on (Kolibri-1 at high)60% / 10% / 29% (49 / 22 / 29)50% / 16% / 34% (35 / 32 / 32)
Mean score per answer, reasoning on0.65 (0.60)0.58 (0.51)
Seconds per answer, median, reasoning on, under test load32182

Score per answer: correct 1, partial ½, wrong 0. Every answer was graded twice. Reasoning off: 238 answers per model. Reasoning on, with Kolibri-1 at high, its most expensive level: 68 answers per model, two runs each. The seconds are medians under our test load, which was heavier on Kolibri-1’s server than on Qwen3.6’s for most of the day: compare them within a model, not between the two. A single-request speed test on quiet servers follows. The same figures question by question are in annex E.

Without reasoning, the two are level, and the tie does not depend on the sampling setting: at temperature 0, at 0.3 and at each maker’s own recommendation the lead changes sides by a few hundredths. Question by question, Qwen3.6 is ahead on eight and Kolibri-1 on five under the first marker, and six against six under the second; the mean difference per question is 0.00 under the first marker and 0.03 in Kolibri-1’s favour under the second, with a 95% interval of about ±0.12 around either (annex D). Kolibri-1 works harder for the same result: a typical answer takes it four model rounds and five tool calls, against two and two for Qwen3.6, and each round resends the instructions, the tools and everything found so far, so it reads about 86,000 prompt tokens per answer in all, against 53,000. That, more than anything the servers do, is why it takes longer.

With reasoning on, Kolibri-1 was slightly ahead, in both of its runs and under both markers. Question by question it leads on six, trails on one or two, and ties on the rest; the mean difference per question is 0.07 to 0.08 in its favour, and the 95% interval ends at zero under the second marker and only just clears it under the first. Its second run narrowed the lead. So “slightly ahead” is what the data supports, “answers better” is not, and with the number of comparisons we ran (annex D) it is a direction the next tests have to confirm. It also costs: under our test load, a median of over five minutes per answer, and about a quarter of its answers over ten minutes. Qwen3.6 with reasoning stays under a minute and a half at the median.

Level on average, not on every question

The tie is an average, and it hides two questions on which the models differ far beyond chance: the only two of seventeen whose difference holds under both markers once we correct for the number of questions compared (annex D). They point in opposite directions.

A customer’s stricter rule: Kolibri-1 rightly reads the second document. How much notice must we give HallvikDemo Trucks before a product change? The general change procedure says 60 days; the customer requirements matrix says 90 for this customer. Kolibri-1 answered correctly in 11 of 14 answers, Qwen3.6 in 1 of 14 (12 against 2 under the second marker), in English and in German alike. Both models can read the contradiction: handed only the right documents, each got it right every time. The difference is how they search. Qwen3.6 typically searched once and answered from the procedure; four times it said the documents do not mention HallvikDemo, with the matrix among its own search results. Kolibri-1 took three or four rounds and found the matrix. Here the extra rounds we counted above as a cost are what bought the right answer, on exactly the kind of question a quality or compliance team worries about. With reasoning on, Qwen3.6 searched more and was right in two of four.

A grease no longer approved: Kolibri-1 states the wrong conclusion. Which grease grade are we buying from ChemcoDemo, and is it the approved one? The supply contract names a grade the specification has withdrawn. Qwen3.6 was right in all 14 answers. Kolibri-1 was fully right in 3. In the other 11 it noticed that the contract names the withdrawn WH-2 and still concluded, usually in its first sentence, that the purchase is approved. A reader who stops at the first line is misled. Handed only the right documents, Kolibri-1 got this one right every time as well, and with reasoning on it was right in all four answers with search too.

A third question separates them under the first marker only: the first order before a supplier’s conflict-minerals report (question 17). That difference measures our answer key, not the models. The key rewards a firm “no” that the library cannot support, and Kolibri-1 mostly said the documents do not answer it (see the limits below).

Fourteen answers per model and question make these behaviours on two questions, not a ranking. What they show is that two models with the same average are not interchangeable: which one serves you better depends on the questions your people ask, and that can only be found out on your own documents.

Why not simply give it everything?

Both models can take in 262,144 tokens at once, more than our whole library. It is tempting to skip the search, put every document into the prompt and let the model find the answer itself. Nothing gets missed by a search that way. We tried it, with both.

Mean score per answer, first marker (second marker)With searchWhole library in the prompt
Kolibri-1, reasoning off0.55 (0.50)0.51 (0.44)
Qwen3.6, reasoning off0.55 (0.47)0.67 (0.54)
Kolibri-1, reasoning on0.65 (0.60)0.73 (0.63)
Qwen3.6, reasoning on0.58 (0.51)0.67 (0.56)

We expected search to win. It did not. With the whole library in view, Qwen3.6 answered clearly better without reasoning, by 0.11 to 0.16 per question, and that holds under both markers; it is the one clear difference in reading between the two. Both models answered better with reasoning. Search won only for Kolibri-1 without reasoning. So the honest lesson is this: searching is cheaper, not more accurate. A question with the whole library attached costs about 150,000 prompt tokens; with search, 53,000 to 86,000, summed over an answer’s rounds. And the whole-library approach stops working the day the library outgrows the window, which for a real company is soon.

The first read of the whole library took about a minute on either server (62 to 68 seconds). After that, because the server keeps what it has already read, answers started within two to four seconds. Whole-library prompts were also where the few runaway answers without reasoning happened: six of 68 at the fixed setting started repeating themselves or ran into the output cap.

Our own homework

Three documents were the trouble, whichever model was answering. A problem-solving (8D) procedure was never found by the search, in 35 attempts. A scanned procedure was found 8 times in 35. A customer requirements matrix, for the question that needs it, 9 times in 35. That is not the model’s fault. It is ours, and it is the next thing we are fixing: the 8D procedure names the person one question asks about only once, and a search by meaning ranked other passages first every time. A search that also matches exact words would most likely find it.

What the models did when the search missed the document is the more useful lesson. Of the 34 answers to that question, all 34 were partial: each model answered from the documents it did find, with a title that was plausible and incomplete. Not one said “I don’t know.” That is what a gap in a library looks like from the user’s side: not a doubtful answer, a confident one.

The held-back document is the counter-example. Asked about a topic the library does not cover, both models mostly said so: Kolibri-1 in 26 of 36 answers, Qwen3.6 in 30 of 42. When nothing is found at all, both admit it. When something nearly right is found, neither does.

The clearest difference is behaviour

Alongside the questions, we ran eight small agent scenarios with mock tools, five trials each, in English and German, with reasoning off and on. They replay failures we have seen in production: a search that finds nothing, a tool that keeps timing out, a tool call written as text instead of made, a long list to return. The scoreboard is in annex E.

Kolibri-1’s reasoning mode keeps retrying when a tool returns nothing. In the two scenarios where the search finds nothing or keeps timing out, Kolibri-1 with reasoning at its maker’s evaluated level and recommended sampling retried the search with new wordings until the harness’s round limit and, since that harness has no fallback answer, gave no answer in 6 of 10 English trials. The calls vary; it is not the same call repeated. We have read the calls, not the reasoning, which the harness did not keep for those trials. At a lower reasoning level the failures halved, to 3 of 10. Without reasoning they almost vanished in English (1 of 20 sampled trials, none at temperature 0), but not in German (6 of 20). Qwen3.6 with reasoning, at its recommended sampling, answered 10 of 10 English trials and 9 of 10 in German. Our product’s loop behaves differently from the harness: it forces a final answer when its budget runs out, and 77 of the 78 answers that reached that budget still answered. The one runaway in the main test, a Kolibri-1 question that took 47 minutes and nine rounds and then gave no answer, ended before that budget was reached.

The agent scenarios also ran with reasoning at temperature 0, because our run-off script runs every scenario at both samplings; the question runs did not. Neither maker recommends that combination: both publish sampling settings for reasoning instead, and greedy decoding in a long chain of thought is known to fall into high-probability cycles. It did what sampling settings exist to prevent. In the two scenarios, Kolibri-1 gave no answer in either (five near-identical trials each, so two observations rather than ten), and in five of the ten trials it reasoned into the 32,768-token cap, the longest after 56 minutes. Qwen3.6 answered all of them. We report it only because a proxy or an application can send that combination by mistake: it is a deployment warning, not a score, and our product never sends it.

A larger token budget does not contain this: the runaways went past 32,768 tokens. What contains it is a budget per answer in time or tokens, a limit on repeated searches that return nothing, and a fallback that answers from what was found. Our loop already has a budget of ten rounds and sixteen tool calls, after which it must answer without tools; 77 answers reached it, and all but one still produced an answer. A time budget is the next addition.

We read this as an early-days exception rather than a verdict. Kolibri-1 is days old and its server plugin newer still. Qwen3.6’s own early faults, which did not show up in this test (below), are the precedent. Models settle, servers catch up, and the application learns what to guard against. What matters is to watch, measure and pre-empt such behaviour with response strategies, which is what the budgets above are.

Language slips, in both directions. Asked in German without reasoning, Kolibri-1 answered in English in 15 of 119 search answers; Qwen3.6 in 4. In the agent scenarios with reasoning, Kolibri-1 answered English prompts in German in 5 of 11 cases where the tool had nothing or failed, with the instructions and the question in English and the tool results in Spanish, the language of our mock data; German was in none of its inputs. A caveat our setup owes Kolibri-1: our answering instructions are in English whatever the question’s language, and 34 of the 48 documents are English. A German question thus arrives with English instructions. That is realistic for a multinational, and it is our product as built; it is not Kolibri-1 in its home language. The markers did not grade language: the rubric has no language criterion, and no grade note mentions it. The slipped answers did score lower than the rest, which suggests the slip accompanies a hard question rather than causing a markdown.

Qwen3.6’s old faults did not show up. On the 4-bit build tested before this one we had seen tool calls written into its reasoning and endless reasoning. On the matched FP8 server, none of its reasoning texts contained tool-call markup and none of its 160 agent trials with reasoning wrote a call as text. We cannot say which change fixed it: weights, server version, cache precision and speculative decoding changed at once.

Qwen3.6 has one small fault of its own. Asked in German without reasoning, it started 10 of 119 search answers with a stray <tool_call> tag, then answered correctly. Users see that tag unless the application strips it. None of its 119 English answers had it. Qwen3.6’s card notes that a higher presence penalty, which we set at its recommended 1.5, “may occasionally result in language mixing and a slight decrease in model performance”; its few wrong-language answers and these tags should be read with that in mind.

Running it: one DGX Spark per model

Each model ran on its own DGX Spark: a desk-sized workstation with NVIDIA’s GB10 Grace Blackwell chip and 128 GB of memory shared by processor and graphics, mounted in our server racks. Kolibri-1 in FP8 takes about 74 GiB of that, which still leaves room for nine conversations at full length.

Installing it was straightforward. Aleph Alpha ships a plugin for the inference server with the model, and Kolibri-1 answered its first questions on a DGX Spark on the day of its release, with no change to our code. The model came well prepared, and it happened to fit our machines. With Qwen3.6, at the time, it did not go that way. The server release that supported it was a release candidate; the only build that fitted was a third party’s 4-bit compression; and that build needed a patched tokenizer file, which cut every input at 4,096 tokens until we found the cause. That cost us days. Part of it was the server software of the time, which has since caught up.

On one DGX SparkKolibri-1Qwen3.6
Memory taken by the model73.6 GiB34.2 GiB
From start to first answer11 minutesabout 8 minutes
Cache for conversations2.4 million tokens7.0 million tokens
Conversations at the full 262,144-token window9about 27
Model rounds / tool calls per search answer, reasoning off, typical (median)4 / 52 / 2
Prompt tokens read per search answer, summed over its rounds (about 14,000 and 9,000 per round)about 86,000about 53,000

Qwen3.6 holds about three times as many long conversations on the same machine. Both models keep a cache that grows with the conversation in only 10 of their layers; the difference is mostly room. Qwen3.6’s weights take half the memory, which leaves twice as much for the cache. The rest comes from how the server lays out each cache, which we have not analysed. To keep the comparison fair, we limited both servers to eight simultaneous requests, the most Kolibri-1’s cache holds at the full window. Nine conversations is the extreme case, each filling the whole 262,144-token window at once. Most conversations are far shorter: a search round in this test is about 14,000 tokens, and at that size the same cache holds well over a hundred of them. The test itself ran only four questions at a time per server, beside the agent scenarios; the cache was never the constraint, and only when Kolibri-1’s server had more than eight requests in flight did the request limit queue them, which is part of the load caveat on every time in this article.

These figures are out of the box. Nothing was tuned by us; the server itself ships a tuned kernel configuration for Qwen3.6’s layer shape on this chip and none for Kolibri-1’s, so any speed difference favours Qwen3.6 for that reason too. The server also warns that its memory-saving cache format runs uncalibrated on both models. There is room to go faster, and possibly to be a little more accurate.

What we left out: pictures

Qwen3.6 can also look at images; Kolibri-1 reads text only. In a document product that matters: photographed delivery notes, scanned procedures, receipts.

We left it out of the comparison on purpose. The three scanned documents in our library were read once, in advance, by an earlier 4-bit build of Qwen3.6 running on our own hardware, and stored as text. Both models then answered from exactly the same text. Kolibri-1 was not marked down for not seeing, and Qwen3.6 got no credit for it.

Whatever the result, reading scans and photos needs a vision model. Some models do both jobs, Qwen3.6 or some of Mistral’s models, for example. With a text-only model such as Kolibri-1, the work is split between two models, one for pictures and one for text, which is no problem on hardware like this: in this test, each had its own machine. The question here is narrower: which model reads, reasons and answers better over text.

One consequence is worth stating. Had the vision model misread a scan, both models would have inherited the mistake. The fact check described above confirmed that every key fact survived the reading.

How we kept it fair

Comparing two models is only meaningful when everything except the model is the same. Our first pass, on 3 October, did not meet that bar: Qwen3.6 was running a third-party, 4-bit compressed version on an older server release, with half the window. It had Kolibri-1 ahead. So we rebuilt Qwen3.6’s server to match Kolibri-1’s and reran every test, for both models. Every model finding in this article comes from that rerun.

In the rerun, both models ran:

  • on the same machine type (DGX Spark), the same inference server (vLLM 0.29.0) and each maker’s own FP8 release;
  • with the same window, the same cache settings, the same memory share and the same limit on simultaneous requests;
  • through our real answering pipeline, with the same prompts, the same search index, the same limits on answer length and the same safeguards, swapping only the model.

They differed only where the model requires it: the format in which each one writes tool calls, its switch for reasoning, and the sampling settings its maker recommends. We ran each model with those settings, with the fixed setting our chat uses, and with a low temperature, our candidate default. With reasoning on, only the makers’ recommended sampling was used: neither maker evaluates reasoning at temperature 0, and our product does not send that combination.

Each question was asked once at the fixed setting, where the answers repeat, and three times at each of the other two; with reasoning, twice for each. The answers were shuffled, stripped of the model’s name and marked blind against the answer key, twice, by two AI models from two different makers.

The exact settings, and the reason behind each, are in the annexes at the end, together with what fits in one DGX Spark.

What this test is not

This is a first encounter with a brand-new model, not a benchmark: a first try, made with as much discipline as we could bring to it, and best treated as one more data point. Read it with these limits in mind.

  • It is small. Seventeen questions, one fictional company, one to three runs per setting. Differences of a few hundredths per answer are noise. The difference with reasoning rests on two runs, 68 answers per model.
  • AI wrote the test and AI marked it. Claude Opus 5.5, at extra-high effort and under our guidance, wrote the library, the answer key and the tests. Each answer was then marked correct, partial or wrong against that key, blind, by two markers: agents running Claude Opus 5.5 at extra-high effort, the same model that wrote the test, and OpenAI’s GPT-6.1 Sol at high reasoning effort, in Codex. They agreed on 78% of grades and never put one answer at correct and the other at wrong. We did not settle their disagreements; every finding is reported under both, and where the two differ in degree we say so: Qwen3.6’s lead with the whole library in the prompt holds under both, its smaller lead when handed only the right documents is within noise under both, and Kolibri-1’s lead with reasoning sits at the edge of what 17 questions can show. Blindness was stronger for the second marker than for the first (annex D), and no person has graded an answer yet; fifty blind human grades are the next step. Neither contestant comes from either maker. The second marker lowers the risk that author and marker share blind spots; it does not remove it.
  • The pipeline knew Qwen first. Our prompts were first written with Qwen as an incumbent, which may favour it.
  • Languages. Kolibri-1 officially supports German and English. Our library also contains Spanish, Polish and Chinese documents, as a real group’s would. Questions were asked in English and in German, with English instructions in both cases. Spanish comes next.
  • Images were not compared. Both models answered from the same text (see above).
  • Two questions need laws the library does not contain. Our answer key expected the model to know them, Chile’s 42-hour week and German co-determination, and both models nearly always failed: one question was wrong in all 79 answers, the other in 72 of 78. Arguably the better answer would have been to say the library does not cover it. That is a flaw in our key rather than in the models, and the two questions are reported apart.
  • One question asks for a summary, of the AI use policy’s rules for confidential data. The first marker accepted almost every answer from both models; the second marker called most of them partial for missing secondary rules. Summarising as a task of its own was not tested.
  • One answer key was wrong. The key for the question about a first order before a supplier’s conflict-minerals report expected a rule that lives in the document we held back. A model that said the documents do not mention it was marked wrong, and a firm “no” was marked right, which favoured Qwen3.6 on that question. The key needs correcting and the question regrading; the figures above include it. Without it and the two outside-law questions, the tie with search holds (Kolibri-1 ahead by 0.04 to 0.06 per question, within noise) and Kolibri-1’s lead with reasoning is slightly larger.
  • The cache precision favours Kolibri-1, if it favours anyone. Both servers keep the conversation cache in FP8, which is Kolibri-1’s evaluated setting and not Qwen3.6’s. If the format costs accuracy, it costs Qwen3.6’s; the details, and the control run we owe, are in annex D.
  • The agent scenarios were built from Qwen’s failures. The eight run-off scenarios replay faults we had seen with Qwen3.6’s earlier build and with another model. They target weaknesses Qwen3.6 is known for; a model that fails in other ways passes them unseen.
  • Not measured yet: raw writing speed on quiet servers, and a library with documents deliberately missing. Both are on the list below.

What we take from it

Four things, in the order a buyer would ask them.

  1. It fits, and it installed without fuss. Kolibri-1 runs on a single DGX Spark, from the first day, with its maker’s plugin and no change to our code.
  2. It answers as well as Qwen3.6 on the path our product uses. With search and without reasoning the two are level, under both markers and at every sampling setting. Qwen3.6 is one of the most widely run open models, so matching it is not a small thing. Kolibri-1 gets there with more rounds and more tokens per answer, so expect it to be slower, and the same machine holds a third as many long conversations. With reasoning on it was slightly ahead in both runs, not yet beyond chance, at several times the wait. In agent work its reasoning mode can keep retrying an empty search until a budget stops it; that is something to budget for, not a verdict.
  3. So the choice becomes a supply-chain question. When quality is level, what decides is who makes the model, where, under which licence, how it is maintained, how it handles your languages, and what capacity it costs per machine. A company that wants its AI supplier in Europe, with German handled by design, can now choose on those grounds without giving up quality on this job.
  4. Pictures still need a second model, and the hardware for it. Kolibri-1 reads text only. Scans, photographed delivery notes and receipts need a vision model beside it, and on a DGX Spark that means a second machine or a smaller model: a vision model of Qwen3.6’s size does not fit next to Kolibri-1 on one Spark.

After this first look, the picture is promising, and Kolibri-1 belongs on the shortlist of any company that runs its documents through a private model. Practice has to confirm it, and that is the part we are looking forward to: Spanish is next, then more questions, more documents and more models in the same frame, and of course use with customers.

Beyond the two models, four lessons hold whatever you run:

  1. Searching is cheaper, not more accurate. On a library that fits in the prompt, a model with everything in view answered as well or better in three of four settings. The case for search is cost and scale: up to three times fewer tokens per question, and it keeps working when the library no longer fits.
  2. Coverage is part of the answer. When the search missed the document that held the answer, every one of 34 answers was a confident partial one. None said “I don’t know.” An assistant needs to know, and say, which documents it is working from, and the search needs checking on your own library before anyone trusts it.
  3. Budgets, not bigger caps. A reasoning model that will not stop is contained by a limit per answer and a fallback, not by giving it more room. Behaviour under failure matters as much as accuracy under success.
  4. Test on your own documents, glitches included. Every document library has its own contradictions. A model that does well on public benchmarks can still take the first value it finds.

Choosing a model is choosing a supplier

Performance is only part of the decision. A company that runs a model on its own hardware is also choosing whom it relies on for the next version, for fixes, and for how well the model will handle its languages. Many of our customers already think this way about their other suppliers: where a component comes from, under which rules it is made, whether there is a second source. An AI model is becoming a component like any other.

There are now serious open models from American, Chinese and French makers, and with Kolibri-1 a German one. That is what we want to give our customers: options. Provenance and maintenance belong on the list, next to accuracy, speed and cost: who makes the model, where, and who will keep it up to date. For many of the companies we talk to, these are not side questions.

Open weights soften the dependency: once you hold a model under a licence like Apache 2.0, it stays yours to run. The rest is architecture. In Iristrace Docs the model sits behind a router, and replacing one with another, as we did for this test, changes the model and nothing else.

The questions we would ask of any model:

  • Does it do the job on your documents, in your languages?
  • What hardware does it need, and what does that cost to run?
  • Under which licence, and what exactly does the licence cover?
  • Who makes it, where, and how likely are they to keep improving it?
  • How easily could you switch to another model later?

What comes next

This was a first look, and it raises as many questions as it answers. The next steps, in order:

  • Summarising, as a task of its own. A policy summarised for a new employee, what changed between two revisions, a Polish or Chinese work instruction summarised in German, checked for facts kept, figures invented, language and length.
  • A library with gaps. The same questions with the key documents deliberately removed, to measure how often each model says “the documents I have do not say” rather than guessing.
  • A rule for contradictions. We will test an instruction that puts a contradiction before the conclusion, for both models and on new questions, so that we do not tune the prompt to these seventeen.
  • Spanish, put through its paces. Many of our customers work in Spanish. Kolibri-1 officially supports German and English, so Spanish gets a thorough grilling next: questions, documents and answers.
  • German from end to end. In this test our answering instructions were in English even when the question was German: realistic for a group, but not Kolibri-1 in its home language. Next it gets German instructions for German questions.
  • Extraction and structured data. Turning invoices, delivery notes and scans into structured data cannot do without a vision model. Questions about records (how many invoices came from one supplier, which one was the largest) are then answered by structured tool calls over that data, and AI agents will be able to work with Iristrace directly through an MCP server. All of this needs its own research and measurement: failure scenarios, many repetitions, reasoning on and off. That takes far more than a weekend.
  • More models, on the same footing. Google’s Gemma 4, OpenAI’s gpt-oss and Mistral’s Ministral, to name a few, served under conditions as identical as possible, so that the results stay comparable.
  • Long-running agents. Aleph Alpha also positions Kolibri-1 for agentic work. Agents that work through a task over many steps, for minutes rather than seconds, are what we want to see for ourselves next, starting with Kolibri-1 at a lower reasoning level and with a hard limit on repeated searches.

And to turn a first encounter into something closer to a measurement, the test itself needs to grow: more questions, some written by people who never saw the answer keys; a larger library with more of the contradictions real libraries collect; more runs; and a share of the AI markers’ work reviewed by people, so that the judges are not only machines.

We are curious and open-minded, and we love working with this technology.


Kolibri-1 runs on a DGX Spark, then, and holds its own there. Applause for Aleph Alpha, and thanks to Heidelberg, for a strong new option in private AI on hardware a mid-sized company can afford. For such companies, that is one more real choice beyond the American and Chinese heavyweights: a capable model of European origin.

If you would like to know more, or to find out whether a private AI setup like this one could serve your operations, get in touch. Alongside our two product lines, Iristrace Checks and Iristrace Docs, we offer training and advice. We would be glad to hear from you and take it further together.

About this article. A first-hand report by Iristrace on a two-day test using AI, written internally and not reviewed externally: not a benchmark and not a general ranking of the models. We have no partnership or other affiliation with Aleph Alpha or Alibaba, and we received no compensation for this article; our only tie to NVIDIA is that we bought the DGX Spark machines as an ordinary customer. None of the three companies saw this text before publication. The test documents belong to fictional companies; every company and person named in them is invented. AI models wrote the test library, marked the answers, ran the statistics and helped write this text; Iristrace reviewed it and is responsible for it. We build and run private AI setups, with whichever model suits the customer.

Annexes, for the technically minded

Annex A — How the two servers were set up, and why

Both models were served by vLLM 0.29.0, each on its own DGX Spark, behind the same LiteLLM proxy and the same answering pipeline. Every setting that had to match was chosen for a reason, written down before the runs.

Same on both serversValueWhy
MachineNVIDIA DGX Spark: GB10 Grace Blackwell, 128 GB unified memoryOne per model, so neither competes with the other for memory
Inference servervLLM 0.29.0The release Aleph Alpha’s plugin supports. The plugin follows one vLLM version at a time, so both servers move together
WeightsFP8: e4m3, 128×128 blocks, dynamic activationsEach maker’s own release, in the same format. No third-party compression
Context window262,144 tokensBoth models’ native maximum. The whole library in one prompt needs about 160,000
KV cacheFP8Aleph Alpha recommends it. At full precision Kolibri-1’s cache would hold half as much
GPU memory share0.9The Spark’s memory is shared with the operating system; this leaves it about 12 GiB
Simultaneous requests8The most Kolibri-1’s cache holds at the full window (9.04), rounded down, so no request is ever paused for lack of memory. Qwen3.6 could hold more; it gets 8 so neither server runs more in parallel
Prompt tokens per scheduling step16,384Our prompts are long (20,000 to 160,000 tokens). Larger steps read them faster, at a small cost to other users’ streaming
Prefix cachingOnA prompt that starts the same way is not read twice. We keep the start of our prompts identical: the date and the question come last
Speculative decodingOffLossless in principle, but it pushes several tokens per step through the streaming tool-call parser, where we have seen trouble before. Speed is not worth a moving part there
Questions in flight4 per serverTimes in this article include waiting behind other work, and Kolibri-1’s server carried more of it; times compare within a model, not between the two
Different by designKolibri-1Qwen3.6
Tool-call and reasoning parserskolibri1qwen3_coder, qwen3
vLLM pluginAleph Alpha’s aleph-alpha-inferencenone
Recommended samplingtemperature 1.0, top-p 0.97, top-k 128reasoning: 1.0, 0.95, 20; otherwise 0.7, 0.8, 20; presence penalty 1.5
Reasoning switchreasoning_effort, run at high, the maker’s evaluated levelenable_thinking
Kernel overridenoneDeepGEMM off: vLLM warns that its FP8 scale format degrades this architecture on Blackwell

Three settings differ from our product. The reasoning cap was 32,768 tokens per round in the test; the product uses 16,384. The test ran four questions in flight per server. And with reasoning on, the test sent Qwen3.6’s presence penalty of 1.5, which the product’s reasoning path does not; in that one respect the product differs from what we tested. A small script reads both running servers and fails if any of seven settings differs: the vLLM version, weight format, context window, cache precision, memory share, prefix caching and speculative decoding. Three more it cannot see from outside, the number of requests served at once, the prompt tokens read per step and the kernels each server chose, we checked in each server’s start-up log. Both logs also show the same machine.

Annex B — What fits in one DGX Spark

The model card says, verbatim: “Model memory footprint: ~78 GB (FP8 weights). Minimum: 2× A100 80 GB, 2× H100 SXM5, 1× H200, 1× B200 or 1× B300. Recommended: 2× H100 SXM5, 2× H200, 1× B200 or 1× B300.” A DGX Spark has 128 GB of memory, shared by processor and graphics. The card’s ~78 GB and the 73.6 GiB in the table below are the same amount in different units. This is how Kolibri-1 fitted, from its final start-up log:

Kolibri-1 on one DGX SparkGiB
Memory visible at start121.7
vLLM’s share (0.9)about 109.5
Model weights in memory, FP873.6
KV cache, FP8: 2,370,457 tokens, 9.04 conversations at the full window34.1
Working memory for the computation itselfabout 2
Left for the operating system and everything elseabout 12
Side by sideKolibri-1Qwen3.6
Parameters, total / active per token78.1 billion / 3.46 billion35 billion / 3 billion
Weights in memory, FP873.6 GiB34.2 GiB
KV cache2,370,457 tokens7,014,337 tokens
Conversations at the full 262,144-token window9.0426.76
From start to first answer11 minutesabout 8 minutes

Why does Qwen3.6 hold three times the conversation? Not because its cache grows in fewer layers: both models keep a cache that grows with every token in only 10 layers (Kolibri-1 in 10 of 50, the other 40 looking back 513 tokens; Qwen3.6 in 10 of 40, the other 30 using linear attention). The difference is mostly room. Qwen3.6’s weights take half the memory, which leaves twice as much for the cache. The rest comes from how the server lays out each cache, which we have not analysed.

Both servers read their weights from disk without read-ahead, at about 120 to 145 MB/s. Kolibri-1’s 74 GiB take about nine of its eleven minutes; Qwen3.6’s 34 GiB take five. In practice a restart of the model server is a planned event, not a blink.

One question remains open. Before its final restart, with the same memory share, Kolibri-1’s server reported a cache of 3.38 million tokens; after it, 2.37 million. The arithmetic narrows it down: the final cache costs about 15.4 KB per token (34.1 GiB for 2.37 million tokens), and at that rate 3.38 million tokens would need 48.6 GiB, more than the 36 GiB left beside the weights. The earlier server must therefore have laid out its cache at a lower cost per token. The capacity did not disappear; the accounting changed with the restart, and which setting changed it is what we have not yet pinned down.

Annex C — The whole library in one prompt

Both models read up to 262,144 tokens natively. Our whole library is about 160,000, so it fits in a single prompt with room to spare. Both model cards describe longer windows of about a million tokens: Kolibri-1’s card says the maker validated quality up to 1,048,576 tokens, Qwen3.6 reaches it with YaRN scaling. We did not test those; both ran at their native window, and we probed its edge.

The edge of the window, on idle servers. Three short values are planted at the start, the middle and the end of an archive of numbered records, and the model is asked for all three. Each prompt opens with a fresh random line, so the server cannot reuse an earlier read; the times are cold, without reasoning, at temperature 0. This was the one test run with nothing else on either server, so its times are comparable between the two.

Prompt tokensKolibri-1: seconds, all three foundQwen3.6: seconds, all three found
about 64,00015 to 16, in 4 of 4 tries17, in 3 of 4 tries
about 129,00045, yes46, yes
about 196,00087, yes89, yes
about 247,000130, yes131, yes
about 260,000145, yes146, yes
266,144, over the windowrefused, with a clean errorrefused, with a clean error

Both recall all three values right up to the edge of the window. The one miss was Qwen3.6’s first try at 64,000 tokens, which returned three record numbers instead of the planted values; three repeats at that size were right. Both read a long prompt at the same speed, within about two seconds of each other at every size: 3,700 to 4,200 tokens a second at 64,000 and about 1,800 at 260,000, so reading the whole window once takes about two and a half minutes. Over the window, both refuse cleanly with an error that says why, which our gateway turns into a message to the user. The test is narrow, three exact strings in uniform filler: it says the full window works and what it costs to read, not how well a model reasons over a long, varied set of documents. The whole-library condition measures that.

What we did observe, over about 370 whole-library prompts of 145,000 to 159,000 tokens:

  • The first read costs a minute. 62 to 68 seconds to the first token on either server, the first time a library is read.
  • The second read is the cheap one. When a long prompt starts exactly like an earlier one, the server reuses what it has already read and only reads the rest. After the first read, answers started within two to four seconds. That is why our prompts keep the documents first and the date and the question last, and why it matters that a model’s chat template adds nothing that changes from one request to the next.
  • Long prompts are where repetition shows. At the fixed setting without reasoning, six of 68 whole-library answers started repeating themselves or ran into the output cap, more than under any other condition. Our repetition stop and output cap caught them. The makers’ own sampling recommendations saw almost none.
Annex D — Choices that can be questioned, and our reasons

A reader who runs models for a living will ask about the settings before the scores. These are the questions we would ask ourselves, each with its answer and, where one is owed, the control we still have to run.

Why these two models? Nearly equal active parameters, 3.46 and 3 billion per token, put them in the same speed class on the same hardware. Both fit one DGX Spark in FP8, both are Apache 2.0, both offer reasoning and tool calling at a 262,144-token window. They differ where a buyer wants to know the price of the difference: Kolibri-1 has more than twice the total parameters and the memory to match; Qwen3.6 covers more languages and reads images.

Why FP8, and each maker’s own release? Full precision does not fit: Kolibri-1’s bfloat16 weights need about 156 GB. The third-party 4-bit Qwen3.6 tested before keeps 93% of full precision on multi-turn tool calling by its own maker’s evaluation, which is why our first pass was unfair. Each maker’s own FP8 release is the highest precision both can run at here. One asymmetry remains: Kolibri-1’s last training stage was quantisation-aware, and Qwen3.6’s FP8 was made after training. Qwen’s card reports its FP8 “nearly identical” to the original; we did not measure that ourselves.

Why an FP8 conversation cache on both? Because Aleph Alpha serves and evaluates Kolibri-1 that way, and because at full precision Kolibri-1’s cache would hold half as much. Qwen3.6’s maker does not document that setting, and the server warns that the scales are uncalibrated on both. If the format costs accuracy, it costs Qwen3.6 more than Kolibri-1. The exposure is bounded, 10 of Qwen3.6’s 40 layers against all 50 of Kolibri-1’s, and a full-precision cache would have been feasible for the test alone: Kolibri-1 would still have held the four whole-library prompts in flight. The control we would run first is Qwen3.6 with a full-precision cache, with search and reasoning off, about a hundred answers.

Why reasoning at high? It is the level Aleph Alpha evaluates and publishes, and Qwen3.6’s reasoning has no level to choose. Comparing each at its maker’s evaluated configuration is the defensible choice, and it means the five minutes per answer are Kolibri-1’s most expensive mode. In the agent scenarios, medium halved the failures to answer; on the questions it has not run yet. It is the second control we owe.

Why three sampling settings without reasoning, and only one with? Temperature 0 is what our product sends; the makers’ recommendations are what they evaluate; 0.3 with top-p 0.9 is our candidate default, sent to both. Kolibri-1’s card gives a single setting, temperature 1.0, with no separate one for reasoning off, so its “recommended” cell runs a mode its maker has not scored; Qwen3.6’s 0.7 is documented, and its card adds that its presence penalty of 1.5, which we used, “may occasionally result in language mixing and a slight decrease in model performance”. Two honest footnotes. First, a request names only some parameters; the rest fall to each server’s defaults, its maker’s top-k, and, for Qwen3.6, to the presence penalty set on its proxy alias, so at temperature 0 and at 0.3 the two were equal in what we set and differed in what we did not, and Qwen3.6’s temperature 0 was not plain greedy. Second, we have not yet verified end to end that every parameter survives the proxy; a ten-minute probe is scheduled. Neither footnote touches the tie with search, which holds at all three settings, the makers’ own included. With reasoning, only the makers’ sampling ran: neither recommends greedy decoding for reasoning, both publish sampling settings instead, and our product never sends that combination. The run-off script ran it anyway; the result is reported above as a deployment warning.

Why one run at temperature 0 and three at the others? Temperature 0 repeats itself almost word for word, so three runs would be one sample copied. At sampled settings each run is a fresh draw, so that is where the repeats are.

Why a 32,768-token reasoning cap when the product allows 16,384? Qwen3.6’s card advises 32,768 “for most queries”, and at 16,384 Kolibri-1 had run out mid-reasoning on two questions in our first pass. Every answer records its tokens, so the data also says what 16,384 would have cut: 13 of Kolibri-1’s first 114 reasoning answers.

Why four questions in flight, and why no speed verdict? To finish in the time we had. Both servers hold eight; both ran four questions at once, and both carried the agent scenarios and repeat runs beside them, Kolibri-1’s server more and for longer: its question runs overlapped with other work all evening, while Qwen3.6’s second reasoning run ran alone. Qwen3.6’s own reasoning answers were 15% slower in its loaded run than in its quiet one, and Kolibri-1’s medians without reasoning ranged from 38 to 91 seconds across runs. Seconds per answer are therefore comparable within a model, not between the two, and the article makes no speed claim beyond what load cannot explain: more rounds and more tokens per answer for Kolibri-1. Single-request speed on quiet servers is a separate measurement, still to run.

Why our own pipeline rather than a neutral harness? Because the question is how each model does in our product. The cost is stated: the answering prompt was refined while Qwen3.6 was the model behind it, so it may suit Qwen3.6 better.

Why three ways of giving the evidence? To separate finding from reading. With search, a wrong answer may be the retriever’s fault. With only the right documents in the prompt, it is the model’s. With the whole library, it is the model’s again, under the load of 160,000 tokens.

How the comparison is computed. Answers to the same question are not independent: a hard question is hard everywhere. So each model’s score is averaged per question first, and the 17 questions are compared in pairs. The mean difference per question, Qwen3.6 minus Kolibri-1, with a 95% bootstrap interval over the questions:

SettingFirst markerSecond marker
With search, reasoning off, all three samplings+0.00 (−0.12 to +0.12)−0.03 (−0.15 to +0.08)
With search, reasoning off, temperature 0.3+0.01 (−0.14 to +0.14)−0.03 (−0.16 to +0.09)
With search, reasoning on−0.07 (−0.15 to −0.02)−0.08 (−0.17 to +0.00)
Whole library in the prompt, reasoning off+0.16 (+0.08 to +0.24)+0.11 (+0.03 to +0.19)
Only the right documents, reasoning off+0.07 (−0.02 to +0.17)+0.04 (−0.06 to +0.14)

Read plainly, Qwen3.6’s whole-library advantage clears zero under both markers, and Kolibri-1’s advantage with reasoning ends at zero under the second marker and only just clears it under the first; a rank test over the 17 questions finds it significant under neither. Read strictly, we ran at least eight such comparisons, so a reader should widen the intervals accordingly, to about 99.4% in place of 95%; at that standard only the whole-library difference under the first marker still clears zero, and everything else, Kolibri-1’s reasoning lead included, is a direction the next runs have to confirm. Without the two outside-law questions and the question whose key was wrong, the figures move slightly towards Kolibri-1 and the picture does not change. The tie with search is a tie under both markers. Treating the 1,428 answers as independent would shrink every interval and overstate the evidence; we do not.

Question by question. The gaps on two questions looked too large to be chance. To check, we asked Claude Opus 5.5, in Claude Code, to test every question, not only the striking ones. It compared the 14 answers per model on each question, with search and without reasoning, in an exact permutation test: if the model made no difference, its name would be an arbitrary label, so the test counts, over every way of splitting the 28 answers into two groups of 14, how often a gap as large as the real one appears. It assumes nothing about how the scores are distributed, which matters with three possible grades and fourteen answers a side. Seventeen questions tested at once need a correction: at the usual 5%, the chance that at least one looks significant by luck alone is about 58%. It corrected with Holm’s method, which keeps the chance of even one false claim across all seventeen at 5%, holds however the questions depend on one another (they share models, runs and library), and never finds less than the simpler Bonferroni correction. Under both markers two differences survive, each with a corrected p below 0.01: question 5, the customer’s 90-day rule (0.86 against 0.14), and question 8, the grease grade no longer approved (0.61 against 1.00). Question 17, the one whose key is wrong, survives under the first marker only (about 0.01; 0.07 under the second). Bonferroni gives the same three. The script is kept with the dataset.

Are the samples big enough? For what we claim, yes; for more, no. A test with a correction keeps false alarms at 5% at any sample size; a small sample only hides smaller differences. After the correction, fourteen answers a side reliably reveal gaps of roughly 0.4 to 0.65 points per answer, so a question without a significant difference may still differ (the per diem, at 0.21, is one candidate). Each finding also rests on one question: its 14 answers repeat it across runs and languages, so they show the behaviour is stable on that question, not that it holds for every question of its kind, and because they are not fully independent the p-values lean optimistic; the two differences are large enough that this does not change the reading. And the overall tie is a tie within about ±0.12 per answer: telling apart models that differ by 0.10 takes about fifty questions, by 0.05 about two hundred.

Why AI markers, and why two? 1,428 answers in the time we had leave no other option, and a rubric fixed in advance, blind marking and a second marker from another maker are the known mitigations. Their agreement, 78% with Cohen’s kappa 0.68, counts as substantial by the usual scale, and the disagreements run one way: the second marker reads the answer keys’ secondary details as required. That vagueness is in our rubric, not in the markers, and it falls on both models alike. Blindness was stronger for the second marker than for the first: the second worked in a folder that held nothing but the rubric, the questions and the answers under new keys; the first marker’s agents were instructed not to open the key, which sat in the same folder tree, and Qwen3.6’s stray tag survived into ten of the blind answers as a fingerprint. For the next round the key moves out of reach and the tag is stripped with the formatting. A person reading a sample of the marks is the next step, and the one no second machine replaces.

What would make us rerun rather than add? Nothing we found. The controls above add cells; none invalidates one. What would force a rerun is a settings error that touched one model only, and the matching script and both start-up logs were kept so that an independent reviewer can look for one.

Annex E — The data: question by question, scenario by scenario, cost per answer

The 17 questions. Mean score per answer with search and reasoning off, correct 1, partial ½, wrong 0, over 14 answers per model and question (seven runs, English and German); first marker, second marker in brackets. The last column counts the searches, of 36 per question, that returned every document the answer key names.

Question, and what makes it hardKolibri-1Qwen3.6Key documents found
1. Who can approve a €120,000 capex request at the Tychy plant?
Policy and approval matrix give different limits; euro against złoty
0.61 (0.68)0.54 (0.64)34 of 36
2. What is the current hotel cap for business travel?
Two revisions, the old one still marked “Effective”
0.89 (0.50)1.00 (0.50)36 of 36
3. Which documents are overdue for review?
Needs every document scanned; the key one is a scan
0.00 (0.00)0.00 (0.00)8 of 36
4. What hardness range applies in Changzhou, and does it match the group spec?
Plant instruction departs from the group spec without a deviation
0.86 (0.86)0.79 (0.57)36 of 36
5. How much notice must we give HallvikDemo Trucks before a product change?
Customer matrix (90 days) stricter than the change procedure (60)
0.86 (0.89)0.14 (0.18)26 of 36
6. Are our Chile HR documents up to date with working-hours law?
Needs Chile’s 2026 law; outside the library
0.00 (0.00)0.00 (0.00)18 of 36
7. Does the USMCA origin analysis match where the rings are made?
Origin table says North America; control plan sources rings from China
0.82 (0.79)1.00 (1.00)36 of 36
8. Which grease grade are we buying from ChemcoDemo, and is it the approved one?
Contract names a grade the specification withdrew
0.61 (0.61)1.00 (1.00)36 of 36
9. Does the remote-work policy apply in Schweinfurt as written?
Needs German co-determination; outside the library
0.00 (0.04)0.07 (0.11)36 of 36
10. List all documents that reference a non-existent document.
Needs every document scanned
0.00 (0.00)0.04 (0.04)33 of 36
11. Summarise the AI use policy’s rules for confidential data.
A summary; the second marker required every secondary rule
1.00 (0.54)1.00 (0.57)36 of 36
12. Can a Grade C engineer spend the full Mexico per diem without approval?
Per diem exceeds the Grade C limit in the approval matrix
0.96 (0.96)0.75 (0.79)36 of 36
13. How long must we keep traceability records for SolanoDemo Automotive parts?
Customer requires 20 years, procedure says 15; the matrix was mostly missed
0.25 (0.25)0.00 (0.00)9 of 36
14. May I accept a €75 gift from a supplier?
Procurement policy (€50) against group code (€100)
0.75 (0.79)0.79 (0.75)32 of 36
15. What is the title of Iker Olaizola Mendia?
Two titles in two documents; the 8D procedure was never found
0.50 (0.50)0.50 (0.50)0 of 36
16. What happens if a supplier does not return its conflict minerals report?
The held-back document; the right answer is “not covered”
0.93 (0.89)1.00 (0.86)—
17. Can we place a first order with a new supplier before it sends its CMRT?
The key rewards the held-back document; being corrected
0.25 (0.18)0.75 (0.46)—

With reasoning on, four answers per model and question (two runs, two languages): Kolibri-1 scores higher on six questions (1, 2, 4, 5, 8 and 13), Qwen3.6 on one (10), and the rest tie. The per-question means are in our dataset.

The agent scenarios. Passes out of 10 trials per cell, five in English and five in German, at each maker’s recommended sampling, Kolibri-1’s reasoning at high. Scenarios 1 and 2 exist only with reasoning on.

Scenario, and what passesKolibri-1 offQwen3.6 offKolibri-1 onQwen3.6 on
S1 A tool call while reasoning
the call is structured, not text
——1010
S2 The same, with a call required
it still calls
——1010
S3 Twenty tickets, all with one date
the date is reported consistently
104910
S4 A list of invoices
the largest is named, and only it
1010109
S5 A maximum returned without its id
it fetches the id or says it cannot; never invents one
1091010
S6 A search that finds nothing
it says so and invents nothing
9959
S7 A tool that keeps timing out
it stops retrying and answers
810210
S8 A 25-item list
no runaway, no repetition
10677

Two cells need a footnote. Qwen3.6’s six failures in S3 without reasoning were not the model’s: asked in German, it used our mock aggregation tool, which answers “no data” for tickets, and reported that faithfully; a gap in our test tool. S8 is flawed as a test: it counts a refusal of the long list as a failure, which is a failure of a different kind. The checks are pattern matches on wording in two languages and were corrected during the runs; every trial keeps its verdict as run alongside the corrected one.

Cost and tail per answer. Tokens, rounds and calls are means over the answers with search. Prompt tokens are summed over an answer’s rounds: each round resends the instructions, the tool definitions and everything found so far, so a five-round answer reads its context five times. The means are pulled up by the answers that hit the round budget; a typical answer takes four rounds for Kolibri-1 and two for Qwen3.6. Seconds are over all answers, of every evidence kind, and were measured under our test load, which was heavier on Kolibri-1’s server; compare them within a model.

Per answerKolibri-1 offKolibri-1 on (high)Qwen3.6 offQwen3.6 on
Prompt tokens read, summed over an answer’s rounds, with searchabout 86,000about 75,000about 53,000about 61,000
Output tokens, with search, reasoning includedabout 710about 6,900about 780about 2,900
Model rounds / tool calls, with search, mean5.0 / 6.04.3 / 6.54.0 / 3.04.5 / 3.9
Seconds, median / 90th percentile / slowest, under test load31 / 114 / 508292 / 1,162 / 2,89422 / 70 / 408120 / 264 / 2,119
Answers over ten minutes0 of 51047 of 2040 of 5106 of 204
Output tokens, median / 90th percentile / most369 / 975 / 7,2844,393 / 17,037 / 48,493355 / 927 / 5,2932,839 / 5,506 / 32,768

Every row of these tables is one line in our dataset: one record per answer with both grades, and one per agent trial with its verdict.

Cookie settings

Choose which cookies you accept. You can change this at any time from “Cookie settings” at the foot of every page. Cookie policy