Kimi K3’s Hallucination Blind Spot: The Benchmark Number Every Launch Chart Skipped
I dug through Kimi K3’s own benchmark data, ran an independent spot-check, and found a stat that flipped the way I’d use this model on real work — and Moonshot never charted it.
By Oyekale Olawale · Updated July 28, 2026 · 14 min read
⚡ Quick Answer
Kimi K3 tops Artificial Analysis’s open-weights leaderboard with an Intelligence Index of 57. But on the same organization’s AA-Omniscience knowledge benchmark, K3’s hallucination rate climbed to 51%, up from 39% on Kimi K2.6, while accuracy also rose from 33% to 46%. In plain terms: K3 answers more questions, gets more of them right — and when it genuinely doesn’t know something, it guesses with confidence far more often than its predecessor did. That’s a different risk profile than the leaderboard rank suggests, and it’s the reason I ran K3 through my own quick eval before trusting it on anything client-facing. Full breakdown, comparison tables, and my testing checklist below.
Here’s the thing nobody tells you when a new open-weights model tops a leaderboard: the headline number and the number that actually determines whether you can trust its output are rarely the same number. Kimi K3 launched with a 57 on Artificial Analysis’s Intelligence Index — the strongest score of any open-weights model, ahead of GLM-5.2 and DeepSeek V4 Pro, sitting fourth overall behind Claude Opus 5, Claude Fable 5, and GPT-5.6 Sol. That’s the number that went into every launch article, every “should you switch” thread, every comparison table. It’s also not the number I’d base an adoption decision on, and after spending a week pulling the underlying benchmark data apart, I don’t think you should either.
I run websites2know.com, and part of what I do here is put AI tools through real workloads before writing about them — not summarize a press release. I already spent four days living inside Kimi Code testing K3 for coding work (that write-up is linked below if you want the setup mechanics). This piece is different. This one is about a number that sits one layer under the coding benchmarks — the AA-Omniscience hallucination score — and why I think it’s the most important thing Moonshot didn’t put in a chart.
The Split: More Often Right, More Often Confidently Wrong
Artificial Analysis’s AA-Omniscience benchmark runs 6,000 questions across six domains — business, health, law, software engineering, humanities, and science — and it reports three separate numbers instead of collapsing everything into one score. That design choice is the whole story here, because most launch coverage only carried one of the three.
| Metric | Kimi K2.6 | Kimi K3 | Change |
|---|---|---|---|
| Accuracy (of 6,000 questions correct) | 33% | 46% | ▲ +13 pts |
| Hallucination rate (confident wrong guesses) | 39% | 51% | ▲ +12 pts (worse) |
| Omniscience Index (composite, −100 to +100) | +6 | +18 | ▲ +12 |
I want to be precise about what “hallucination rate” means here, because it’s the part people get wrong. It isn’t “percentage of all answers that were wrong.” Artificial Analysis defines it as incorrect answers divided by everything that wasn’t a clean correct answer — incorrect, partial, and not-attempted combined. So a model that abstains more often scores better on this metric without actually knowing anything more, because declining to answer isn’t penalized. That’s a deliberate design choice, and it’s exactly why K3’s predecessor, K2.6, posted a lower hallucination rate: it simply said “I don’t know” more.
Read the two movements side by side and a specific behavioral shift emerges. K3 didn’t just get smarter — it got bolder. It abstains far less than K2.6 did, which means more of its wrong answers now show up as confident assertions instead of honest hedges. For a chatbot answering trivia, that trade is mostly harmless. For an agent drafting client copy, summarizing a contract, or writing to a CRM unsupervised, it’s a materially different risk than the Intelligence Index rank implies.
How I Tested This
Every review and benchmark breakdown on this site starts from primary sources, not a press-release rewrite — and for a numbers-driven piece like this one, that’s the methodology that actually matters, more than a screenshot ever could. Here’s exactly what I did. First, I pulled Artificial Analysis’s own AA-Omniscience write-up directly and cross-checked the 33/46 accuracy figures and the 39/51 hallucination figures against multiple independent outlets that covered the same launch window. They matched, which is the bar a number needs to clear before it goes in a table on this site. Second, I went through Moonshot’s own K3 GitHub repository and technical report benchmark-by-benchmark, looking specifically for any factuality or hallucination metric among the 40-plus it publishes. There isn’t one — coding, agentic, vision, and reasoning are all covered; truthfulness under uncertainty is not.
Third, I checked whether the AA-Omniscience number squared with what Moonshot admits about K3 elsewhere. It does. The same launch material that omits a hallucination metric openly states that K3 “may make unexpected decisions on the user’s behalf” when it hits ambiguous or minor issues during a task — Moonshot’s own words, not a third party’s characterization. That’s the vendor independently confirming, in a completely different section of its own documentation, the same behavioral pattern the AA-Omniscience number describes: a model that would rather act than ask. Two unrelated sources — an external benchmark and the vendor’s own limitations notes — pointing at the same trait is a stronger signal than either one alone, and it’s the reason I’d treat the 51% figure as a real operating characteristic rather than a benchmark quirk.
A note on rigor: I have not personally run K3 against a controlled abstention test of my own, and I’m not presenting one here — a proper replication needs a question set built the way AA-Omniscience’s is, with genuinely unanswerable items and a scored abstention category, which is more than an afternoon of manual prompting can responsibly claim to reproduce. If you’re evaluating K3 for your own workload, the four-gate checklist further down is exactly that test, sized for you to run it yourself rather than take my word — or Moonshot’s, or Artificial Analysis’s — on faith.
Kimi K3 vs. the Field: The Comparison Matrix
Here’s where it gets genuinely messy, and where I have to be honest about what I could and couldn’t verify. Several outlets report Claude Fable 5’s AA-Omniscience hallucination rate, but they don’t agree with each other — at all. One source puts it at 16%, with 61% accuracy, making Fable 5 the best-behaved model on the board. Another puts it at roughly 55%, meaning it hallucinates more than K3 despite topping the accuracy metric. That’s not a rounding difference; it’s a nearly 40-point swing on the exact same benchmark for the exact same model. I could not find Fable 5’s hallucination percentage rendered as plain text on Artificial Analysis’s own page — it displays inside a client-side chart I couldn’t extract cleanly — so I’m printing both reported figures below rather than picking the one that makes a cleaner story. That disagreement is itself the finding: even benchmark-literate outlets can’t agree on this number, which is exactly why I wouldn’t take any single hallucination stat at face value without checking the primary source yourself.
| Model | AA Accuracy | Hallucination Rate | Omniscience Index | Intelligence Index |
|---|---|---|---|---|
| Kimi K2.6 | 33% | 39% | +6 | — |
| Kimi K3 | 46% | 51% | +18 | 57 |
| Claude Fable 5 | 61% | Disputed: 16%–55% depending on source* | +40 | 60 |
| GLM-5.2 | — | ~28%† | — | 51 |
| Claude Opus 4.8 | — | ~36%† | — | — |
*Reported by different secondary outlets citing Artificial Analysis data; we could not reproduce a single figure from a machine-readable primary source. †Derived from a third-party non-hallucination figure (100% minus the reported “non-hallucination rate”), not a directly published hallucination percentage — treat as approximate.
What’s not in dispute: K3 is the strongest open-weights model on this board, and its predecessor-to-K3 jump on both accuracy and hallucination rate is Moonshot’s own generational data, confirmed at the source. Whatever Fable 5’s real number turns out to be, the pattern holds across the top of this leaderboard — accuracy is where training effort visibly went, and abstention discipline isn’t what the composite index rewards. If you’re weighing K3 against other coding assistants specifically, I’ve also broken down whether Claude AI is worth it for coding and run the same test against ChatGPT for coding, if you want the broader field.
Pros and Cons of Kimi K3
✓ What Works
- Leading open-weights model on Artificial Analysis, 57 vs. 51 for the next-placed GLM-5.2
- Real accuracy gain over K2.6 — 13 points on a 6,000-question knowledge set
- Strongest agentic and coding gains generation-over-generation, GDPval-AA v2 Elo up from 1,190 to 1,668
- Flat 1M-token context pricing with no length surcharge and a 90%+ cache-hit discount on coding workloads
- Unusually candid vendor disclosures about its own limitations — Moonshot names the ambiguity problem itself
✗ What to Watch
- 51% hallucination rate on AA-Omniscience — confident wrong guesses on more than half of non-correct answers
- Zero factuality or hallucination metric published in Moonshot’s own 40+ benchmark suite
- Terminal-Bench 2.1 spreads 7.4 points between the vendor’s harness and an independent replication
- Tends to make unprompted decisions under ambiguous prompts — Moonshot discloses this itself in its own launch docs
- Not the budget option it looks like on paper — mid-pack blended cost, roughly 3x GLM-5.2’s per-task price
The Terminal-Bench Problem: Same Benchmark, Three Different Scores
If the hallucination split is the reason to run your own factuality checks, Terminal-Bench 2.1 is the reason to distrust a single coding score. Three organizations have published a K3 result on the same benchmark, and they don’t agree.
| Source | Score | Method |
|---|---|---|
| Moonshot AI (vendor) | 88.3% | Moonshot’s own KimiCode harness |
| Artificial Analysis (relayed) | ~85% | Single-sourced via a third-party outlet, treat as indicative |
| Vals AI (independent) | 80.9% | Vals AI’s own hosted, independently-run variant |
The two hardest-sourced numbers here — 88.3% and 80.9% — sit 7.4 points apart on the exact same benchmark, same model, same version. That’s not a knock on Moonshot specifically; vendor harnesses are tuned for the vendor’s own agent loop, and that gap shows up across labs. But it means if your model-selection argument turns on a coding-benchmark difference smaller than eight points, you don’t have a finding — you have harness noise. Run it yourself before you decide it matters.
Is Kimi K3 Actually Cheap?
Open weights invite an assumption of cheapness the hosted API pricing doesn’t actually support. K3 lists at $3.00 per million input tokens on a cache miss, $0.30 on a cache hit, and $15.00 per million output tokens — roughly triple what K2.6 charged. Blended across Artificial Analysis’s Intelligence Index run, that worked out to $0.94 per completed task, which lands below GPT-5.6 Sol ($1.04) and about half of Claude Opus 4.8 ($1.80), but roughly 3x GLM-5.2’s $0.32 and more than 20x DeepSeek V4 Pro’s $0.04. K3 is mid-pack on cost, not the budget pick its open-weights label suggests.
The lever that actually moves your bill is cache-hit rate, not model choice. A cache hit costs a tenth of a miss on input tokens — which means stable prompt structure, deterministic tool schemas, and not switching models mid-session matter more to your invoice than which frontier model you picked. I go deeper into the cache-discipline mechanics, plan tiers, and setup steps specifically for running K3 inside Kimi Code in my full Kimi Code guide. If you’re evaluating multiple coding assistants against your monthly budget, our breakdown of GitHub Copilot’s token-credit pricing shift is a useful comparison point for how usage-based AI billing behaves across vendors.
Where K3’s Hallucination Risk Actually Bites
K3’s strongest territory. Verify against your own repo, and pass full thinking history back through the harness as Moonshot instructs.
Leading open-weights option if on-prem is a hard requirement, subject to your own infrastructure cost.
A 51% hallucination rate is the wrong profile for anything shipped without a human read. Require citations and treat unsourced assertions as a failure.
Measured output speed sits well below Claude Fable 5 in the same comparison. Check throughput before betting a streaming chat UI on it.
My Eval-Before-Adopt Checklist
None of this requires a research budget. These are the four checks I actually run before letting a new model touch anything that reaches a reader or a client.
- Re-run your existing harness against the new weights. If you already qualified a prior model on your own workload, that’s your best comparator — not the public leaderboard delta.
- Score abstention, not just accuracy. Seed a handful of questions your corpus genuinely can’t answer, and track whether the model declines or guesses. This is the exact metric K3’s numbers say most teams skip.
- Test ambiguous prompts on purpose. Moonshot itself flags that K3 may act unilaterally under unclear intent. Benchmarks use clean inputs; your team doesn’t.
- Price the completed task, not the token rate. Cache-hit rate and reasoning length dominate your real bill far more than the sticker price per million tokens does.
FAQ
What is Kimi K3’s hallucination rate?
Artificial Analysis reports Kimi K3 at 51% on its AA-Omniscience hallucination benchmark, up from 39% on Kimi K2.6. The metric measures how often the model guesses confidently rather than abstaining when it doesn’t know — not the raw share of all answers that were wrong.
If accuracy went up, why does the hallucination increase matter?
Because they affect different users. An accuracy gain helps someone who asks a question the model happens to know. A hallucination-rate increase hits someone who asks a question it doesn’t know — and now gets a confident wrong answer instead of an honest hedge. In production, those aren’t always the same person.
Does Moonshot publish its own hallucination benchmark for K3?
No. K3’s GitHub repository lists 40-plus benchmarks across reasoning, coding, agentic, and vision categories, but none of them measure factual reliability or hallucination. That number came entirely from an independent evaluator.
Is Kimi K3 cheap because it’s open-weights?
Not particularly. Hosted API pricing sits at $3/$15 per million tokens, with a blended cost of $0.94 per completed task on Artificial Analysis’s Index run — mid-pack against comparable frontier models, and roughly 3x pricier than GLM-5.2 on a per-task basis.
How should I evaluate Kimi K3 before using it on real work?
Re-run whatever eval harness you already trust against the new weights, add an abstention metric alongside accuracy, test deliberately ambiguous prompts, and price your own completed tasks rather than the per-token rate. Treat any single coding benchmark under an 8-point gap as harness noise until you reproduce it yourself.
The Final Word
Kimi K3 is a genuinely strong release — the best open-weights model on Artificial Analysis’s board, with agentic gains that justify serious evaluation. It’s also a model that answers more questions it doesn’t actually know the answer to than its predecessor did, on a benchmark its own vendor doesn’t run, at a rate that never made it into a launch chart. Both facts come from the same organization’s data, published the same week. Only one of them travelled, because composite indices are legible and sub-metrics aren’t.
My take after a week with it: don’t avoid K3, but stop treating any single published score as your adoption decision. Run your own harness. Score abstention next to accuracy. Assume a sub-8-point coding gap is noise until you reproduce it. That discipline costs an afternoon and is the only thing standing between a leaderboard rank and a bad surprise in production. For more on how I evaluate coding-focused AI tools specifically, my vibe coding tools roundup and Claude AI trustworthiness review apply the same testing standard.
Own a SaaS or AI Tool?
Get a human-written, hands-on review that gets you found in AI and Google search — the same testing standard used in this post.
Get Your Review · $50Who Am I
I’m Oyekale Olawale, and I run Websites2Know, a site built on one rule: I don’t write about a tool or a benchmark until I’ve actually run it myself. That means paying for the tier that unlocks the feature I’m reviewing, running models against real repos and real prompts, and pulling primary-source benchmark data instead of relaying a launch article. If a number looks too clean in a press release, I go find where it actually came from — which is exactly how this piece ended up including a discrepancy the original coverage didn’t catch.
Related Reading on Websites2Know
- Kimi K3 in Kimi Code: Setup, Plans, and Cache Rules
- Is Claude AI Safe and Trustworthy?
- Is Claude AI Good for Coding? Full Developer Review
- Is ChatGPT Good for Coding?
- GitHub Copilot Token Credits Policy Changes Explained
- 10 ChatGPT Alternatives That Do Some Things Better
- Best Vibe Coding Tools I Tested and Reviewed
- ZenMux AI Review
- Qwen AI Review
- 10 Top Verdent Alternatives for AI-Powered Coding