FormalyFormaly
FeaturesWhyComparePricingBlog
All articles
Comparison/ Jul 17, 2026

Claude Fable 5 vs. Kimi K3: Real-World Agentic Benchmarks Compared

Kimi K3 is the first open-weight model to beat Claude Fable 5 outright on several real coding and agentic benchmarks, not just compete on price. Here is what the sourced numbers actually say.

8 min read

Arindam Majumder

Arindam Majumder

Formaly

Share this article

HN

Moonshot AI released Kimi K3 on July 16, 2026, a 2.8 trillion parameter mixture-of-experts model the company is calling the largest open-weight model ever shipped. That claim is normal launch-day noise. What isn't normal is the benchmark table underneath it: K3 beats Claude Fable 5 outright on several real coding and agentic tests, not by a rounding error, and not just because it's cheap.

That's a different story from the last comparison I wrote. MiniMax M3 competed with Fable 5 on price and lost on every capability benchmark. Kimi K3 competes on price and wins some of the capability benchmarks. That combination hasn't really shown up from an open model before.

The short version: Fable 5 still leads on raw intelligence and on a handful of coding benchmarks. K3 leads on terminal work, long-horizon agent evaluations, and browsing, and it does it at roughly a third of the price. The gap between the two is now closer than it is wide.

A note on methodology

Same disclaimer as last time. This is not a live task-execution study where both models get wired into identical tools and run through the same tasks side by side. Neither model's API was available to run that kind of test for this piece.

Every number below comes from Moonshot's own release benchmarks, Artificial Analysis, and independent coverage, cross-checked across sources and cited at the bottom. Where two trackers report different numbers for the same benchmark, and that happens more than once here, both are shown instead of picked at random.

The two models, side by side

Claude Fable 5Kimi K3
MakerAnthropicMoonshot AI
ReleasedJun 9, 2026Jul 16, 2026
AccessAPI, Bedrock, Vertex, FoundryAPI live, weights promised for Jul 27
WeightsClosedOpen (pending release)
ParametersNot disclosed2.8T total, 16 experts active (MoE)
Context window1M tokens1,048,576 tokens
Cached input pricenot published separately$0.30 / M tokens
Input price$10.00 / M tokens$3.00 / M tokens (uncached)
Output price$50.00 / M tokens$15.00 / M tokens
AA Intelligence Index60 (#1 of 189)57 (#4 of 189)

That last row is the headline difference from the MiniMax comparison. Against M3, Fable 5's Intelligence Index lead was 60 to 44, a wide gap. Against K3, it's 60 to 57. That's a genuinely close race between a closed frontier model and an open one that isn't fully released yet.

Gap 1: coding benchmarks, and a big asterisk

Moonshot published a three-way comparison against Fable 5 and GPT-5.6 Sol at launch. The catch, and it's a real one, is that each vendor ran its model inside its own harness (Kimi Code, Claude Code, Codex), so these aren't apples to apples in the way a shared test suite would be.

BenchmarkKimi K3Claude Fable 5
DeepSWE67.570.0
Program Bench77.876.8
Terminal-Bench 2.188.384.6*
FrontierSWE81.286.6
SWE Marathon42.035.0

* Fable 5's Terminal-Bench 2.1 score moves depending on who's reporting it. Anthropic and Artificial Analysis have put it as high as 88.0%, other independent runs report 83.4%, and Moonshot's own comparison table (shown above) has it at 84.6 in the specific configuration it tested against K3. Treat the K3-vs-Fable-5 gap on this one as roughly a wash rather than a clean win either way.

Fable 5 still wins DeepSWE and FrontierSWE by a real margin. K3 wins Terminal-Bench 2.1, Program Bench, and SWE Marathon, also by a real margin. This isn't one model sweeping the other, it's a split decision depending on what kind of coding task you're running.

Gap 2: agentic and knowledge-work benchmarks

This is where K3's story gets more interesting than "cheap and decent." On Artificial Analysis's own long-horizon agent evaluations, reported at launch by VentureBeat, K3 landed third overall behind Fable 5 Max and GPT-5.6 Sol Max, but ahead of Claude Opus 4.8:

BenchmarkKimi K3Fable 5 MaxGPT-5.6 Sol MaxClaude Opus 4.8
GDPval-AA v2 (real-world tasks, 44 occupations)1,6871,8151,747.81,600
AA-Briefcase (long-horizon knowledge work)1,5271,5871,495not reported

Moonshot's own comparison table, published separately, shows slightly different numbers for the same two benchmarks (K3 at 1,668 and 1,548, Fable 5 at 1,760 and 1,583). The two trackers don't agree down to the point, but they agree on the shape: K3 sits close behind Fable 5 on these, not far behind like M3 did.

Where K3 actually pulls ahead is task automation. Moonshot reports K3 ranking first in four of eight automation benchmarks, including Automation Bench and BrowseComp:

BenchmarkKimi K3Fable 5GPT-5.6 Sol
Automation Bench30.829.129.7
JobBench52.957.446.5
BrowseComp91.288.090.4

K3 wins Automation Bench and BrowseComp. Fable 5 wins JobBench by a wider margin than either of K3's wins. Again: a split, not a sweep.

Gap 3: Frontend Code Arena, the one clean win

The most repeated headline out of K3's launch is a genuinely unambiguous result. On Arena.AI's Frontend Code Arena, a benchmark that scores models on the frontend code they generate, K3 took the number one spot with a score of 1,679, ahead of both Fable 5 and GPT-5.6 Sol by a clear margin. This is the one benchmark in this whole comparison where there's no asterisk, no harness caveat, and no disagreement between trackers. An open-weight Chinese model is currently the best-scoring model on a major public leaderboard for a task Anthropic has spent a lot of its positioning on.

Gap 4: speed tells the same story as last time

The throughput numbers are close. K3 runs at about 62 tokens per second, Fable 5 at 66.1. Neither is fast for its price tier.

Time-to-first-token is where the real difference shows up again, and it's the same shape as the MiniMax comparison:

  • Kimi K3: 1.99 seconds to first token
  • Claude Fable 5: 149.8 seconds to first token

K3's thinking mode is real but far shorter. Fable 5's extended-thinking budget is enormous by comparison. If your agent is planning a long multi-tool session, that upfront cost on Fable 5's side can be worth it. If a person is waiting on a response, it usually isn't.

Gap 5: cost, but the gap is much smaller than MiniMax's

This is the part that separates K3 from the last comparison. MiniMax M3 was roughly 33 to 42 times cheaper than Fable 5. K3 isn't playing that game.

Claude Fable 5Kimi K3Multiple
Input tokens$10.00 / M$3.00 / M (uncached)~3.3x cheaper
Output tokens$50.00 / M$15.00 / M~3.3x cheaper
Cached inputnot published$0.30 / Mn/a

Roughly a third of the price, not a thirtieth. That changes the pitch. K3 isn't "good enough at a fraction of the cost," it's closer to "competitive, sometimes better, at a real but modest discount." That's a harder sell for switching an existing pipeline, and an easier one for a new build that doesn't have Fable 5 baked in yet.

The part that isn't in any benchmark: the weights aren't out yet

Kimi K3's API went live on July 16. The actual model weights, the thing that makes an "open-weight" model meaningfully different from a closed one, are scheduled for release on July 27, 2026, based on technical documentation reviewed by outside researchers. As of this writing, that hasn't happened yet.

That matters because "open-weight" is doing a lot of work in K3's pitch. Until the weights actually ship, self-hosting it, fine-tuning it, or treating it as independent of Moonshot's own API availability is a promise, not a fact. MiniMax M3 could make that claim on day one. K3 can't, yet.

What this actually tells us

Fable 5 is still the more intelligent model by Artificial Analysis's composite measure, 60 versus 57, and it wins outright on DeepSWE, FrontierSWE, and JobBench. That's real and worth taking seriously if peak reasoning is what you need.

But this is the first open-weight model that beats Fable 5 outright on more than one real benchmark: Terminal-Bench 2.1, Program Bench, SWE Marathon, Automation Bench, BrowseComp, and a clean win on Frontend Code Arena. At roughly a third of the price. That's a meaningfully different competitive position than MiniMax M3 held a month earlier.

Use caseBetter fit
Peak reasoning and DeepSWE / FrontierSWE-style codingFable 5
Long multi-tool agent sessions where upfront planning pays offFable 5
Vendor-managed compliance and safety guardrailsFable 5
Terminal-driven agent work and task automationKimi K3
Frontend code generationKimi K3
Browsing and retrieval-heavy agentsKimi K3
Interactive latency (time-to-first-token)Kimi K3
Needs weights in hand today, not in ten daysNeither yet, check back after Jul 27

The honest read: this isn't a value pick beating a premium one anymore. It's two frontier-tier models trading wins, with one of them costing a third as much and about to open its weights.

On this page

A note on methodologyThe two models, side by sideGap 1: coding benchmarks, and a big asteriskGap 2: agentic and knowledge-work benchmarksGap 3: Frontend Code Arena, the one clean winGap 4: speed tells the same story as last timeGap 5: cost, but the gap is much smaller than MiniMax'sThe part that isn't in any benchmark: the weights aren't out yetWhat this actually tells us
› Keep readingAll articles →
Essay

Completion Rate Is a Vanity Metric

7 min read
Essay

The Form Is Dying. The Interview Is Replacing It.

7 min read
Engineering

How We Chose Our AI Provider: Why Formaly Runs on Nebius Token Factory

9 min read
FormalyFormaly

Talk to build. Talk to answer.

© 2026 Formaly

AboutDocsBlogComparePrivacyTermsContact

Made by Arindam

Formaly