Inkling
Thinking Machines·Released Jul 15, 2026Open weights
Updated Oct 8, 11:55 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
10 results · 36.0- ProgramBenchVals AINear the floor0.0%Reading at most 65.7
- CyberBench PatchVals AI78.6%Reading 46.8
- LiveBench CodingLiveBench · xhigh effort60.2%Reading 41.9
- SciCodeArtificial Analysis · xhigh effort47.0%Reading 41.5
- Terminal-Bench 4.0 (AA run)Artificial Analysis · xhigh effortNear the floor1.0%Reading at most 34.5
- Terminal-Bench 4.0Vals AINear the floor0.5%Reading at most 33.0
- APEX-SWEMercor · high effort26.1%Reading 29.8
- Code MigrationVals AI11.8%Reading 29.6
- Vibe Code BenchVals AI19.2%Reading 29.2
- Vibe Code Bench 1–100Vals AI7.3%Reading 28.7
Research and reasoning
6 results · 34.6- Humanity's Last Exam (AA run)Artificial Analysis · xhigh effort31.9%Reading 45.2
- LiveBench ReasoningLiveBench · xhigh effort83.4%Reading 44.7
- Chess PuzzlesEpoch AI · xhigh effort21.0%Reading 32.3
- FrontierMath Tiers 1–3Epoch AI · xhigh effort33.3%Reading 29.5
- ProofBenchVals AINear the floor0.0%Reading at most 23.0
- FrontierMath Tier 4Epoch AI · xhigh effortNear the floor4.9%Reading at most 22.2
Professional work
9 results · 40.3- LiveBench Data AnalysisLiveBench · xhigh effort72.8%Reading 50.9
- τ-Bench Banking (AA run)Artificial Analysis · xhigh effort29.1%Reading 47.9
- Legal Research BenchVals AI28.4%Reading 47.7
- Finance AgentVals AI46.6%Reading 47.1
- MedCodeVals AI41.2%Reading 38.1
- Tax Agent BenchVals AI15.3%Reading 37.1
- APEX-AgentsMercor · high effort33.8%Reading 37.0
- EMBVals AI38.4%Reading 32.7
- Harvey Legal Agent BenchmarkVals AINear the floor2.1%Reading at most 24.8
Knowledge and accuracy
4 results · 46.4- LiveBench Instruction FollowingLiveBench · xhigh effort70.1%Reading 62.9
- AA-LCRArtificial Analysis · xhigh effort77.3%Reading 44.8
- LiveBench LanguageLiveBench · xhigh effort73.5%Reading 41.1
- SimpleQA VerifiedEpoch AI · xhigh effort40.3%Reading 40.8
Human preference
1 result · 38.2- Arena TextLMArena1,441.4Reading 37.9
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- IOIVals AI · CodingReference only14.9%
- LiveCodeBenchVals AI · CodingReference only85.5%
- SkillsBenchVals AI · CodingWatching26.1%
- SRE BenchVals AI · CodingWatching0.0%
- SWE-bench VerifiedVals AI · CodingReference only77.6%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only88.9%
- APEX-AccountingMercor · Professional workWatching0.8%
- CorpFinVals AI · Professional workReference only68.6%
- LegalBenchVals AI · Professional workReference only82.9%
- MedScribeVals AI · Professional workWatching85.4%
- TaxEvalVals AI · Professional workReference only75.3%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only88.3%
- GPQA DiamondVals AI · Knowledge and accuracyReference only87.1%
- MMLU-ProVals AI · Knowledge and accuracyReference only86.3%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only25.0
- LiveBench averageLiveBench · Composite indicesReference only71.9%
- Vals IndexVals AI · Composite indicesReference only28.7%
- MMMU ProVals AI · VisionReference only78.5%
- Arena WebDevLMArena · Writing and designReference only1,412.2
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- Mystery Game Puzzles · Epoch AI
- MysteryMechanism · Vals AI
- Terminal-Bench Science · Vals AI
- τ²-Bench Telecom (AA run) · Artificial Analysis
- BullshitBench · BullshitBench
- IFBench (AA run) · Artificial Analysis