Inkling Small
Thinking Machines·Released Jul 15, 2026Open weights
Updated Oct 8, 9:18 PM ET
The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.
By domain
Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.
Every result behind the score
The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".
Coding
7 results · 34.3- ProgramBenchVals AINear the floor0.5%Reading at most 65.7
- SciCodeArtificial Analysis49.7%Reading 48.8
- Terminal-Bench 4.0 (AA run)Artificial AnalysisNear the floor1.0%Reading at most 34.5
- CyberBench PatchVals AI73.2%Reading 33.9
- Terminal-Bench 4.0Vals AINear the floor0.0%Reading at most 33.0
- Code MigrationVals AI13.7%Reading 32.9
- Vibe Code BenchVals AI19.1%Reading 29.1
Research and reasoning
6 results · 28.8- Humanity's Last Exam (AA run)Artificial Analysis33.3%Reading 47.3
- FrontierMath Tier 4Epoch AI · xhigh effort17.1%Reading 40.4
- FrontierMath Tiers 1–3Epoch AI · xhigh effort46.3%Reading 39.3
- Chess PuzzlesEpoch AI · xhigh effort18.0%Reading 26.6
- ProofBenchVals AI6.0%Reading 25.1
- Mystery Game PuzzlesEpoch AI · xhigh effort6.0%Reading 4.0
Professional work
7 results · 31.9- Legal Research BenchVals AI25.5%Reading 43.7
- Tax Agent BenchVals AI15.4%Reading 37.3
- Finance AgentVals AI41.3%Reading 36.5
- τ-Bench Banking (AA run)Artificial Analysis18.8%Reading 30.1
- MedCodeVals AI37.9%Reading 26.6
- EMBVals AI31.8%Reading 24.8
- Harvey Legal Agent BenchmarkVals AINear the floor1.7%Reading at most 24.8
Knowledge and accuracy
2 results · 24.6- AA-LCRArtificial Analysis75.7%Reading 42.3
- SimpleQA VerifiedEpoch AI · xhigh effort19.1%Reading 3.8
Human preference
1 result · 28.4- Arena TextLMArena1,405.2Reading 27.2
Shown, not scored
Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.
- IOIVals AI · CodingReference only9.3%
- LiveCodeBenchVals AI · CodingReference only85.9%
- SkillsBenchVals AI · CodingWatching33.6%
- SWE-bench VerifiedVals AI · CodingReference only82.2%
- OTIS Mock AIMEEpoch AI · Research and reasoningReference only90.0%
- CorpFinVals AI · Professional workReference only69.6%
- LegalBenchVals AI · Professional workReference only83.0%
- MedScribeVals AI · Professional workWatching84.1%
- TaxEvalVals AI · Professional workReference only75.5%
- GPQA DiamondEpoch AI · Knowledge and accuracyReference only88.5%
- GPQA DiamondVals AI · Knowledge and accuracyReference only83.6%
- MMLU-ProVals AI · Knowledge and accuracyReference only85.6%
- AA Intelligence IndexArtificial Analysis · Composite indicesReference only25.7
- Vals IndexVals AI · Composite indicesReference only25.5%
- Arena VisionLMArena · VisionReference only1,206.3
- Arena WebDevLMArena · Writing and designReference only1,408.6
No result yet
Scored evaluations this model has not taken. A missing result neither adds nor subtracts.
- LiveBench Coding · LiveBench
- Vibe Code Bench 1–100 · Vals AI
- APEX-SWE · Mercor
- LiveBench Reasoning · LiveBench
- MysteryMechanism · Vals AI
- Terminal-Bench Science · Vals AI
- APEX-Agents · Mercor
- LiveBench Data Analysis · LiveBench
- τ²-Bench Telecom (AA run) · Artificial Analysis
- LiveBench Language · LiveBench
- LiveBench Instruction Following · LiveBench
- BullshitBench · BullshitBench
- IFBench (AA run) · Artificial Analysis