Skip to content

Inkling Small

Thinking Machines·Released Jul 15, 2026Open weights

Updated Oct 8, 9:18 PM ET

Overall score
31.0
Rank #59 overall
Evidence
23 results
64% of scored kinds covered
Context window
524K tokens
From OpenRouter
API price, USD per 1M tokens
$0.45 / $1.20
Input / output · cached input $0.10

The score is on the board's scale (reference average 50, about 15 points per standard deviation), not a percentage.

By domain

Each domain is scored on the same scale as the overall score. Coverage is the share of the domain's kinds of evaluation with a result; it says how much evidence there is, not how good the model is.

Coding#55
34.3
7 results · 67% coverage
Research and reasoning#59
28.8
6 results · 63% coverage
Professional work#56
31.9
7 results · 78% coverage
Knowledge and accuracy#63
24.6
2 results · 33% coverage
Human preference
28.4
1 result · 100% coverage

Every result behind the score

The result as the evaluator publishes it, and the reading it gives on the board's scale. A hard evaluation can give a high reading for a modest result; results near 95% or 5% only say "at least" or "at most".

Coding

7 results · 34.3

Research and reasoning

6 results · 28.8

Professional work

7 results · 31.9

Knowledge and accuracy

2 results · 24.6

Human preference

1 result · 28.4

Shown, not scored

Composite indices, saturated or older evaluations, evaluations still being checked, and vision, writing and multilingual results. They do not move the score.

  • IOIVals AI · CodingReference only9.3%
  • LiveCodeBenchVals AI · CodingReference only85.9%
  • SkillsBenchVals AI · CodingWatching33.6%
  • SWE-bench VerifiedVals AI · CodingReference only82.2%
  • OTIS Mock AIMEEpoch AI · Research and reasoningReference only90.0%
  • CorpFinVals AI · Professional workReference only69.6%
  • LegalBenchVals AI · Professional workReference only83.0%
  • MedScribeVals AI · Professional workWatching84.1%
  • TaxEvalVals AI · Professional workReference only75.5%
  • GPQA DiamondEpoch AI · Knowledge and accuracyReference only88.5%
  • GPQA DiamondVals AI · Knowledge and accuracyReference only83.6%
  • MMLU-ProVals AI · Knowledge and accuracyReference only85.6%
  • AA Intelligence IndexArtificial Analysis · Composite indicesReference only25.7
  • Vals IndexVals AI · Composite indicesReference only25.5%
  • Arena VisionLMArena · VisionReference only1,206.3
  • Arena WebDevLMArena · Writing and designReference only1,408.6

No result yet

Scored evaluations this model has not taken. A missing result neither adds nor subtracts.

  • LiveBench Coding · LiveBench
  • Vibe Code Bench 1–100 · Vals AI
  • APEX-SWE · Mercor
  • LiveBench Reasoning · LiveBench
  • MysteryMechanism · Vals AI
  • Terminal-Bench Science · Vals AI
  • APEX-Agents · Mercor
  • LiveBench Data Analysis · LiveBench
  • τ²-Bench Telecom (AA run) · Artificial Analysis
  • LiveBench Language · LiveBench
  • LiveBench Instruction Following · LiveBench
  • BullshitBench · BullshitBench
  • IFBench (AA run) · Artificial Analysis
How readings and scores are computed