Skip to content
Read the original: Simon Willison· Published 22/100AI score22/100

Qwen 3.8 27B tested on long-number addition, reasoning vs non-reasoning modes

Original titleI re-ran an experiment @colin_fraser ran against GPT-4o a while back to see how good it was at adding long numbers, only this time I trie...

AISummary

Simon Willison re-ran Colin Fraser's earlier GPT-4o long-number addition experiment, this time testing Qwen 3.8 27B running locally. He compared the model's performance in both reasoning and non-reasoning modes.

Read the original x.com

Source: Simon Willison · x.comPublished · added here