Qwen 3.8 27B tested on long-number addition, reasoning vs non-reasoning modes
Original titleI re-ran an experiment @colin_fraser ran against GPT-4o a while back to see how good it was at adding long numbers, only this time I trie...
AISummary
Simon Willison re-ran Colin Fraser's earlier GPT-4o long-number addition experiment, this time testing Qwen 3.8 27B running locally. He compared the model's performance in both reasoning and non-reasoning modes.
Source: Simon Willison · x.comPublished · added here