Muse Spark 1.1 beats GPT-5.6 Sol on radiology benchmark RadLE 2.0
AIOn Radiology's Last Exam, Meta's Muse Spark 1.1 outperforms OpenAI's GPT-5.6 Sol and Gemini 3.1, but still trails Fable and human experts. The result comes from a post by Jason Wei, with the benchmark context coming from a separate post about RadLE 2.0, an uncertainty-aware radiology diagnosis benchmark.



