AI Daily · Oct 11, 2026 · Sunday
BAD INFLUENCE
Xiaomi's MiMo-V2.6 report explains scaling RL along batch, environments, and grading
Xiaomi's MiMo-V2.6 technical report argues that scaling reinforcement learning, not more pretraining data, is the main lever for frontier capability, along batch size, environment diversity, and grader strength. The article summarizes the report's methods, including groupwise agentic grading, a frozen-router fix for expert load collapse, and reward hacking defenses. It reports RL post-training costs of $2.6 million for MiMo-V2.6-Pro and $0.9 million for MiMo-V2.6-Flash.
Model releases & updates
VeriLoop E2, a 27B open-weight reasoning model, is now available
ModelScope announces VeriLoop E2, a 27B open-weight model built on Qwen3.8-27B with a 262K native context window. The company reports 76.2% on SWE-bench Pro, 88.8% on Terminal-Bench 2.1, 98.3% on AIME 2026, and 93.9% on GPQA Diamond. Model weights and public inference utilities are released under Apache 2.0, while the production VeriLoop Harness is not included.

Industry news
TypeSafe AI's Jev model lifts valuation to $7.5 billion in 24 days
TypeSafe AI, parent of the decision-only model Jev, raised $870 million in a Series A at a $7.5 billion post-money valuation, according to the company. The model handles judgment tasks rather than chat or code, and reportedly processes over one trillion tokens a day, mostly from automated backend calls. The company also says about a third of Fortune 500 companies use or have access to it.
In brief
- FT says OpenAI and Anthropic revenue figures now move public marketsRohan PaulFollow-up
Get AI Daily by email
One email every morning with the day's top AI news. Unsubscribe anytime.
We use your address only for this list. Privacy