Voxtral Realtime paper is out !
AIThe model is released under the Apache 2 license, and achieves state-of-the-art transcription performance at sub-500ms latency.
Updated
Updated
AIThe model is released under the Apache 2 license, and achieves state-of-the-art transcription performance at sub-500ms latency.
AIMiniMax describes Forge, its internal reinforcement learning framework for training real-world agents, which was used during the development of MiniMax M2.5. The post explains a Windowed FIFO scheduler, prefix tree merging that the post says yields a 40x training speedup, and CISPO-based training across more than one hundred thousand agent scaffolds and environments.
Why it matters: The post details how the Forge framework balances throughput, stability, and agent flexibility, with concrete scheduling and prefix-merging methods for training agent RL at scale.
AIYi Tay introduces Aletheia, a math research agent powered by an advanced version of Gemini Deep Think. The post says it produced two publishable papers, one fully automatic and one human-AI collaboration, and solved multiple open Erdős problems. The attached image shows a Google DeepMind paper titled "Towards Autonomous Mathematics Research" with a generator, verifier, and reviser loop.
AI…and discover obscure mathematical links across TCS, physics, and more. Link to paper:
AIQuoc Le announced a paper describing Aletheia, an agent built on Gemini Deep Think that goes beyond Olympiad problems to PhD-level mathematics. The post says Aletheia iteratively generates and verifies proofs, collaborates on human-AI research, autonomously generates a paper on eigenweights, and solves open Erdős conjectures.
AIAnthropic reports that resource configuration alone can move Terminal-Bench 2.0 scores by up to 6 percentage points, with infra error rates falling from 5.8% under strict enforcement to 0.5% when uncapped. Above about 3x the per-task specs, extra headroom starts letting agents solve tasks they previously could not, so limits can change what the eval measures.
Why it matters: The source shows how container resource limits shift agentic coding scores, which helps readers interpret small leaderboard gaps and set up evals more consistently.
AIQuoc Le announced a case study using Gemini to systematically evaluate 700 conjectures labeled open in the Erdős Problems database. The team addressed 13 problems, finding 5 novel autonomous solutions and identifying 8 existing solutions missed by previous literature.
AIThe Beijing Academy of Artificial Intelligence (BAAI) published its Emu3 multimodal large model research in Nature, which the post describes as the first large-model achievement led by a Chinese research institution in that journal. Emu3 learns from text, image, and video at scale using next-token prediction alone, reaching generation and perception performance comparable to task-specific methods. The authors frame this as a step toward scalable, unified multimodal intelligence systems.
AIAi2's Open Coding Agents family, with SERA as its first release, was built by Tim Dettmers and collaborators on 32 GPUs. The method generates synthetic bug trajectories with soft verification, comparing patches by line overlap instead of running tests. The post reports that a 32B model fine-tuned on about 7,000 trajectories for one private repository matched its GLM 4.5-Air teacher, and that the baseline costs $500 to run.
AIBerkeley AI Research proposes an information-based framework that evaluates and optimizes imaging systems using mutual information estimated directly from noisy measurements. The team reports that the metric predicts decoder performance across color photography, radio astronomy, lensless imaging, and microscopy, and that optimized designs match end-to-end methods while requiring less memory and compute.
AI1M× faster screening — 10 trillion protein–molecule pairs/day. ⚡ Full genome-scale drug-target mapping — screened 10k proteins × 500M molecules, identifying 2M+ candidates. DrugCLIP bridges AlphaFold’s structures to real drug candidates — launching the post‑AlphaFold era of scalable discovery. 📄Paper: 🌐Platform: #AI #Science
AIApple has released SHARP, a model that generates a 3D Gaussian representation of a scene from a single photograph in less than a second on a standard GPU. The output renders in real time as high-resolution photorealistic views of nearby camera positions, with metric absolute scale, and the paper reports reductions of 25–34% in LPIPS and 21–43% in DISTS versus the best prior model.
AIThinking Machines published a post on on-policy distillation, a training approach combining the error-correcting relevance of RL with the reward density of SFT. The quoted post reports that in math reasoning and an internal chat assistant, on-policy distillation can outperform other approaches at a fraction of the cost.
AIThinking Machines Lab describes on-policy distillation, which samples rollouts from a student model and has a teacher grade each token with reverse KL. The authors report that this matches Qwen3-style reasoning results at a fraction of RL's cost, with AIME'24 reaching 70% in about 150 steps from a 400k SFT checkpoint. The method also helps recover instruction-following behavior lost during fine-tuning on internal documents.
Why it matters: The post explains why on-policy distillation gives dense per-token feedback, letting a small model match RL results at much lower compute cost.
AIStanford and Cognition AI researchers introduced Kevin-32B, a 32B-parameter model trained with multi-turn reinforcement learning to write CUDA kernels. On KernelBench, it solves 89% of tasks at best@16 and achieves 65% average correctness over eight refinement steps, versus 53% for o4-mini and 51% for o3. Its best@16 speedup is 1.41x, and multi-turn training outperforms single-turn training as refinement steps increase.
AILiquid AI reports STAR, an evolutionary algorithm that synthesizes tailored neural network architectures from a numerical genome representation. The authors say it produced hundreds of designs that outperform Transformer and hybrid architectures in quality, with smaller caches and parameter counts, and can optimize for latency on target hardware. The full method is described in the arXiv technical report 2411.17800.
Why it matters: The post explains how evolutionary search over a new architecture design space produced designs beating Transformers and hybrids, giving a concrete method for quality versus latency and memory trade-offs.
AICognition reports that its agent Devin resolved 79 of 570 sampled SWE-bench issues, a 13.86% success rate, without being given the files to edit. The report says this exceeds the best previous unassisted baseline of 1.96% and the best assisted result of 4.80%. It also describes the adapted evaluation setup, a 45-minute runtime limit, and cases where Devin failed on multi-file edits.
Why it matters: The report explains how SWE-bench was adapted for end-to-end agent evaluation, with failure cases that clarify where the 13.86% result comes from and its limits.
AIThe team proposed a simple yet efficient machine translation solution named VOLT, which is now open source to developers. Check it out: