Skip to content

All AI news

Oct 8

TodayOct 8Thu67 items
  1. GoodfireAI score28

    We built cybersecurity monitors for Kimi K3 and GLM 5.3 — more accurate, 50x faster, & 50x cheaper than an optimized LLM judge. External red-teaming from @farairesearch found that our monitor greatly reduces universal jailbreaks. 🧵

    We built cybersecurity monitors for Kimi K3 and GLM 5.3 — more accurate, 50x faster, & 50x cheaper than an optimized LLM judge. External red-teaming from @farairesearch found that our monitor greatly reduces universal jailbreaks. 🧵

  2. Google ResearchAI score14

    Missed yesterday's demo on adaptive agent environments? Stop by the #COLM2026 Google booth #107 today at 11:00 AM PT to catch Zifeng Wang presenting EnvHarness — a flexible framework enabling co-evolution between LLM agents and their training environments.

    Missed yesterday's demo on adaptive agent environments? Stop by the #COLM2026 Google booth #107 today at 11:00 AM PT to catch Zifeng Wang presenting EnvHarness — a flexible framework enabling co-evolution between LLM agents and their training environments.

  3. Goodfire ResearchAI score57

    Goodfire deploys probe-based cyber monitors on Kimi K3 with a judge cascade

    Goodfire Research describes probe-based cyber monitors for Kimi K3 and GLM 5.3 deployed on a production inference stack. The probe filters suspicious exchanges before an LLM judge reviews them, reaching about 93% recall at a 5.5% benign-session interruption rate at roughly 50x lower judge cost. In FAR.AI's red-teaming, the monitor reduced universal jailbreaks to zero across 140 tested strategies.

  4. GoogleAI score40

    In our prospective clinical study, our models were able to pinpoint the gestational age within 4 days of accuracy which is a tight enough zone that can have a really meaningful clinical impact. If we're able to expand these tools to low resource settings, we can start to move the needle on maternal deaths and bridge that gap in care that we see all over the world.

    In our prospective clinical study, our models were able to pinpoint the gestational age within 4 days of accuracy which is a tight enough zone that can have a really meaningful clinical impact. If we're able to expand these tools to low resource settings, we can start to move the needle on maternal deaths and bridge that gap in care that we see all over the world.

  5. ZyphraAI score34

    Our approach is lossless. The same tokens visit the same experts. The architecture, routing decisions and training objective stay unchanged. The model computes the same function. By organizing where experts and tokens live we reduce the communication needed to do the same work.

    Our approach is lossless. The same tokens visit the same experts. The architecture, routing decisions and training objective stay unchanged. The model computes the same function. By organizing where experts and tokens live we reduce the communication needed to do the same work.

  6. ZyphraAI score38

    The gains depend on the training configuration, and are largest when each token uses more experts and when the experts span several nodes. Results in Megatron-LM on 8 to 64 GPUs: - Token exchange: 1.16x to 2.63x faster - Full training step: up to 1.41x faster

    The gains depend on the training configuration, and are largest when each token uses more experts and when the experts span several nodes. Results in Megatron-LM on 8 to 64 GPUs: - Token exchange: 1.16x to 2.63x faster - Full training step: up to 1.41x faster

  7. ZyphraAI score22

    Crucially, we find that these patterns emerge early enough to be useful during training and are cheap to measure. Using just 4,096 tokens, our placement and routing predictions come within roughly one percentage point of results using 524,000 tokens.

    Crucially, we find that these patterns emerge early enough to be useful during training and are cheap to measure. Using just 4,096 tokens, our placement and routing predictions come within roughly one percentage point of results using 524,000 tokens.

  8. ZyphraAI score18

    Routing is also predictable across layers: the experts a token used in one layer tell us which it will likely need next. Token shuffling moves each token to the GPU holding those experts, inside a transfer that already runs after attention, so it adds no network traffic.

    Routing is also predictable across layers: the experts a token used in one layer tell us which it will likely need next. Token shuffling moves each token to the GPU holding those experts, inside a transfer that already runs after attention, so it adds no network traffic.

  9. ZyphraAI score37

    So we put experts that are picked together on the same GPU, and send each token to a GPU once no matter how many of its experts live there. We call this correlated expert placement. On 8 GPUs it removes up to 58% of the token copies sent.

    So we put experts that are picked together on the same GPU, and send each token to a GPU once no matter how many of its experts live there. We call this correlated expert placement. On 8 GPUs it removes up to 58% of the token copies sent.

  10. ZyphraAI score32

    MoE models route each token sparsely to just a few experts. As these models grow, the experts are spread across GPUs and nodes, and tokens have to be sent to their experts and back. In our runs that exchange took 13-24% of step time on one node and 45-60% across four nodes.

    MoE models route each token sparsely to just a few experts. As these models grow, the experts are spread across GPUs and nodes, and tokens have to be sent to their experts and back. In our runs that exchange took 13-24% of step time on one node and 45-60% across four nodes.

  11. ZyphraAI score23

    Those routing decisions have patterns. For instance, in one layer, just 64 of 8,128 possible expert pairs accounted for 42% of tokens. If experts were picked independently, those pairs would carry 1.6%. We show how to exploit this structure to improve where experts are placed.

    Those routing decisions have patterns. For instance, in one layer, just 64 of 8,128 possible expert pairs accounted for 42% of tokens. If experts were picked independently, those pairs would carry 1.6%. We show how to exploit this structure to improve where experts are placed.

  12. ZyphraAI score22

    In Mixture of Expert (MoE) models the cost of moving tokens to their experts can dominate total runtime. At Zyphra research, we use patterns in how tokens are routed to experts to make that communication faster by up to 2.63x on @AMD MI300X GPUs, with the model unchanged.

    In Mixture of Expert (MoE) models the cost of moving tokens to their experts can dominate total runtime. At Zyphra research, we use patterns in how tokens are routed to experts to make that communication faster by up to 2.63x on @AMD MI300X GPUs, with the model unchanged.

  13. OpenBMBAI score36

    ReJev fine-tunes MiniCPM5-2B to lift decision accuracy to 80.50%

    ReJev, an independent community project, applied LoRA post-training to OpenBMB's MiniCPM5-2B for bounded agent decisions: state, question, and candidate options yield one choice. On its sealed 1,892-sample holdout, accuracy rose from 51.11% to 80.50% (+29.39 percentage points) with 0% invalid outputs, at about $5.31 in cumulative Modal billing including earlier experimental overhead. The authors describe this as an early, task-specific result, not parity with Jev.

  14. QbitAI (量子位)AI score44

    PaperBenchX Shows Top Model Reproduces Only 13.98% of 93 Scientific Papers End-to-End

    UniPat AI's PaperBenchX benchmark found the strongest model, GPT-6 Astra, fully reproduced only 13.98% of 93 real research-paper tasks across 12 scientific fields. Reproduction was judged by regenerating outputs in an isolated environment, with 3,168 expert-verified scoring items. UniPat has open-sourced 12 test tasks and kept 81 tasks closed to preserve long-term evaluation validity.

  15. MarkTechPostAI score45

    NVIDIA's PivotOPD Trains Multi-Turn AI Agents to Recover From Pivotal Mistakes

    NVIDIA, Princeton University, and the University of Maryland introduced PivotOPD, an on-policy distillation method that teaches multi-turn LLM agents to recover from their most damaging early mistake. Tested on Qwen3-1.7B and Qwen3-8B students, it posts the best average against 13 baselines on ALFWorld, WebShop, and Search-based QA. It recovers from 72.7% of replayed pivotal mistakes, versus 20.3% for standard OPD, with no added inference cost.

  16. PandailyAI score46

    Galbot and Tsinghua's LATENT Wins IROS Award for Humanoid Tennis Forehand

    A Galbot, Tsinghua University and collaborators paper won IROS 2026's Best Entertainment and Amusement Paper Award for LATENT, a humanoid tennis-return method trained on imperfect amateur motion-capture clips. In simulation, the full forehand policy succeeded on 96.52 percent of returns, versus 71.85 percent for PULSE. On a real Unitree G1, the paper reports 90.90 percent forehand success across 20 consecutive rallies, with motion capture still used rather than the robot's own cameras.

  17. PandailyAI score38

    Huawei Presents Experimental XMFS Shared-Memory Filesystem at LPC 2026

    Huawei engineers presented XMFS, an experimental Linux kernel prototype filesystem, at the Linux Plumbers Conference in Prague on October 5. It aims to let applications reach cross-node shared memory on CXL 3.0 or Huawei unified bus servers through standard POSIX file calls. The code exists only on openEuler, not in the mainline Linux kernel.

  18. Anthropic ResearchAI score62

    Anthropic researcher builds first complete UV sky map with Claude Science

    Johns Hopkins astrophysicist Brice Ménard, working as an Anthropic researcher, used Claude Science to produce the first complete map of the sky in ultraviolet light. Claude orchestrated agents to merge GALEX, Swift, and FIMS/SPEAR data, then predicted roughly a third of the sky that no UV telescope had observed, using relationships to visible, infrared, and radio data. Hidden test regions were reconstructed to within about 10% of real measurements, and each pixel is labeled measured or predicted with uncertainty estimates.

    AIWhy it matters: The post shows how an astrophysicist used Claude Science agents to merge UV surveys and predict missing sky regions, with a validation step that makes the method reusable.

  19. Artificial Analysis ArticlesAI score50

    Harvey LAB-AA v1.1 adds hallucination checks to legal AI benchmark

    Harvey LAB-AA v1.1 adds hallucination checks that audit every model deliverable against task source documents, with material hallucinations zeroing a task's score. GPT-6 Astra averaged 0.03 material hallucinations per task across 120 tasks, while Gemini 3.8 Flash averaged 13.96. Harvey uses GPT-6 Sol (high) as the hallucination checker, separate from its three-judge rubric panel.

Oct 7

Oct 7Wed
  1. Khazix (数字生命卡兹克)AI score88

    OpenAI Releases 722 Unpublished AI-Generated Math Manuscripts on GitHub

    OpenAI published 722 math manuscripts covering 372 result groups in a new GitHub repository, openai/math, all produced by an unreleased internal model. The author describes the results as including a near-Riemann hypothesis claim pushed to 0.875, and notes that 25 Fields Medal winners criticized the company's approach to AI math research.

    AIWhy it matters: The piece traces how AI math results moved from benchmarks to open problems, offering context on verification and the mathematicians' pushback.

  2. Elvis SaraviaAI score67

    Tool-using multimodal models refuse harmful requests less often, NVIDIA study finds

    A NVIDIA study accepted at NeurIPS 2026 reports that multimodal models refuse harmful requests less reliably when they call tools. Refusal failures rise by up to 68.7% relative and by 17.7% on average across the models tested, including Claude Opus 4.6 and 4.7 and Gemini Agentic Vision. The authors attribute this to tool outputs crowding out the original harmful intent and to attention shifting toward describing tool results. Re-inserting the original request and image before the final response restores part of the lost refusals.

  3. vLLMAI score46

    vLLM-Omni technical report unifies serving for omni-modality generation

    The vLLM team released a technical report on vLLM-Omni, a unified serving runtime for omni-modality generation spanning multi-stage autoregressive pipelines, iterative diffusion, and stateful sessions. Current LLM servers and diffusion stacks each cover only one of these patterns, pushing deployments to stitch disjoint runtimes together. vLLM-Omni offers a shared control plane in which an orchestrator advances requests across stages, specialized engines handle compute, and a connector carries payloads.

  4. Apple Machine Learning ResearchAI score42

    Apple's Normalizing Trajectory Models generate images in four steps with exact likelihood

    Apple researchers introduced Normalizing Trajectory Models (NTM), which model each reverse diffusion step as a conditional normalizing flow trained with exact likelihood. The model matches or outperforms strong image generation baselines on text-to-image benchmarks in just four sampling steps while retaining exact likelihood over the generative trajectory.

  5. Waymo BlogAI score42

    Sober Drivers Still Face Nearly 4x Nighttime Fatal Crash Risk, Waymo Study Finds

    Waymo research found that even fully sober human drivers face nighttime fatal crash risk 3.1 to 3.9 times higher than daytime risk, pointing to systemic hazards beyond impairment. The study used an exposure reconstruction model across the 50 most populous U.S. urban areas, showing removing alcohol-involved drivers lowers the average urban fatal crash rate by 23%, from 1.42 to 1.10 per 100 million miles.

  6. Epoch AIAI score67

    Epoch tests six AI models on real Epoch work and finds they cannot yet fully automate it

    Epoch gave six models 11 real work tasks from its own operations, including graphic design, data insights, and research design, and graded outputs against employee standards. Fable 5.1 and GPT-6 Astra led on average task performance, reliably handling well-defined work such as coding and computational analysis. The report finds that all models still fail on open-ended judgment, including matching Epoch's standards, designing informative experiments, and generating diverse ideas, so the authors conclude AI cannot yet replace workers at Epoch.

    AIWhy it matters: The report separates well-defined task reliability from open-ended judgment failures, which benchmark scores on easily verifiable tasks would miss.

  7. Google ResearchAI score23

    Join Alex Bie at the @COLM_conf Google booth #107 today at 5:00 PM for a walkthrough of ContinuousBench, a standardized benchmark designed to measure knowledge transfer in differentially private (DP) synthesis. Don't miss the chance to explore if DP synthetic data truly preserve information, or just style? @GoogleDeepMind Read the paper: https://arxiv.org/abs/2606.01849

    Join Alex Bie at the @COLM_conf Google booth #107 today at 5:00 PM for a walkthrough of ContinuousBench, a standardized benchmark designed to measure knowledge transfer in differentially private (DP) synthesis. Don't miss the chance to explore if DP synthetic data truly preserve information, or just style? @GoogleDeepMind Read the paper: https://arxiv.org/abs/2606.01849

  8. Google ResearchAI score42

    Today we announce new findings from Visiting Fellow David Autor on how AI impacts how professionals build expertise. In a three-month randomized controlled trial with practicing patent attorneys, we test both short-term productivity & longer-term skill building that occur as a result of AI usage. More: http://goo.gle/4AVeOWf

    Today we announce new findings from Visiting Fellow David Autor on how AI impacts how professionals build expertise. In a three-month randomized controlled trial with practicing patent attorneys, we test both short-term productivity & longer-term skill building that occur as a result of AI usage. More: http://goo.gle/4AVeOWf

  9. Google ResearchAI score62

    Google Research finds AI boosts patent drafting but junior lawyers' gains vanish without it

    A Google Research field experiment with 133 patent lawyers found AI tool access raised drafting scores by 0.34 to 0.38 standard deviations over three months. When the tool was removed for a redlining task, only senior lawyers kept an advantage of 0.45 SD, while junior lawyers showed no discernible improvement. The authors argue that tools which boost current output must not stop junior professionals from building the judgment that senior experts rely on.

    AIWhy it matters: The field experiment separates AI's short-term productivity gains from skill retained after the tool is removed, which matters for training junior professionals.

  10. NVIDIA AIAI score26

    An AI agent makes a mistake early in a task, then keeps going in the wrong direction. Our researchers built PivotOPD to teach agents how to avoid those mistakes and recover when they happen. During training, a teacher model shows the agent a better action and how to get back on track over the next few steps. Read the paper and watch how it works: https://research.nvidia.com/labs/lpr/pivotopd

    An AI agent makes a mistake early in a task, then keeps going in the wrong direction. Our researchers built PivotOPD to teach agents how to avoid those mistakes and recover when they happen. During training, a teacher model shows the agent a better action and how to get back on track over the next few steps. Read the paper and watch how it works: https://research.nvidia.com/labs/lpr/pivotopd