Skip to content

#Data/Training

Oct 8

TodayOct 8Thu30 items
  1. PandailyAI score46

    ByteDance Seed Finds Periodic Weak Spots in Chunked KV-Cache Compression

    ByteDance Seed researchers found that language models compressing their KV cache in fixed-size chunks retrieve the same information unevenly depending on token position. In a 128K-token needle-in-a-haystack test, base DeepSeek-V4 checkpoints differed by up to 40.2 percentage points by phase, and post-training narrowed but did not eliminate the gaps. The authors urge evaluating such models across positional phases, since high average accuracy can hide systematic failures.

  2. PandailyAI score41

    Simplexity Robotics Trains One Robot to Tend Two CNC Lathes With 600 Trajectories

    Simplexity Robotics says it trained a single robot to load and unload two CNC lathes on its own, using 600 real-robot trajectories and reporting a 100% success rate on the precision CNC insertion task. The work, presented at IROS 2026 on September 29, combines the SimpleWAM world action model, a DRAM memory module and DPE action scoring, with force and torque feedback for insertion recovery. The company did not say how many trials the 100% figure covers, and it does not describe the yield of the whole cell.

  3. QbitAI (量子位)AI score58

    AgentGarten lets agents evolve through code-built worlds and neural rendering

    MirroS released AgentGarten, which pairs executable code environments with a real-time neural renderer running above 30 fps so agents can act, observe, and learn. In a one-on-one hide-and-seek setup, the hider learned to block passages by round 4 and the seeker learned to climb ramps by round 10, guided by notes the agents wrote after each round. The authors report applying the same loop to four other tasks, including a dog-companion game, a narrow-bridge car passing task, herding, and quarry loading.

  4. PandailyAI score55

    Chinese Team Publishes 3D Cell Atlas of Rice's Full Life Cycle in Cell

    A Chinese-led team published in Cell a three-dimensional spatiotemporal cell atlas covering rice from germinating seed to grain fill, along with a public portal and the RICE scGPT single-cell foundation model. The atlas combines single-nucleus RNA sequencing with BGI's Stereo-seq spatial transcriptomics across 10 organ and tissue types and 61 stages, defining 119 cell types and 133 subtypes.

  5. AnthropicAI score57

    Astrophysicist uses Claude to build first complete ultraviolet sky map

    An astrophysicist worked with Claude Science to create the first complete ultraviolet map of the sky, covering regions never observed in UV. Claude located existing datasets, combined them, and filled gaps with statistical inference, taking a few days rather than weeks of human work. The map is presented as a teaching tool and an example of low-priority scientific work that AI now makes feasible.

  6. Epoch AI · The Epoch BriefAI score49

    Epoch AI's October 2026 Brief Covers AI Agents, Falling Costs, and China's Chip Exposure

    Epoch AI estimates the AI chips shipped through 2027 could support about 30 to 170 million concurrent frontier-model agents, or nearly 2 billion with cheaper models. Its researchers find the cost of a fixed level of AI performance has fallen about 47% per quarter over the past three years. The newsletter also reports China's semiconductor supply-chain exposure is 2.7 times that of the US.

  7. Google ResearchAI score20

    Interested in privacy-preserving synthetic data? Catch Alex Bie at the @COLM_conf Google booth (#107) today, at 2:00 PM PT for an encore presentation of ContinuousBench, evaluating knowledge transfer and data contamination in DP synthesis.@GoogleDeepMind Join the conversation at #COLM2026!

    Interested in privacy-preserving synthetic data? Catch Alex Bie at the @COLM_conf Google booth (#107) today, at 2:00 PM PT for an encore presentation of ContinuousBench, evaluating knowledge transfer and data contamination in DP synthesis.@GoogleDeepMind Join the conversation at #COLM2026!

  8. GoodfireAI score44

    The Alzheimer’s Translation Challenge is built on a new 150M-cell atlas of neurons, astrocytes, and microglia across different genetic backgrounds under combinatorial perturbations, with multi-modal readouts. The data will be made available through the AD workbench and Prima Mente's modeling platform.

    The Alzheimer’s Translation Challenge is built on a new 150M-cell atlas of neurons, astrocytes, and microglia across different genetic backgrounds under combinatorial perturbations, with multi-modal readouts. The data will be made available through the AD workbench and Prima Mente's modeling platform.

  9. Elvis SaraviaAI score46

    RSIGym gives research agents services, lifting SWE-bench Verified to 50.33%

    RSIGym provides a research agent with training, inference, evals, and sandboxes as callable services, so it spends its budget on experiments rather than rebuilding infrastructure. With Opus 5 as the researcher, the improved system rose from 17.67% to 50.33% on SWE-bench Verified. The post also highlights a way to measure co-evolution between harnesses and models.

  10. Google ResearchAI score14

    Missed yesterday's demo on adaptive agent environments? Stop by the #COLM2026 Google booth #107 today at 11:00 AM PT to catch Zifeng Wang presenting EnvHarness — a flexible framework enabling co-evolution between LLM agents and their training environments.

    Missed yesterday's demo on adaptive agent environments? Stop by the #COLM2026 Google booth #107 today at 11:00 AM PT to catch Zifeng Wang presenting EnvHarness — a flexible framework enabling co-evolution between LLM agents and their training environments.

  11. GoogleAI score40

    In our prospective clinical study, our models were able to pinpoint the gestational age within 4 days of accuracy which is a tight enough zone that can have a really meaningful clinical impact. If we're able to expand these tools to low resource settings, we can start to move the needle on maternal deaths and bridge that gap in care that we see all over the world.

    In our prospective clinical study, our models were able to pinpoint the gestational age within 4 days of accuracy which is a tight enough zone that can have a really meaningful clinical impact. If we're able to expand these tools to low resource settings, we can start to move the needle on maternal deaths and bridge that gap in care that we see all over the world.

  12. ZyphraAI score34

    Our approach is lossless. The same tokens visit the same experts. The architecture, routing decisions and training objective stay unchanged. The model computes the same function. By organizing where experts and tokens live we reduce the communication needed to do the same work.

    Our approach is lossless. The same tokens visit the same experts. The architecture, routing decisions and training objective stay unchanged. The model computes the same function. By organizing where experts and tokens live we reduce the communication needed to do the same work.

  13. ZyphraAI score38

    The gains depend on the training configuration, and are largest when each token uses more experts and when the experts span several nodes. Results in Megatron-LM on 8 to 64 GPUs: - Token exchange: 1.16x to 2.63x faster - Full training step: up to 1.41x faster

    The gains depend on the training configuration, and are largest when each token uses more experts and when the experts span several nodes. Results in Megatron-LM on 8 to 64 GPUs: - Token exchange: 1.16x to 2.63x faster - Full training step: up to 1.41x faster

  14. ZyphraAI score22

    Crucially, we find that these patterns emerge early enough to be useful during training and are cheap to measure. Using just 4,096 tokens, our placement and routing predictions come within roughly one percentage point of results using 524,000 tokens.

    Crucially, we find that these patterns emerge early enough to be useful during training and are cheap to measure. Using just 4,096 tokens, our placement and routing predictions come within roughly one percentage point of results using 524,000 tokens.

  15. ZyphraAI score18

    Routing is also predictable across layers: the experts a token used in one layer tell us which it will likely need next. Token shuffling moves each token to the GPU holding those experts, inside a transfer that already runs after attention, so it adds no network traffic.

    Routing is also predictable across layers: the experts a token used in one layer tell us which it will likely need next. Token shuffling moves each token to the GPU holding those experts, inside a transfer that already runs after attention, so it adds no network traffic.

  16. ZyphraAI score37

    So we put experts that are picked together on the same GPU, and send each token to a GPU once no matter how many of its experts live there. We call this correlated expert placement. On 8 GPUs it removes up to 58% of the token copies sent.

    So we put experts that are picked together on the same GPU, and send each token to a GPU once no matter how many of its experts live there. We call this correlated expert placement. On 8 GPUs it removes up to 58% of the token copies sent.

  17. ZyphraAI score32

    MoE models route each token sparsely to just a few experts. As these models grow, the experts are spread across GPUs and nodes, and tokens have to be sent to their experts and back. In our runs that exchange took 13-24% of step time on one node and 45-60% across four nodes.

    MoE models route each token sparsely to just a few experts. As these models grow, the experts are spread across GPUs and nodes, and tokens have to be sent to their experts and back. In our runs that exchange took 13-24% of step time on one node and 45-60% across four nodes.

  18. ZyphraAI score23

    Those routing decisions have patterns. For instance, in one layer, just 64 of 8,128 possible expert pairs accounted for 42% of tokens. If experts were picked independently, those pairs would carry 1.6%. We show how to exploit this structure to improve where experts are placed.

    Those routing decisions have patterns. For instance, in one layer, just 64 of 8,128 possible expert pairs accounted for 42% of tokens. If experts were picked independently, those pairs would carry 1.6%. We show how to exploit this structure to improve where experts are placed.

  19. ZyphraAI score22

    In Mixture of Expert (MoE) models the cost of moving tokens to their experts can dominate total runtime. At Zyphra research, we use patterns in how tokens are routed to experts to make that communication faster by up to 2.63x on @AMD MI300X GPUs, with the model unchanged.

    In Mixture of Expert (MoE) models the cost of moving tokens to their experts can dominate total runtime. At Zyphra research, we use patterns in how tokens are routed to experts to make that communication faster by up to 2.63x on @AMD MI300X GPUs, with the model unchanged.

  20. OpenBMBAI score36

    ReJev fine-tunes MiniCPM5-2B to lift decision accuracy to 80.50%

    ReJev, an independent community project, applied LoRA post-training to OpenBMB's MiniCPM5-2B for bounded agent decisions: state, question, and candidate options yield one choice. On its sealed 1,892-sample holdout, accuracy rose from 51.11% to 80.50% (+29.39 percentage points) with 0% invalid outputs, at about $5.31 in cumulative Modal billing including earlier experimental overhead. The authors describe this as an early, task-specific result, not parity with Jev.

  21. MarkTechPostAI score45

    NVIDIA's PivotOPD Trains Multi-Turn AI Agents to Recover From Pivotal Mistakes

    NVIDIA, Princeton University, and the University of Maryland introduced PivotOPD, an on-policy distillation method that teaches multi-turn LLM agents to recover from their most damaging early mistake. Tested on Qwen3-1.7B and Qwen3-8B students, it posts the best average against 13 baselines on ALFWorld, WebShop, and Search-based QA. It recovers from 72.7% of replayed pivotal mistakes, versus 20.3% for standard OPD, with no added inference cost.

  22. PandailyAI score46

    Galbot and Tsinghua's LATENT Wins IROS Award for Humanoid Tennis Forehand

    A Galbot, Tsinghua University and collaborators paper won IROS 2026's Best Entertainment and Amusement Paper Award for LATENT, a humanoid tennis-return method trained on imperfect amateur motion-capture clips. In simulation, the full forehand policy succeeded on 96.52 percent of returns, versus 71.85 percent for PULSE. On a real Unitree G1, the paper reports 90.90 percent forehand success across 20 consecutive rallies, with motion capture still used rather than the robot's own cameras.

  23. PandailyAI score38

    Huawei Presents Experimental XMFS Shared-Memory Filesystem at LPC 2026

    Huawei engineers presented XMFS, an experimental Linux kernel prototype filesystem, at the Linux Plumbers Conference in Prague on October 5. It aims to let applications reach cross-node shared memory on CXL 3.0 or Huawei unified bus servers through standard POSIX file calls. The code exists only on openEuler, not in the mainline Linux kernel.

  24. Anthropic ResearchAI score62

    Anthropic researcher builds first complete UV sky map with Claude Science

    Johns Hopkins astrophysicist Brice Ménard, working as an Anthropic researcher, used Claude Science to produce the first complete map of the sky in ultraviolet light. Claude orchestrated agents to merge GALEX, Swift, and FIMS/SPEAR data, then predicted roughly a third of the sky that no UV telescope had observed, using relationships to visible, infrared, and radio data. Hidden test regions were reconstructed to within about 10% of real measurements, and each pixel is labeled measured or predicted with uncertainty estimates.

    AIWhy it matters: The post shows how an astrophysicist used Claude Science agents to merge UV surveys and predict missing sky regions, with a validation step that makes the method reusable.

Oct 7

Oct 7Wed
  1. Apple Machine Learning ResearchAI score42

    Apple's Normalizing Trajectory Models generate images in four steps with exact likelihood

    Apple researchers introduced Normalizing Trajectory Models (NTM), which model each reverse diffusion step as a conditional normalizing flow trained with exact likelihood. The model matches or outperforms strong image generation baselines on text-to-image benchmarks in just four sampling steps while retaining exact likelihood over the generative trajectory.

  2. Epoch AIAI score67

    Epoch tests six AI models on real Epoch work and finds they cannot yet fully automate it

    Epoch gave six models 11 real work tasks from its own operations, including graphic design, data insights, and research design, and graded outputs against employee standards. Fable 5.1 and GPT-6 Astra led on average task performance, reliably handling well-defined work such as coding and computational analysis. The report finds that all models still fail on open-ended judgment, including matching Epoch's standards, designing informative experiments, and generating diverse ideas, so the authors conclude AI cannot yet replace workers at Epoch.

    AIWhy it matters: The report separates well-defined task reliability from open-ended judgment failures, which benchmark scores on easily verifiable tasks would miss.

  3. Google ResearchAI score23

    Join Alex Bie at the @COLM_conf Google booth #107 today at 5:00 PM for a walkthrough of ContinuousBench, a standardized benchmark designed to measure knowledge transfer in differentially private (DP) synthesis. Don't miss the chance to explore if DP synthetic data truly preserve information, or just style? @GoogleDeepMind Read the paper: https://arxiv.org/abs/2606.01849

    Join Alex Bie at the @COLM_conf Google booth #107 today at 5:00 PM for a walkthrough of ContinuousBench, a standardized benchmark designed to measure knowledge transfer in differentially private (DP) synthesis. Don't miss the chance to explore if DP synthetic data truly preserve information, or just style? @GoogleDeepMind Read the paper: https://arxiv.org/abs/2606.01849

  4. Google ResearchAI score42

    Today we announce new findings from Visiting Fellow David Autor on how AI impacts how professionals build expertise. In a three-month randomized controlled trial with practicing patent attorneys, we test both short-term productivity & longer-term skill building that occur as a result of AI usage. More: http://goo.gle/4AVeOWf

    Today we announce new findings from Visiting Fellow David Autor on how AI impacts how professionals build expertise. In a three-month randomized controlled trial with practicing patent attorneys, we test both short-term productivity & longer-term skill building that occur as a result of AI usage. More: http://goo.gle/4AVeOWf

  5. NVIDIA AIAI score26

    An AI agent makes a mistake early in a task, then keeps going in the wrong direction. Our researchers built PivotOPD to teach agents how to avoid those mistakes and recover when they happen. During training, a teacher model shows the agent a better action and how to get back on track over the next few steps. Read the paper and watch how it works: https://research.nvidia.com/labs/lpr/pivotopd

    An AI agent makes a mistake early in a task, then keeps going in the wrong direction. Our researchers built PivotOPD to teach agents how to avoid those mistakes and recover when they happen. During training, a teacher model shows the agent a better action and how to get back on track over the next few steps. Read the paper and watch how it works: https://research.nvidia.com/labs/lpr/pivotopd

  6. Google ResearchAI score10

    LLM agents learn by interacting with environments, but static setups limit their growth. Today at 2:00 PM, join Zifeng Wang at the #COLM2026 Google booth (#107) to learn about EnvHarness, a plug-in architecture that dynamically reshapes environment behaviors to improve reinforcement learning and agent adaptability.

    LLM agents learn by interacting with environments, but static setups limit their growth. Today at 2:00 PM, join Zifeng Wang at the #COLM2026 Google booth (#107) to learn about EnvHarness, a plug-in architecture that dynamically reshapes environment behaviors to improve reinforcement learning and agent adaptability.

  7. Epoch AIAI score22

    We instructed the AI models that their technique should improve performance on several benchmarks. We already knew that all of these could be improved by a recent human-authored post-training innovation: on-policy self-distillation (SDPO).

    We instructed the AI models that their technique should improve performance on several benchmarks. We already knew that all of these could be improved by a recent human-authored post-training innovation: on-policy self-distillation (SDPO).

  8. Epoch AIAI score50

    AI developers aim to create an automated AI researcher. How close are they? To find out, we built InnovationEval, which tests whether AI can produce post-training innovations comparable in magnitude to a recently published advance. So far, agents’ results are underwhelming.

    AI developers aim to create an automated AI researcher. How close are they? To find out, we built InnovationEval, which tests whether AI can produce post-training innovations comparable in magnitude to a recently published advance. So far, agents’ results are underwhelming.