Skip to content

#Data/Training

Jul 25

Jul 25Sat
  1. Ali GhodsiAI score26

    This is one of those unintuitive things. Agents that cook longer are often worse. Genie just gets to the results faster. Ontology will be key to getting these agents the context they need to get the answers right quickly.

    This is one of those unintuitive things. Agents that cook longer are often worse. Genie just gets to the results faster. Ontology will be key to getting these agents the context they need to get the answers right quickly.

Jul 23

Jul 23Thu
  1. Sequoia CapitalAI score58

    Western AI Builders Depend on Chinese Open Models Through Distillation

    The essay argues that Western companies increasingly rely on Chinese open-weight models like Qwen and Kimi for post-training, while Western labs cannot lawfully distill from American frontier models. It says Qwen's share of new open-model fine-tunes rose from 1% in January 2024 to 69% by February 2026, citing ATOM's Report. The authors propose controlled teacher access and tighter enforcement against foreign distillation as a domestic alternative.

  2. Matei ZahariaAI score36

    AI-based research and engineering is one of the main topics we're researching in the STAR Lab at Berkeley, and something I'm seeing used in industry more and more too. This is nice work from my grad students packaging multiple AI-based "autoresearch" algorithms into one API in the GEPA package so you can mix-and-match them to get the best results. These can be used for anything from writing prompts to designing agents to optimizing code.

    AI-based research and engineering is one of the main topics we're researching in the STAR Lab at Berkeley, and something I'm seeing used in industry more and more too. This is nice work from my grad students packaging multiple AI-based "autoresearch" algorithms into one API in the GEPA package so you can mix-and-match them to get the best results. These can be used for anything from writing prompts to designing agents to optimizing code.

  3. Ahmad Al-DahleAI score62

    Ahmad Al-Dahle outlines five myths about AI model distillation

    Al-Dahle argues that distillation is a standard training method used inside labs, under licenses, or without authorization, so it does not by itself show theft. He says a few million conversations are small against trillion-token runs, yet can matter in late-stage training, reinforcement learning bootstrapping, or training a grader. He also argues that model outputs are hard to trace after paraphrasing or mixing, and that transferred capability is difficult to measure.

Jul 22

Jul 22Wed

Jul 21

Jul 21Tue
  1. Meta AI BlogAI score44

    Meta's SAM 3 and DINOv3 Power SYNAPS-I's Genesis Mission Imaging Pipeline

    SYNAPS-I, a multi-lab Genesis Mission project led by Lawrence Berkeley National Laboratory, uses Meta's open-source SAM 3 and DINOv3 models to segment X-ray and micro-CT scientific imagery. The fine-tuned pipeline, run on 300 A100 GPUs, reduced a grapevine xylem analysis from a month of expert annotation per time step to about 15 minutes. The team can deploy the open models inside secure national lab infrastructure, where research data must remain.

Jul 15

Jul 15Wed
  1. Fei-Fei LiAI score60

    RoboTTT scales robot policy context to 8,000 timesteps using test-time training

    Stanford SVL and NVIDIA Robotics introduced RoboTTT, which uses test-time training to give robot policies up to 8,000 timesteps of context at constant inference cost. The source reports that 8K-context pretraining beats 1K by 62%, and that performance keeps improving from 128 to 8K timesteps with no sign of saturation. The authors also describe one-shot imitation from human video and in-episode error recovery.

  2. Liquid AI NewsletterAI score38

    Liquid AI Releases Antidoom and IFStruct to Fix Reasoning Loops and Schema Errors

    Liquid AI released Antidoom, an open-source method that retrains a single overtrained token to eliminate "doom loops" in small reasoning models. On LFM2.5-2.6B and Qwen3.5-4B, loop rates fell from 10.2% to 1.4% and from 22.9% to 1%, respectively. The company also released IFStruct, an open-source benchmark measuring whether model outputs satisfy a schema, where LFM2.5-350M rose from 21.10% to 44.90% after training.

  3. Jim FanAI score62

    RoboTTT scales robot policy context to 8,000 timesteps with constant inference cost

    Jim Fan introduced RoboTTT, a robot model that uses test-time training to compress history into a tiny inner model updated at each sensor reading. The post reports closed-loop performance rising steadily from 128 to 8K timesteps, and 8K-context pretraining beating 1K by 62%. It also claims one-shot in-context learning from human video and mid-episode error recovery, with learning continuing after deployment.

Jul 12

Jul 12Sun
  1. ByteDance · new models on Hugging FaceAI score41

    ByteDance releases UniVR-34B-Planning for visual-space reasoning and planning

    ByteDance's UniVR-34B-Planning, built on Emu3.5 at 34B parameters, learns visual reasoning, physical dynamics, and long-term planning from visual demonstrations using a next-token objective and two-stage training on the VR-X dataset with VR-GRPO reinforcement learning. On the VR-X benchmark it scores 58.2 overall, up 18.4 points from the Emu3.5 34B baseline of 39.8. The Planning checkpoint is available on Hugging Face under CC BY 4.0, alongside a General checkpoint.

Jul 9

Jul 9Thu
  1. Thinking Machines LabAI score44

    Thinking Machines Argues the Future Worth Building Keeps Humans Central to AI Decisions

    Thinking Machines Lab says AI should extend human will and judgment, with people shaping its goals through continuous feedback rather than relying on models trained once and frozen. The company outlines three technical directions: training strong models, building tools for customization including training model weights, and developing interfaces that let personal judgment influence AI work. It also says it will publish research for the scientific community.

  2. Benedict EvansAI score60

    Benedict Evans argues AI token prices face unstable, commodity-leaning equilibrium

    Benedict Evans argues that token prices are unstable amid a supply crunch, and that foundation models may end up as low-margin commodity infrastructure rather than holding lasting pricing power. He cites inference gross margins of 40-50% that exclude training costs, which currently exceed revenue, and compares the outlook with mobile data and semiconductor manufacturing. He concludes that the outcome remains uncertain and that value capture above the model layer would require changes not yet visible.

Jul 8

Jul 8Wed
  1. Cognition Blog (Devin, Windsurf)AI score62

    Cognition releases SWE-1.7, a coding model trained with long-horizon RL

    Cognition launched SWE-1.7, which it says reaches frontier-level coding performance at lower cost, trained from a Kimi K2.7 base. The post describes RL methods including top-p sampling replay to preserve entropy, compressed weight deltas across multi-cluster training, and self-compaction for rollouts up to six hours. SWE-1.7 is available in Devin via Cerebras at 1000 TPS.

    AIWhy it matters: The post details entropy preservation, multi-cluster weight sync, and self-compaction, offering concrete RL training techniques for long-horizon coding agents to compare against one's own pipeline.

Jul 7

Jul 7Tue
  1. Berkeley AI ResearchAI score62

    Berkeley researchers outline how data systems must change as agents take over knowledge work

    Berkeley AI Research authors argue that near-free inference will make agents the dominant workload for data systems, requiring redesign for agentic speculation, agent-run state and coordination, and agent-synthesized systems. The post cites inference prices falling 9x to 900x per year with a median near 50x, and reports that about 80-90% of sub-queries in a text-to-SQL benchmark were duplicates. It frames the three directions as data systems for, of, and by agents.

    AIWhy it matters: The piece maps three concrete data-system challenges posed by near-free inference, useful for anyone designing infrastructure for agent workloads and memory.

Jul 3

Jul 3Fri
  1. Arthur MenschAI score34

    Mistral urges enterprises to adopt open-source models and own their AI data

    Mistral CEO Arthur Mensch argues enterprises should use open-source models and store their own data to avoid dependence on closed providers that retain customer data. He says companies should build continuous training loops from employee and user interactions, and shrink models to cut deployment costs. Mistral positions its Studio control plane and Forge training platform, deployed on customer infrastructure or via zero-data-retention hosting, as tools for this shift.

Jul 1

Jul 1Wed
  1. Jim FanAI score51

    Jim Fan introduces ASPIRE, a self-evolving robot skills library for continual learning

    Jim Fan announces ASPIRE, a system where coding agents use multimodal sensory traces from simulation and real robots to run evolutionary search over control programs and add the results to a growing skills library. The post claims up to a roughly 10x reduction in transfer learning tokens for sim-to-real and single-arm to bimanual transfer, and says the full stack will be open-sourced.

Jun 30

Jun 30Tue
  1. John SchulmanAI score38

    People sometimes ask why fine-tune when general-purpose models keep getting better. Bridgewater's work is a good reminder that with the right data -- here, expert judgements -- you can beat prompting-only approaches by a lot. @ddkang and the Bridgewater AIA Labs team are great -- glad to see them sharing this.

    People sometimes ask why fine-tune when general-purpose models keep getting better. Bridgewater's work is a good reminder that with the right data -- here, expert judgements -- you can beat prompting-only approaches by a lot. @ddkang and the Bridgewater AIA Labs team are great -- glad to see them sharing this.

  2. Jim FanAI score60

    ASPIRE lets robots build an evolving skills library that transfers across tasks

    Jim Fan introduces ASPIRE, a system in which coding agents observe multimodal sensory traces and run evolutionary search over control programs to distill skills into a growing library. The post says ASPIRE shares know-how rather than pixels or weights across the sim-to-real gap, reducing transfer learning tokens by up to about 10x. The author also says the full stack will be open-sourced and provides a gallery of 150+ tasks and 90+ skills.

Jun 29

Jun 29Mon
  1. Hamel HusainAI score54

    Why Hard-to-Eval AI Products Need Designs That Support Verification

    Hamel Husain argues that an AI product whose output is hard to verify is a product design problem, not just an evaluation problem. He shows before-and-after sketches for an AI data agent, a PE lesson planner, and a workers' compensation report tool, each adding provenance, scoped edits, and checkable evidence. He notes that designing for verification also makes evals easier to build and grade.

  2. Meta AI BlogAI score68

    Meta's Brain2Qwerty v2 decodes sentences from non-invasive brain recordings

    Meta released Brain2Qwerty v2, an end-to-end deep learning pipeline that decodes sentences in real time from non-invasive brain recordings. The model reached 61% word accuracy across participants, compared with 8% for other non-invasive methods, and 78% for the best participant. Meta also released the v1 and v2 training code, and partner BCBL released the v1 dataset.

    AIWhy it matters: The source reports word accuracy and data-scaling results for non-invasive decoding, offering a benchmark against surgical brain-computer interfaces and prior non-invasive methods.

Jun 28

Jun 28Sun
  1. PaddlePaddleAI score46

    👏Excited to see Unlimited-OCR now running in vLLM! Huge thanks to the @vllm_project community for the support and collaboration in bringing efficient long-context OCR to more developers. Check out the recipe and give it a try 👇

    👏Excited to see Unlimited-OCR now running in vLLM! Huge thanks to the @vllm_project community for the support and collaboration in bringing efficient long-context OCR to more developers. Check out the recipe and give it a try 👇

Jun 27

Jun 27Sat
  1. PaddlePaddleAI score36

    PaddleFormers 1.2 adds DeepSeek-V4 training with 128K+ context support

    PaddleFormers 1.2 is released with support for training DeepSeek-V4 and 128K+ long-context training. The update adds Context Parallel, Packing, Document Mask Attention, and the Muon optimizer, plus ultra-fused mHC, CSA, and HCA operators, DeepEP/HybridEP communication, and lossless FP8 training with AutoSubbatch memory balancing. The project is presented as fully open-source and is available on GitHub.

Jun 26

Jun 26Fri
  1. Qwen · new models on Hugging FaceAI score44

    Qwen3-ForcedAligner-0.6B-hf Adds Timestamp Alignment for Speech Transcripts

    Qwen released Qwen3-ForcedAligner-0.6B-hf, a Transformers-format forced aligner that predicts timestamps for arbitrary units within up to 5 minutes of speech in 11 languages. The model accepts transcripts from any ASR system, and the documentation shows it paired with Qwen3-ASR-0.6B and NVIDIA Parakeet CTC. Until it ships in an official Transformers release, users must install Transformers from source.

Jun 25

Jun 25Thu
  1. Lilian WengAI score40

    A super long overdue (3+ years?) post on scaling laws. Compute is expensive. Scaling laws are a way to help us reason about the optimal compute allocation between data and model size before committing to a large run. The post covers what scaling laws predict, how compute-optimal allocation works, why Kaplan et al. and Chinchilla disagree, and how data limits + fitting details make extrapolation tricky. https://lilianweng.github.io/posts/2026-06-24-scaling-laws/

    A super long overdue (3+ years?) post on scaling laws. Compute is expensive. Scaling laws are a way to help us reason about the optimal compute allocation between data and model size before committing to a large run. The post covers what scaling laws predict, how compute-optimal allocation works, why Kaplan et al. and Chinchilla disagree, and how data limits + fitting details make extrapolation tricky. https://lilianweng.github.io/posts/2026-06-24-scaling-laws/

  2. PaddlePaddleAI score38

    PP-OCRv6 recognition uses CTC and NRTR heads to curb hallucination

    PP-OCRv6's recognition module uses a CTC plus NRTR dual-head design so text is decoded from visual features rather than language priors, reducing hallucination. In hallucination tests, PP-OCRv6_medium reaches 93.2%, versus 85.0% for the best VLM, and recognition accuracy across 15 scenarios is 83.2%, above PP-OCRv5_server's 78.1%. NRTR is used only during training, adding language regularization at no inference cost, and it contributes +1.16% accuracy.

Jun 24

Jun 24Wed
  1. PaddlePaddleAI score30

    PP-OCRv6 Detection Module Outperforms VLMs on Text Localization Benchmarks

    PaddlePaddle says its PP-OCRv6_medium text detector reached an 86.2% detection Hmean in benchmarks, versus 46.8% for Gemini-3.1-Pro and 38.3% for GPT-5.5. The detector's design uses RepLKFPN with 7×7 kernels to cut FPN neck parameters from 172K to 118K, auxiliary deep supervision heads on P2–P4, and Focal Loss paired with Dice Loss, which adds +1.15% Hmean in ablation.

Jun 23

Jun 23Tue
  1. Lil'Log (Lilian Weng)AI score40

    Scaling Laws, Carefully: Early Empirical Power-Law Studies of Loss, Data and Model Size

    Lil'Log examines early empirical work showing that deep learning generalization error follows power-law curves as training data and model size grow. Hestness et al. (2017) found the exponent reflects the problem domain rather than the architecture, while Rosenfeld et al. (2020) modeled loss jointly as a function of model size N and data size D, fitting parametric forms on small configurations to extrapolate to larger ones.

Jun 19

Jun 19Fri
  1. AI Futures ProjectAI score60

    Forecast puts China's commercial EUV lithography in late 2030s

    The post argues that China's commercial-scale EUV machines should be forecast for the late 2030s and immersion DUV for the mid-2030s, using ASML's development timeline as a reference. It also weighs factors that could push these estimates earlier or later, including state funding, espionage, talent flows, and the use of AI in R&D. The authors note that forecasts placing either milestone in the 2020s would need strong justification.

Jun 18

Jun 18Thu
  1. OpenAI Alignment Research BlogAI score62

    OpenAI study finds beneficial-trait RL improves alignment across untrained domains

    OpenAI reports that reinforcement learning on realistic conversations targeting traits such as honesty, epistemic humility, and corrigibility improved a model across 44 out-of-distribution alignment evaluations. Gains included reward hacking, deception, and health benchmarks, and training only on health conversations still improved non-health alignment scores. The trained model was also harder to steer toward harmful behavior with adversarial persona prompts or harmful fine-tuning.

    AIWhy it matters: The post tests whether reinforcement learning on beneficial traits in one domain transfers to unrelated alignment benchmarks and holds up under adversarial steering.

Jun 17

Jun 17Wed
  1. John SchulmanAI score40

    PPO had a second wave in the LLM era for reasons unanticipated by the original paper - the importance-ratio objective fixes biases from numeric error, async training, and forward pass noise - the clipping objective affects entropy through a mechanism that we didn't know about at the time of publication (DAPO, https://arxiv.org/abs/2509.26114)

    PPO had a second wave in the LLM era for reasons unanticipated by the original paper - the importance-ratio objective fixes biases from numeric error, async training, and forward pass noise - the clipping objective affects entropy through a mechanism that we didn't know about at the time of publication (DAPO, https://arxiv.org/abs/2509.26114)

Jun 16

Jun 16Tue
  1. OpenAI Alignment Research BlogAI score60

    WildChat-based simulation predicts OpenAI production misalignment rates within roughly 3x

    OpenAI's alignment team found that re-generating 100,000 WildChat conversations with five recent OpenAI models predicted production failure rates across four orders of magnitude, with 95% of predictions within 1.04 orders of magnitude. The approach was weaker for agentic misalignment categories, where errors were about 37 times larger, and it still held roughly without access to chain-of-thought reasoning, with mean multiplicative error rising from 3.6x to 4.0x.

    AIWhy it matters: The post tests whether public chat data can predict real production failure rates, and where that prediction breaks down for agentic behavior.

  2. Arthur MenschAI score20

    We're working with companies and governments around the world to make sure their AI systems are up and running outside of external control, improving with each model release, and with an efficient cost structure. Forge allows to continuously train models based on recorded human-AI interaction, a key unlock for efficiency.

    We're working with companies and governments around the world to make sure their AI systems are up and running outside of external control, improving with each model release, and with an efficient cost structure. Forge allows to continuously train models based on recorded human-AI interaction, a key unlock for efficiency.

  3. Arthur MenschAI score22

    We've built Studio (for deployment) and Forge (for training) as portable products, and are now hosting them on infrastructure we control. We'll run in your VPC, your datacenter, or on our infrastructure that is decoupled from US service providers. We have capacity online, it's growing fast, and we can help you secure it.

    We've built Studio (for deployment) and Forge (for training) as portable products, and are now hosting them on infrastructure we control. We'll run in your VPC, your datacenter, or on our infrastructure that is decoupled from US service providers. We have capacity online, it's growing fast, and we can help you secure it.

Jun 10

Jun 10Wed
  1. ByteDance · new models on Hugging FaceAI score52

    ByteDance open-sources Bernini-Diffusers for semantic video generation and editing

    ByteDance open-sourced inference code and model weights for Bernini-Diffusers, a full video generation and editing pipeline with an MLLM-based semantic planner and a DiT-based renderer. The release bundles a Qwen2.5-VL planner and Wan2.2 diffusion components in one self-contained directory, and the source recommends it over the renderer-only Bernini-R for complex instruction following.

Jun 8

Jun 8Mon
  1. Xiaomi MiMoAI score62

    Xiaomi MiMo open-sources a 1T model running over 1,000 tps on 8 GPUs

    Xiaomi MiMo and the TileRT team say a 1T model exceeds 1,000 tps on a single standard 8-GPU node using general-purpose GPUs. The speedup comes from FP4 quantization and DFlash, a block-masked parallel speculative decoding method that accepts more tokens per verification, with TileRT tailoring its compiler and kernels to these techniques. Open weights for the FP4 + DFlash checkpoint are available on Hugging Face.

Jun 6

Jun 6Sat
  1. Ahead of AI (Sebastian Raschka)AI score32

    Raschka Lists 2026 LLM Research Papers from January Through May, Heavy on Reasoning and Efficiency

    Sebastian Raschka has published a curated list of LLM research papers he bookmarked from January through May 2026, not a complete survey of the field. The list is weighted toward reasoning models, reinforcement learning, and efficient inference, with added interest in agent harnesses, long context, and diffusion language models. He highlights Nvidia's Nemotron 3 Super, a 120B-A12B hybrid model alternating attention and Mamba-2 layers, as a must-read, and notes a 4B Nano variant and the 550B-A55B Nemotron 3 Ultra released two days earlier.

May 29

May 29Fri

May 26

May 26Tue
  1. One Useful Thing (Ethan Mollick)AI score40

    Mollick Warns AI Writing Defaults Erode Learning and Human Thinking

    Ethan Mollick argues that using AI as a default for writing, without thinking, risks undermining the human effort that builds skill and style. He cites two Wharton-linked studies: a Turkish high school experiment where ChatGPT access hurt test performance, and a Taipei Python course where a personalized AI tutor raised exam scores by 0.15 standard deviations. Mollick calls the difference how AI is used, not whether, and notes that the tools for tutor-style learning are not intuitive to access.