Skip to content

#Data/Training

Jan 27

Jan 27Tue
  1. Tim DettmersAI score72

    Tim Dettmers Details How SERA Built an Open Coding Agent on 32 GPUs

    Ai2's Open Coding Agents family, with SERA as its first release, was built by Tim Dettmers and collaborators on 32 GPUs. The method generates synthetic bug trajectories with soft verification, comparing patches by line overlap instead of running tests. The post reports that a 32B model fine-tuned on about 7,000 trajectories for one private repository matched its GLM 4.5-Air teacher, and that the baseline costs $500 to run.

Jan 22

Jan 22Thu
  1. BAAIAI score38

    BAAI releases RoboCOIN, a large bimanual robot manipulation dataset

    BAAI's RoboCOIN is a bimanual robot dataset with more than 180,000 trajectories across 421 tasks, collected from 15 robot platforms in 16 real-world scenarios. Its three-tier annotations at trajectory, segment, and frame levels help robots learn both what to do and how to do it. The post says integrating these annotations raised success rates on complex tasks by up to 50% for models such as π₀.

Jan 14

Jan 14Wed
  1. Black Forest Labs · new models on Hugging FaceAI score46

    FLUX.2 [klein] 9B Base Released on Hugging Face as Undistilled Open-Weight Model

    Black Forest Labs has released FLUX.2 [klein] 9B Base, a 9 billion parameter undistilled rectified flow transformer with open weights for text-to-image generation and multi-reference editing. The model is intended for fine-tuning, LoRA training, and research, and fits in about 29GB VRAM on NVIDIA RTX 4090-class GPUs. A reference implementation is available on GitHub, and the model works with ComfyUI and Diffusers.

Jan 10

Jan 10Sat
  1. Berkeley AI ResearchAI score36

    Information-Driven Design Framework Evaluates Imaging Systems by Mutual Information

    Berkeley AI Research proposes an information-based framework that evaluates and optimizes imaging systems using mutual information estimated directly from noisy measurements. The team reports that the metric predicts decoder performance across color photography, radio astronomy, lensless imaging, and microscopy, and that optimized designs match end-to-end methods while requiring less memory and compute.

Jan 9

Jan 9Fri
  1. BAAIAI score47

    🚀 Science! Tsinghua AIR & BAAI DrugCLIP redefines drug discovery with AI. ⚡ 1M× faster screening — 10 trillion protein–molecule pairs/day. ⚡ Full genome-scale drug-target mapping — screened 10k proteins × 500M molecules, identifying 2M+ candidates. DrugCLIP bridges AlphaFold’s structures to real drug candidates — launching the post‑AlphaFold era of scalable discovery. 📄Paper: https://www.science.org/doi/10.1126/science.ads9530 🌐Platform: https://www.drugclip.com/ #AI #Science

    🚀 Science! Tsinghua AIR & BAAI DrugCLIP redefines drug discovery with AI. ⚡ 1M× faster screening — 10 trillion protein–molecule pairs/day. ⚡ Full genome-scale drug-target mapping — screened 10k proteins × 500M molecules, identifying 2M+ candidates. DrugCLIP bridges AlphaFold’s structures to real drug candidates — launching the post‑AlphaFold era of scalable discovery. 📄Paper: https://www.science.org/doi/10.1126/science.ads9530 🌐Platform: https://www.drugclip.com/ #AI #Science

Dec 2, 2025

Dec 2, 2025Tue
  1. Apple · new models on Hugging FaceAI score36

    Apple releases CLaRa-7B-Instruct for compressed-document retrieval-augmented QA

    Apple has published CLaRa-7B-Instruct on Hugging Face, an instruction-tuned unified RAG model with built-in semantic document compression at 16× and 128× ratios. The model answers instruction-following questions directly from compressed document representations, and its paper, GitHub repository, and transformers usage example are referenced in the release.

Nov 28, 2025

Nov 28, 2025Fri

Nov 17, 2025

Nov 17, 2025Mon
  1. Andrej KarpathyAI score60

    Karpathy argues verifiability predicts which tasks AI automates fastest

    Karpathy argues that verifiability, not specifiability, is the most predictive feature for AI automation, since verifiable tasks can be optimized directly or through reinforcement learning. He says a task is suited to this approach when the environment is resettable, efficient, and rewardable. This explains the jagged frontier of LLM progress, with verifiable domains like math and code advancing rapidly while creative and strategic tasks lag behind.

Nov 5, 2025

Nov 5, 2025Wed

Oct 27, 2025

Oct 27, 2025Mon

Oct 26, 2025

Oct 26, 2025Sun
  1. Thinking Machines LabAI score70

    Thinking Machines Lab explains on-policy distillation for cheaper LLM post-training

    Thinking Machines Lab describes on-policy distillation, which samples rollouts from a student model and has a teacher grade each token with reverse KL. The authors report that this matches Qwen3-style reasoning results at a fraction of RL's cost, with AIME'24 reaching 70% in about 150 steps from a 400k SFT checkpoint. The method also helps recover instruction-following behavior lost during fine-tuning on internal documents.

    AIWhy it matters: The post explains why on-policy distillation gives dense per-token feedback, letting a small model match RL results at much lower compute cost.

Sep 3, 2025

Sep 3, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score38

    Eight Sleep Uses Devin AI as Data Analyst to Clear Ad-Hoc Requests

    Eight Sleep integrated Cognition's Devin into its data workflows, letting staff tag Devin in Slack to query Snowflake, dbt, and Looker and check Amplitude. The company says it is now shipping 3x as many data features and investigations each week, with its ad-hoc data request queue near zero. Devin was used to trace a suspicious revenue spike to a better-than-expected email campaign.

Aug 27, 2025

Aug 27, 2025Wed

May 5, 2025

May 5, 2025Mon
  1. Cognition Blog (Devin, Windsurf)AI score39

    Kevin-32B Uses Multi-Turn Reinforcement Learning to Write Faster CUDA Kernels

    Stanford and Cognition AI researchers introduced Kevin-32B, a 32B-parameter model trained with multi-turn reinforcement learning to write CUDA kernels. On KernelBench, it solves 89% of tasks at best@16 and achieves 65% average correctness over eight refinement steps, versus 53% for o4-mini and 51% for o3. Its best@16 speedup is 1.41x, and multi-turn training outperforms single-turn training as refinement steps increase.

Nov 30, 2024

Nov 30, 2024Sat
  1. Liquid AI BlogAI score60

    Liquid AI's STAR uses evolutionary search to synthesize tailored model architectures

    Liquid AI reports STAR, an evolutionary algorithm that synthesizes tailored neural network architectures from a numerical genome representation. The authors say it produced hundreds of designs that outperform Transformer and hybrid architectures in quality, with smaller caches and parameter counts, and can optimize for latency on target hardware. The full method is described in the arXiv technical report 2411.17800.

    AIWhy it matters: The post explains how evolutionary search over a new architecture design space produced designs beating Transformers and hybrids, giving a concrete method for quality versus latency and memory trade-offs.