Skip to contentSkip to stories

Updated

#Data/Training

Sep 24

Sep 24Thu
  1. LangChain BlogAI score50

    LangSmith Fine-Tuning and smithtune Turn Agent Trajectories Into Custom Models

    AILangChain launched LangSmith Fine-Tuning and smithtune, a CLI that turns LangSmith agent trajectories into fine-tuned models through dataset creation, training with Fireworks or Baseten, and evaluation in LangSmith. smithtune currently supports supervised fine-tuning, training models on recorded examples of good agent behavior by updating model weights. The tool lets teams train specialized models without building the data pipeline by hand.

Sep 23

Sep 23Wed
  1. Google Developers BlogAI score62

    Google reproduces Olmo 3 7B pre-training in MaxText on TPUs

    AIGoogle Developers reproduced Ai2's Olmo 3 7B from scratch in MaxText on Google Cloud TPUs, covering both the stage-1 pre-training run and the stage-2 mid-training anneal. The match was checked on held-out C4 loss, an 8-task accuracy suite, multi-domain perplexity, and token-level KL, not just the training loss curve. The post also describes a data-loader bug that made training loss look better than the reference while held-out metrics did not move.

    Why it matters: The post documents how a faithful reproduction was verified on held-out metrics, including a data bug that training loss alone would have hidden.

  2. Google DeepMindAI score62

    Google DeepMind details server-side memory for Private AI Compute

    AIGoogle DeepMind describes a persistent memory layer for its Private AI Compute platform that stores user context encrypted in the cloud. The encryption keys are held on the user's devices, and data is decrypted only inside hardware-isolated secure enclaves before being re-encrypted. The company says it is publishing a tamper-proof public record of its server software and an independent audit.

    Why it matters: The post explains how persistent cloud memory can keep personal AI context encrypted under keys held on the user's device, a concrete privacy design.

  3. Baseten BlogAI score62

    Baseten launches NVIDIA Nemotron 3 Diarization with four latency profiles

    AIBaseten has made NVIDIA Nemotron 3 Diarization available as batch, streaming, and real-time diarized transcription presets. The single checkpoint serves four algorithmic latencies from 0.32 to 30.4 seconds, and the post reports DER of 9.8% on AISHELL-4 at the low profile versus 27.2% for Streaming Sortformer v2.1.

    Why it matters: The post shows one checkpoint serving four latency profiles with DER figures against named baselines, useful for judging real-time speaker labeling tradeoffs.

  4. Anthropic NewsroomAI score73

    Claude agents discover a novel CRISPR-like enzyme system in bacteriophages

    AIAnthropic's new life sciences group reports that Claude autonomously identified a previously uncharacterized enzyme system, called array-associated reverse transcriptase (ART), in bacteriophages. Claude agents searched over 200,000 reverse transcriptases, narrowed 3,500 candidates to 20, and one agent flagged a CRISPR-like repeat array after about 21 hours. Human scientists then validated the finding in the lab, and the function of ART remains unknown.

    Why it matters: The post shows how Claude agents surveyed DNA sequence data, flagged a candidate, and then led to lab validation, which is a concrete workflow for AI-assisted biology research.

  5. Prime Intellect BlogAI score60

    Prime Intellect makes Prime Sandboxes generally available as microVMs for agentic RL

    AIPrime Intellect has made Prime Sandboxes generally available, offering each sandbox as a full Linux virtual machine with its own kernel and support for Docker Compose. The product is available through its CLI/SDK and RL suite, with accounts starting at 1,024 concurrent sandboxes, and pricing listed at $0.02 per vCPU-hour, $0.0125 per GiB-hour of memory, and $0.0002 per GiB-hour of disk, valid through December 22. The company says GPU microVMs, snapshotting, sandbox forking, and persistent workspaces are planned next.

    Why it matters: The post explains why full VMs rather than gVisor containers matter for agentic RL, since silent environment differences can reward behaviors that fail to transfer.

Sep 22

Sep 22Tue
  1. Together AI BlogAI score38

    How to train your own Jev classifier for $17 with Together AI

    AIThe Together AI blog shows how to fine-tune a Qwen3.5 4B base model into a classification model using about 38,000 examples sampled from six Hugging Face datasets, at a training cost of roughly $17.0. The tutorial covers cloning the tev1 repository, normalizing data with provided scripts, launching a Together AI fine-tuning job that takes about 25 minutes, and deploying the result to a dedicated H100 endpoint.

  2. Fireworks AI BlogAI score46

    Fireworks ARCv3 cuts RL weight-update payloads nearly 50% for cross-region training

    AIFireworks released ARCv3, a lossless compressor for BF16 weight-update deltas sent from trainers to RL rollout machines. Across 1,000 production RL deltas, ARCv3 produced payloads nearly 50% smaller than ARCv2, averaging about 0.19% of the BF16 weight size versus 0.36%. ARCv3 is available through the Fireworks Training API as fireworks-delta-compression.

  3. Black Forest Labs · new models on Hugging FaceAI score58

    Black Forest Labs releases open-weights FLUX 3 Action SO-101 robot policy

    AIBlack Forest Labs has published FLUX 3 Action SO-101 on Hugging Face as an open-weights 7B world action model. It takes two camera frames, the robot state, and a text instruction, then returns the next 42 actions with predicted video frames, with 32 executed at 30 Hz before replanning. The card also provides a rank-32 LoRA fine-tuning recipe for user datasets and states that the application must enforce joint velocity, force, and workspace limits.

Sep 21

Sep 21Mon
  1. Google Developers BlogAI score38

    Google Colab premium benefits now included in Google AI plans

    AIGoogle AI subscribers now get premium Colab benefits, including priority access to faster accelerators and more powerful machines. Google AI Ultra subscribers also get uninterrupted background execution and Premium GPU access for long training runs. The benefits roll out over the next few weeks in Colab-supported countries, and existing Colab subscriptions are unchanged.

  2. Apple · new models on Hugging FaceAI score46

    Apple releases LensVLM-9B, a vision-language model for compressed text images

    AIApple has released LensVLM-9B on Hugging Face, a 9B-parameter Vision Language Model that scans compressed images of text and selectively expands relevant pages to their uncompressed form. The repository provides a demo script and supports compression settings of 5x, 10x, and 15x. Model files are under the Apple Machine Learning Research Model License, and the accompanying source code is distributed separately under the Apple Sample Code License.

  3. Amazon ScienceAI score60

    Amazon Science reports AI models for designing and characterizing antibodies

    AIAmazon Science describes three papers on AI for antibody discovery: MochiBind ranks antibody binding strength from sequence alone, CA-MAP predicts developability properties using batch-aware context, and an agent-guided pipeline designed nanobody binders against a novel cancer target. In the pipeline, 116 candidates survived lab screening, and 46 were identified as strong binders, which are being used to train the next design cycle.

    Why it matters: The source reports the method, benchmark setup, and experimental validation in a single design workflow, showing how predictors, agents, and lab screening connect in antibody discovery.

  4. Microsoft ResearchAI score50

    Microsoft Research open-sources RetroChimera, a retrosynthesis model published in Nature

    AIMicrosoft Research published RetroChimera, a retrosynthesis framework that combines the R-SMILES 2 Transformer model and the NeuralLoc graph neural network through learned ensembling to propose synthesis routes for small molecules. In blind tests, PhD-level chemists preferred its individual reaction predictions over those from preceding models and recorded literature reactions. The implementation and weights are open-sourced for researchers developing new medicinal molecules and materials.

Sep 18

Sep 18Fri
  1. Google · AI blogAI score29

    Google adds Nobel laureate Philippe Aghion and new directors to AI & Economy research team

    AIGoogle's AI & Economy Research Program has added Nobel laureate Philippe Aghion as an Academic Advisor and Ajay Agrawal as a Visiting Fellow, with Anu Madgavkar and Daniel Rock named Directors. The program covers the future of work, productivity and growth, global technology diffusion, and AI's impact on scientific discovery, and builds on the recently launched AI & Economy ATLAS v1.0.

Sep 17

Sep 17Thu
  1. xAI News (Grok)AI score42

    Grok Voice Transcribe 2.0 Doubles Accuracy of Predecessor at Same Price

    AIxAI released Grok Voice Transcribe 2.0, a speech-to-text model that is twice as accurate as Grok Voice Transcribe 1.0 at the same price, and ranks first for accuracy among 32 streaming models on the Artificial Analysis leaderboard. Batch transcription costs $0.10 per hour of audio and streaming $0.20 per hour, with diarization, timestamps, and key terms included. Existing Speech-to-Text API integrations gain the improvement with no code changes, and developers must pin grok-voice-transcribe-1.0 to stay on the older model during the transition.

  2. Google · AI blogAI score38

    UN System Data Commons unifies global statistics into an AI-ready open platform

    AIThe United Nations system launched UN System Data Commons, an open-source platform built on Data Commons by Google that integrates siloed global statistics into one AI-ready knowledge graph. Users can query it in natural language, browse by location or theme, and use MCP-enabled AI agents to fetch verified figures and draft charts or reports. The UN plans to add more datasets, aiming to include 80% of UN system statistical datasets by 2027.

  3. Ai2 (Allen Institute for AI)AI score42

    Crowdsourced Game Steering Arena Shows Olmo 3 Prosocial Scores Can Be Gamed

    AINortheastern University MS student Soham Padia used Ai2's open Olmo 3-32B model to build Steering Arena, a public game in which players submit text prefixes to steer prosocial behavior. About 600 submissions from a few dozen people showed the top 36 entries were unreadable token strings, while the best plain-English entry ranked 37th at about 2.7 times lower score. The results suggest that once an evaluation metric is exposed, it becomes an optimization target.

Sep 15

Sep 15Tue
  1. Tencent · new models on Hugging FaceAI score44

    Tencent releases WeVisDoc-4B, a document parser that leads OmniDocBench v1.6

    AITencent's WeVisDoc-4B, fine-tuned from Qwen3-VL-4B-Instruct, converts page images into structured Markdown with LaTeX formulas and HTML tables. It scores 95.38 Overall on OmniDocBench v1.6 and a mean Overall of 75.54 across three PureDocBench tracks, ranking first among compared end-to-end parsers in all four reported settings. The model is available on Hugging Face and runs through vLLM, which requires version 0.11.1 or later.

  2. Tencent · new models on Hugging FaceAI score37

    Tencent Releases WeVisDoc-2B and WeVisDoc-4B Document Parsing Models on Hugging Face

    AITencent's WeVisDoc-4B, fine-tuned from Qwen3-VL-4B-Instruct, scores 95.38 Overall on OmniDocBench v1.6 and 75.54 mean Overall across three PureDocBench tracks. The end-to-end parser converts page images into structured Markdown with LaTeX formulas and HTML tables, and the 2B variant is also available. The repository provides vLLM serving scripts with a 32768-token default context and a Python client for batch processing.

  3. Google · Innovation & AIAI score44

    Google says its language tools now support over 300 languages used by 7 billion people

    AIGoogle says its technologies now support more than 300 languages spoken by 7 billion people, representing 86% of the global population. The company also released its AI & Economy ATLAS, which it describes as a look at how people are using AI globally. The post highlights recent AI science work, including AlphaGenome Atlas, WeatherNext 3, and a Planetary Prediction Engine.

  4. Baseten BlogAI score40

    LangChain uses Baseten Loops to train custom models for LangSmith Engine

    AILangChain is using Baseten Loops, a managed fine-tuning service, to train custom models for LangSmith Engine, its in-platform agent that debugs and improves AI agents. The article says LangChain fine-tunes large open-weight models on agent traces and trains smaller open-weight models such as Qwen for tasks like failure-mode categorization. Baseten Loops supports supervised fine-tuning, reinforcement learning, and long-context workloads, and lets checkpoints be evaluated and deployed directly to inference.

Sep 14

Sep 14Mon
  1. vLLM BlogAI score62

    How vLLM Speculators trained a DSpark draft model for Kimi K3 on GB300 NVL72

    AIThe vLLM team trained a DSpark speculative decoding draft model for Kimi K3, a 2.8T-parameter model, using the Speculators library on GB300 NVL72 hardware. They added a MooncakeHiddenStatesConnector to stream hidden states from disaggregated vLLM inference nodes to training nodes across multiple machines. The released speculator raises single-stream interactivity from about 110 to about 435 tokens per second per user on math reasoning, with up to about 3.5x higher output throughput under concurrent load.

    Why it matters: The post shows how hidden-state extraction and Mooncake transfers let a 2.8T-parameter model's speculator be trained across multiple nodes, a reusable pattern for similar setups.

  2. Tencent · new models on Hugging FaceAI score44

    Tencent Releases SAS Sparse-Attention Gate Checkpoints for Qwen3 Models on Hugging Face

    AITencent released Simple-Attention-Sparsification (SAS) gate checkpoints for Qwen3-4B, Qwen3-8B, and Qwen3-14B, which learn to rank and select KV blocks using continuous gates optimized with the language-modeling loss. The router-only packages, 64 MiB to 81 MiB each with 33.0M to 42.0M gate parameters, require the frozen Qwen3 base model and the seer_attn backend in a forked sglang-blocksparse build. The default sparse decode budget is 2,048 tokens, and the checkpoints can be evaluated at 1,024, 2,048, or 4,096 budgets without retraining.

Sep 10

Sep 10Thu
  1. Together AI BlogAI score52

    Together AI expands Fine-Tuning with live metrics, expert LoRA, and early stopping

    AITogether AI expanded its Fine-Tuning service with support for newer open-weight models, live metrics tracking, and finer training controls. Expert LoRA adapters can be applied to Mixture-of-Experts expert layers, and early stopping keeps the checkpoint with the best validation loss. Dataset previews, sample weights, pre-flight validation, and lower prices on selected models are also included.

  2. Google Developers BlogAI score55

    Google details autonomous LLM post-training loops using Tunix on TPUs

    AIGoogle Developers Blog describes autofinetune, a project applying autonomous agent loops to LLM post-training with Tunix, Gemma, and Cloud TPUs. In an SFT case study on FunctionGemma, an agent ran 20 automated experiments on a Cloud TPU v5e-1 to adjust LoRA settings, optimizers, and learning rates. In a GRPO case study on Gemma 3 1B for GSM8K math reasoning, the agent ran 40 experiments on a Cloud TPU v6e-1 and improved total reward by about 10%.

  3. Google LabsAI score38

    Google Labs' Dreambeans personalized daily story app now available to all U.S. accounts

    AIGoogle Labs has made Dreambeans, its experimental app that creates personalized daily story collections, available to all U.S. accounts aged 18 and over on Android and iOS. Each daily collection combines personalized topics with information distilled from connected Google apps, including Calendar, Gmail, Photos, Search, YouTube, and Gemini. Users can dive deeper into stories, bookmark them, share them, and give feedback to improve future collections.

  4. Amazon ScienceAI score55

    Research agents avoid overfitting when their winning strategies compress into few tokens

    AIAmazon Science researchers found that LLM research agents running benchmark hill-climbing rarely overfit, because their winning strategies can be compressed into prompts of about 32 tokens. A fresh reproducer agent with no access to the validation set matched the explorer's performance on most of eight datasets from that short prompt alone. The team also used the test to flag overfitting, since validation-specific gains did not survive compression.

  5. Replit BlogAI score36

    Replit and Databricks Integration Becomes Generally Available with Lakebase Support

    AIReplit's integration with Databricks is now generally available, adding native Databricks Lakebase support that lets Replit Agent automatically provision a Lakebase database when an app is ready to deploy. Apps built with Replit can read live Databricks warehouse data while storing new app data in Lakebase, inheriting existing Unity Catalog security and governance controls. The update also adds automated preview deploys that keep test data isolated from live business data.

  6. Mistral AIAI score36

    Cloudera and Mistral Partner to Deliver Sovereign AI on Enterprise Data

    AIMistral AI and Cloudera announced a partnership that integrates Mistral's models with Cloudera's hybrid data platform for enterprise AI. Customers can run inference across private and public cloud, on-prem, and fully air-gapped environments, and train custom models on proprietary data while retaining ownership. Cloudera cited 30 exabytes of customer-managed data on its platform.

Sep 9

Sep 9Wed
  1. BAAI · new models on Hugging FaceAI score24

    BAAI open-sources EPT, UniPath, and MiSI AIDD molecular and crystal modeling resources

    AIBAAI released open-source resources for three AIDD projects on Hugging Face: EPT, an equivariant pretrained transformer for unified 3D molecular representation learning, and UniPath, a learnable-time flow matching method for crystal structure and energy prediction. The repository mirrors their GitHub source code and READMEs, with setup, preprocessing, training, and evaluation documentation. The MiSI benchmark is released separately on Hugging Face.

  2. Fireworks AI BlogAI score58

    Fireworks AI outlines a staged path from closed APIs to owned specialized models

    AIFireworks AI describes a four-stage path for teams moving from renting closed frontier models to training their own, starting with API use and prompt, context, and harness engineering. The post uses the UIPad computer-use dataset to show that Kimi K3 ties GPT 5.6 Sol overall at 87.7 but wins three of four categories while costing about half as much, suggesting routing. After roughly three hours of training on the training split, the tuned Kimi K3 outperforms GPT 5.6 Sol on the held-out test set.

  3. Fireworks AI BlogAI score60

    Genspark's Gen-1 Slides matches Opus 5 decks at about one-tenth the cost per deck

    AIGenspark and Fireworks Lab post-trained the open-weight MiniMax M3 into Gen-1 Slides, a model that plans, writes, and checks slide decks end-to-end. On Genspark's evaluation it matches Claude Opus 5 at about 1/17 of its input-token list price, roughly 90% less per finished deck. In production it cut low-rated decks from 18% to 3.6% over the base model.

    Why it matters: The post explains a post-training pipeline with reward design, curriculum, and numerical fixes, showing how a cheaper model was tuned toward a frontier quality bar.

  4. Mistral AIAI score54

    Mistral details how AI agents migrated 40,000 lines of Fortran to C++

    AIMistral AI helped a European energy operator migrate 40,000 lines of Fortran 77 to C++ for a reservoir simulator with no test suite. The post explains a parity harness that checks numerical agreement between the two codebases, and a workflow where agents coder, tester, and reviewer migrate modules under human review. Its authors note the approach covered the self-contained first sprint of 40,000 of 300,000 lines and that dependent systems would bring additional challenges.

  5. Ai2 (Allen Institute for AI)AI score39

    Goodfire Traces Olmo Safety Regression to Preference Training Data

    AIGoodfire used Ai2's open post-training stack, including the Dolci preference dataset, intermediate Olmo checkpoints, and OLMES evaluations, to trace a safety regression in Olmo. Preference training made Olmo more likely to comply with harmful requests on a refusal benchmark, and Goodfire linked part of this to specific Dolci examples where the preferred response encouraged compliance. Because Ai2 publishes the individual preferred and rejected responses, researchers could test targeted changes to reduce the regression.

Sep 8

Sep 8Tue
  1. Google DeepMindAI score74

    Google DeepMind launches AlphaGenome Atlas to predict 9 billion DNA variant effects

    AIGoogle DeepMind has introduced AlphaGenome Atlas, a platform with predicted molecular effects for 9 billion single-nucleotide variants in the human genome. It is free for academic research through a web portal, and the AlphaGenome Variant Impact score condenses predictions from AlphaGenome and AlphaMissense into one number for ranking variants. The source says collaborators used it to identify variants in unsolved rare disease cases and to find rare non-coding variants linked to traits.

    Why it matters: The source details how precomputed variant predictions, a single impact score, and linked feature attributions make genome-wide mutation effects searchable for researchers without coding skills.

  2. Google DeepMind · The KeywordAI score72

    Google DeepMind launches AlphaGenome Atlas, a database of DNA variant effect predictions

    AIGoogle DeepMind has released AlphaGenome Atlas, a web portal that predicts the regulatory effects of all 9 billion possible single-letter genetic changes in the human genome. The Atlas provides an AlphaGenome Variant Impact (AVI) score that combines coding and non-coding predictions to help researchers prioritize variants. The source says the portal requires no coding skills and is available to researchers and biologists worldwide.

    Why it matters: The source details how the Atlas's AVI score is used in real rare disease and UK Biobank analyses, showing a practical route for prioritizing non-coding variants.

  3. Google DeepMind · YouTubeAI score78

    DeepMind releases AlphaGenome Atlas, a predictive map of every possible DNA letter change

    AIGoogle DeepMind has used AlphaGenome to predict the molecular impact of every possible single-letter change in the human genome, around nine billion variants. The resulting AlphaGenome Atlas is a 1PB dataset that assigns each variant an AlphaGenome Variant Impact (AVI) score, covering both coding and non-coding variations, and is available to researchers worldwide. The video notes that AlphaGenome has not been validated or approved for any clinical use.

    Why it matters: The release supplies a precomputed impact score for every possible single-letter genome change, which lets researchers look up variants without running the model themselves.

  4. NVIDIA · new models on Hugging FaceAI score46

    NVIDIA Releases NV-Reason-CT, a 3D Vision-Language Model for Chest and Abdominal CT

    AINVIDIA's NV-Reason-CT is a 3D vision-language model for CT image analysis that combines a native 3D vision encoder with a language model. It is designed for radiology report generation, question answering, and multi-step reasoning across chest and abdominal CT volumes. The model converts a 384×384×384-mm input into 13,824 visual tokens without spatial downsampling and is available on Hugging Face under the OpenMDW-1.1 License.

Sep 7

Sep 7Mon
  1. OpenBMB (MiniCPM) · new models on Hugging FaceAI score45

    openbmb/JustRL-II-base-model: RL starting checkpoint for long-CoT math reasoning

    AIOpenBMB released JustRL-II-base-model, the pre-RL starting checkpoint for the JustRL II math-reasoning case study, scoring about 61% on AIME 2025 before reinforcement learning. The full JustRL II recipe reaches 81% on AIME 2025 in about 300 RL steps from this checkpoint, versus about 74% for a standard GRPO baseline. The Llama-architecture weights are available on Hugging Face and are intended for reproducing the recipe and research on long-CoT RL, not general assistant use.