Skip to contentSkip to stories

Updated

#Data/Training

Sep 15

Sep 15Tue
  1. Tencent · new models on Hugging FaceAI score37

    Tencent Releases WeVisDoc-2B and WeVisDoc-4B Document Parsing Models on Hugging Face

    AITencent's WeVisDoc-4B, fine-tuned from Qwen3-VL-4B-Instruct, scores 95.38 Overall on OmniDocBench v1.6 and 75.54 mean Overall across three PureDocBench tracks. The end-to-end parser converts page images into structured Markdown with LaTeX formulas and HTML tables, and the 2B variant is also available. The repository provides vLLM serving scripts with a 32768-token default context and a Python client for batch processing.

  2. Lewis TunstallAI score30

    Periodic Labs advances toward cracking condensed matter physics superconductor problem

    AIPeriodic Labs, the team behind high-throughput materials labs in Menlo Park, reports progress on one of condensed matter physics' hardest problems. Its open-source model Neon, trained with mid-training and RL on 1,300 H200s plus months of lab data, surpasses GPT-6 Astra on the company's analysis benchmark. The work targets materials science challenges including superconductors, magnets, and semiconductors.

  3. Sundar PichaiAI score42

    Google outlines AI for science, weather, languages, and economic research

    AIGoogle says it is focusing AI efforts on health, disaster and weather resilience, learning, and economic opportunity. Recent examples include AlphaGenome Atlas, which maps all 9B possible single-letter genetic changes across the human genome and is openly available to researchers, and WeatherNext 3, described as its most accurate and capable global weather AI model to date. The post also cites AI & Economy ATLAS, an open-access look at global AI usage, and says its translation services now cover nearly 300 languages spoken by 7B people.

  4. Jason WeiAI score40

    Jason Wei says wet-lab data lets a specialized model beat GPT-6 Astra

    AIJason Wei argues that specialized, often private wet-lab data can let a task-specific model outperform a general frontier model on scientific tasks. He cites Neon, an open-source model that Liam Fedus says was mid-trained and RL-tuned on experimental data using 1,300 H200s to surpass GPT-6 Astra on an analysis benchmark. The post frames this data as a potential moat as work moves toward the frontier of science.

  5. Google · Innovation & AIAI score44

    Google says its language tools now support over 300 languages used by 7 billion people

    AIGoogle says its technologies now support more than 300 languages spoken by 7 billion people, representing 86% of the global population. The company also released its AI & Economy ATLAS, which it describes as a look at how people are using AI globally. The post highlights recent AI science work, including AlphaGenome Atlas, WeatherNext 3, and a Planetary Prediction Engine.

  6. Baseten BlogAI score40

    LangChain uses Baseten Loops to train custom models for LangSmith Engine

    AILangChain is using Baseten Loops, a managed fine-tuning service, to train custom models for LangSmith Engine, its in-platform agent that debugs and improves AI agents. The article says LangChain fine-tunes large open-weight models on agent traces and trains smaller open-weight models such as Qwen for tasks like failure-mode categorization. Baseten Loops supports supervised fine-tuning, reinforcement learning, and long-context workloads, and lets checkpoints be evaluated and deployed directly to inference.

  7. Tencent HunyuanAI score38

    EvolveScaler benchmarks AI on evolving world-state reasoning, frontier models struggle

    AITencent Hunyuan introduced EvolveScaler, a benchmark that builds worlds as executable state machines and renders them into natural language with 117 prototypes, 159 question operators, and five difficulty tiers. On the hardest tier, 14 frontier models' median avg@5 falls to 11.3. Training on EvolveScaler data yields a +5.25 average gain across 8 out-of-distribution benchmarks.

Sep 14

Sep 14Mon
  1. vLLM BlogAI score62

    How vLLM Speculators trained a DSpark draft model for Kimi K3 on GB300 NVL72

    AIThe vLLM team trained a DSpark speculative decoding draft model for Kimi K3, a 2.8T-parameter model, using the Speculators library on GB300 NVL72 hardware. They added a MooncakeHiddenStatesConnector to stream hidden states from disaggregated vLLM inference nodes to training nodes across multiple machines. The released speculator raises single-stream interactivity from about 110 to about 435 tokens per second per user on math reasoning, with up to about 3.5x higher output throughput under concurrent load.

    Why it matters: The post shows how hidden-state extraction and Mooncake transfers let a 2.8T-parameter model's speculator be trained across multiple nodes, a reusable pattern for similar setups.

  2. LlamaIndexAI score29

    LlamaIndex proposes two-pass just-in-time OCR for agent document pipelines

    AILlamaIndex proposes a two-pass just-in-time OCR pattern for agents working through document collections, avoiding parsing every page upfront. LiteParse, an open-source Rust tool supporting 50+ formats, performs a fast layout-aware first pass with bounding boxes, headings, tables, and a per-page complexity flag, processing a full data room in 32 seconds. LlamaParse then parses only the pages needing deeper analysis, returning cell-level tables, bounding boxes, and confidence scores.

  3. Tencent · new models on Hugging FaceAI score44

    Tencent Releases SAS Sparse-Attention Gate Checkpoints for Qwen3 Models on Hugging Face

    AITencent released Simple-Attention-Sparsification (SAS) gate checkpoints for Qwen3-4B, Qwen3-8B, and Qwen3-14B, which learn to rank and select KV blocks using continuous gates optimized with the language-modeling loss. The router-only packages, 64 MiB to 81 MiB each with 33.0M to 42.0M gate parameters, require the frozen Qwen3 base model and the seer_attn backend in a forked sglang-blocksparse build. The default sparse decode budget is 2,048 tokens, and the checkpoints can be evaluated at 1,024, 2,048, or 4,096 budgets without retraining.

Sep 13

Sep 13Sun

Sep 11

Sep 11Fri
  1. Dwarkesh PatelAI score18

    Dwarkesh Patel on why Sonnet 5 and Opus 5 trail GLM 5.3

    AIDwarkesh Patel said a discussion with John, Beren, and Charlie questioned why Sonnet 5 and Opus 5 feel weaker than GLM 5.3, even though Anthropic could use raw logit distillation from Fable and train on Fable's environments. The discussion raised questions about the value of distillation, what makes it effective, and which model behaviors are hard to extract through it.

  2. Dwarkesh PatelAI score42

    Dwarkesh Patel releases podcast with AI researchers on frontier progress

    AIDwarkesh Patel announced a new episode featuring John Schulman, Chris O'Neill, and Beren Millidge, three AI researchers from openish companies. The discussion covers the case against recursive self-improvement, drivers of Chinese labs' progress, training of automated AI researchers, long-horizon RL, the sim-to-real gap, and the role of data and RL in recent progress.

Sep 10

Sep 10Thu
  1. Together AI BlogAI score52

    Together AI expands Fine-Tuning with live metrics, expert LoRA, and early stopping

    AITogether AI expanded its Fine-Tuning service with support for newer open-weight models, live metrics tracking, and finer training controls. Expert LoRA adapters can be applied to Mixture-of-Experts expert layers, and early stopping keeps the checkpoint with the best validation loss. Dataset previews, sample weights, pre-flight validation, and lower prices on selected models are also included.

  2. Google Developers BlogAI score55

    Google details autonomous LLM post-training loops using Tunix on TPUs

    AIGoogle Developers Blog describes autofinetune, a project applying autonomous agent loops to LLM post-training with Tunix, Gemma, and Cloud TPUs. In an SFT case study on FunctionGemma, an agent ran 20 automated experiments on a Cloud TPU v5e-1 to adjust LoRA settings, optimizers, and learning rates. In a GRPO case study on Gemma 3 1B for GSM8K math reasoning, the agent ran 40 experiments on a Cloud TPU v6e-1 and improved total reward by about 10%.

  3. Sherwin WuAI score62

    OpenAI launches ChatGPT for Financial Services with GPT-6 Astra reasoning

    AIOpenAI has made ChatGPT for Financial Services available, a tailored ChatGPT Work experience that combines built-in financial data with GPT-6 Astra's reasoning. Teams can use it to develop research, build financial models, and create customized client materials. The author says it integrates financial data sources including Daloopa, PitchBook, and LSEG.

  4. Google LabsAI score38

    Google Labs' Dreambeans personalized daily story app now available to all U.S. accounts

    AIGoogle Labs has made Dreambeans, its experimental app that creates personalized daily story collections, available to all U.S. accounts aged 18 and over on Android and iOS. Each daily collection combines personalized topics with information distilled from connected Google apps, including Calendar, Gmail, Photos, Search, YouTube, and Gemini. Users can dive deeper into stories, bookmark them, share them, and give feedback to improve future collections.

  5. John SchulmanAI score40

    Schulman says user data gains in math are unlikely; disclosure norms needed

    AIJohn Schulman argues that training on user data contributes little to frontier math gains, which come mainly from scaling pretraining and RLVR. He says user data is more likely used to find failure modes that hired annotators struggle to recreate. He calls for stronger norms on disclosing how companies train on user data, including the methods and capabilities targeted.

  6. Amazon ScienceAI score55

    Research agents avoid overfitting when their winning strategies compress into few tokens

    AIAmazon Science researchers found that LLM research agents running benchmark hill-climbing rarely overfit, because their winning strategies can be compressed into prompts of about 32 tokens. A fresh reproducer agent with no access to the validation set matched the explorer's performance on most of eight datasets from that short prompt alone. The team also used the test to flag overfitting, since validation-specific gains did not survive compression.

  7. Replit BlogAI score36

    Replit and Databricks Integration Becomes Generally Available with Lakebase Support

    AIReplit's integration with Databricks is now generally available, adding native Databricks Lakebase support that lets Replit Agent automatically provision a Lakebase database when an app is ready to deploy. Apps built with Replit can read live Databricks warehouse data while storing new app data in Lakebase, inheriting existing Unity Catalog security and governance controls. The update also adds automated preview deploys that keep test data isolated from live business data.

  8. Mistral AIAI score36

    Cloudera and Mistral Partner to Deliver Sovereign AI on Enterprise Data

    AIMistral AI and Cloudera announced a partnership that integrates Mistral's models with Cloudera's hybrid data platform for enterprise AI. Customers can run inference across private and public cloud, on-prem, and fully air-gapped environments, and train custom models on proprietary data while retaining ownership. Cloudera cited 30 exabytes of customer-managed data on its platform.

  9. Tencent HunyuanAI score60

    Tencent Hunyuan releases open-source AuK audio model for speech generation and editing

    AITencent Hunyuan has released AuK, an open-source foundation model for unified speech generation and editing that takes natural-language instructions and reference audio. It supports tasks including zero-shot TTS, timbre, style and emotion editing, denoising, and music separation. A companion AuK-Flash variant runs 4-step inference and is about 4.5 times faster under matched conditions, with code, weights, and a demo now available.

Sep 9

Sep 9Wed
  1. BAAI · new models on Hugging FaceAI score24

    BAAI open-sources EPT, UniPath, and MiSI AIDD molecular and crystal modeling resources

    AIBAAI released open-source resources for three AIDD projects on Hugging Face: EPT, an equivariant pretrained transformer for unified 3D molecular representation learning, and UniPath, a learnable-time flow matching method for crystal structure and energy prediction. The repository mirrors their GitHub source code and READMEs, with setup, preprocessing, training, and evaluation documentation. The MiSI benchmark is released separately on Hugging Face.

  2. Fireworks AI BlogAI score58

    Fireworks AI outlines a staged path from closed APIs to owned specialized models

    AIFireworks AI describes a four-stage path for teams moving from renting closed frontier models to training their own, starting with API use and prompt, context, and harness engineering. The post uses the UIPad computer-use dataset to show that Kimi K3 ties GPT 5.6 Sol overall at 87.7 but wins three of four categories while costing about half as much, suggesting routing. After roughly three hours of training on the training split, the tuned Kimi K3 outperforms GPT 5.6 Sol on the held-out test set.

  3. Fireworks AI BlogAI score60

    Genspark's Gen-1 Slides matches Opus 5 decks at about one-tenth the cost per deck

    AIGenspark and Fireworks Lab post-trained the open-weight MiniMax M3 into Gen-1 Slides, a model that plans, writes, and checks slide decks end-to-end. On Genspark's evaluation it matches Claude Opus 5 at about 1/17 of its input-token list price, roughly 90% less per finished deck. In production it cut low-rated decks from 18% to 3.6% over the base model.

    Why it matters: The post explains a post-training pipeline with reward design, curriculum, and numerical fixes, showing how a cheaper model was tuned toward a frontier quality bar.

  4. LlamaIndexAI score23

    LlamaParse now available as a ChatGPT connector for document parsing

    AILlamaIndex has made LlamaParse available in the ChatGPT plugin directory, following its earlier Claude integration. The connector parses scanned, table-heavy, and chart-filled documents into Markdown, JSON, or HTML, extracts fields into a user-defined schema, searches document collections, and classifies and splits files into sections.

  5. Mistral AIAI score54

    Mistral details how AI agents migrated 40,000 lines of Fortran to C++

    AIMistral AI helped a European energy operator migrate 40,000 lines of Fortran 77 to C++ for a reservoir simulator with no test suite. The post explains a parity harness that checks numerical agreement between the two codebases, and a workflow where agents coder, tester, and reviewer migrate modules under human review. Its authors note the approach covered the self-contained first sprint of 40,000 of 300,000 lines and that dependent systems would bring additional challenges.