Skip to contentSkip to stories

Updated

#Data/Training

Showing low-relevance items too. Hide low-relevance items

Sep 17

Sep 17Thu
  1. Google · AI blogOfficialAI score38

    UN System Data Commons unifies global statistics into an AI-ready open platform

    AIThe United Nations system launched UN System Data Commons, an open-source platform built on Data Commons by Google that integrates siloed global statistics into one AI-ready knowledge graph. Users can query it in natural language, browse by location or theme, and use MCP-enabled AI agents to fetch verified figures and draft charts or reports. The UN plans to add more datasets, aiming to include 80% of UN system statistical datasets by 2027.

  2. LlamaIndex 🦙OfficialAI score13

    LlamaIndex's Jerry Liu on document parsing challenges for enterprise agents

    AILlamaIndex CEO Jerry Liu spoke at Connected Stack's Founder Flash Talks about messy, complex documents that general-purpose models struggle to read, a problem enterprise agents eventually face. The company says it is building document infrastructure for agents, which it describes as the new knowledge workers.

    Image from @llama_index's post
  3. Ali GhodsiXAI score44

    Databricks CEO Ali Ghodsi shares 10 leadership lessons from interview

    AIDatabricks CEO Ali Ghodsi discussed leadership in an interview with @bhalligan, and the main post calls it a fun and very different conversation. The quoted background notes that Ghodsi never wanted to be CEO and was handed the interim title in 2015, when revenue was $1.5M. The source's listed takeaways emphasize focusing on the biggest bottleneck, embracing conflict, and studying competitors' weaknesses.

  4. Google DeepMindOfficialAI score47

    AlphaGenome Atlas boosts rare genetic signal detection in UK Biobank study

    AIResearchers at the University of Exeter, working with Google DeepMind, used AlphaGenome Atlas on data from more than 54,000 UK Biobank participants. The approach raised detection of rare genetic signals by over 22% and uncovered new DNA variants that influence levels of PLA2G7, a protein linked to metabolic health.

    Image from @GoogleDeepMind's post
  5. Google DeepMindOfficialAI score33

    Researchers use AlphaGenome Atlas to identify disease-causing DNA variants

    AIResearchers at the Broad Institute, the University of Exeter, and other institutions are already using AlphaGenome Atlas to identify potential disease-causing DNA variants and interpret their role. The post is a thread announcement, and it gives no further details on methods or results.

    Video from @GoogleDeepMind's post
  6. OpenBMBOfficialAI score29

    Kahya-TTS: Turkish speech model fine-tuned from VoxCPM2 on 100 hours

    AIDeveloper Alican Kiraz fine-tuned OpenBMB's open-source VoxCPM2 voice model on nearly 100 hours of natural Turkish speech, creating Kahya-TTS for Turkish text-to-speech. The project shows how open-source voice models can be adapted to new languages and specialized datasets. The model is available on Hugging Face.

    Image from @OpenBMB's post
  7. JetBrains AI BlogOfficialAI score44

    Building a RAG Pipeline for Semantic Code Search: A Developer Diary

    AIJetBrains describes building Air Context, a RAG pipeline that gives LLM agents semantic code search over real repositories instead of grep. The first installment covers parsing, chunking, and vectorization, arguing that fixed-size line chunks split related code and that structure-aware chunking using language grammar produces better retrieval units.

  8. Ai2 (Allen Institute for AI)OfficialAI score42

    Crowdsourced Game Steering Arena Shows Olmo 3 Prosocial Scores Can Be Gamed

    AINortheastern University MS student Soham Padia used Ai2's open Olmo 3-32B model to build Steering Arena, a public game in which players submit text prefixes to steer prosocial behavior. About 600 submissions from a few dozen people showed the top 36 entries were unreadable token strings, while the best plain-English entry ranked 37th at about 2.7 times lower score. The results suggest that once an evaluation metric is exposed, it becomes an optimization target.

Sep 16

Sep 16Wed
  1. TinkerOfficialAI score32

    Sundial trains Inkling-Small to fix LaTeX errors in under a second

    AISundial fine-tuned Thinking Machines' Inkling-Small with RLVR on 3,978 verified TeX.StackExchange fixes, using rewards for compilation and PDF match and penalties for removed content. The trained model fixes 83.7% of LaTeX errors in under one second at $0.0013 per fix, according to the post. Sundial says it is rolling out the model in its editor, applying fixes as suggestions and rebuilding the PDF.

  2. LlamaIndex 🦙OfficialAI score14

    LlamaIndex webinar on insurance document pipelines with LlamaParse and Extract

    AILlamaIndex Solutions Architect Abrar Mahi will host a webinar on turning insurance documents such as accord forms and policy documents into structured data for underwriting, policy review, and claims. The session covers extracting policy, property, and claims history into a defined schema, verifying values with citations and bounding boxes, and using confidence scores with validation rules to route items to human review.

    Image from @llama_index's post

Sep 15

Sep 15Tue
  1. Tencent · new models on Hugging FaceOfficialAI score44

    Tencent releases WeVisDoc-4B, a document parser that leads OmniDocBench v1.6

    AITencent's WeVisDoc-4B, fine-tuned from Qwen3-VL-4B-Instruct, converts page images into structured Markdown with LaTeX formulas and HTML tables. It scores 95.38 Overall on OmniDocBench v1.6 and a mean Overall of 75.54 across three PureDocBench tracks, ranking first among compared end-to-end parsers in all four reported settings. The model is available on Hugging Face and runs through vLLM, which requires version 0.11.1 or later.

  2. Tencent · new models on Hugging FaceOfficialAI score37

    Tencent Releases WeVisDoc-2B and WeVisDoc-4B Document Parsing Models on Hugging Face

    AITencent's WeVisDoc-4B, fine-tuned from Qwen3-VL-4B-Instruct, scores 95.38 Overall on OmniDocBench v1.6 and 75.54 mean Overall across three PureDocBench tracks. The end-to-end parser converts page images into structured Markdown with LaTeX formulas and HTML tables, and the 2B variant is also available. The repository provides vLLM serving scripts with a 32768-token default context and a Python client for batch processing.

  3. Stability AIOfficialAI score25

    Stability AI's team presents color consistency research at ECCV

    AIStability AI's interactive research team presented new work at the 19th European Conference on Computer Vision (ECCV) aimed at keeping colors consistent across shots and AI-generated reference photos. The post frames color consistency as a persistent friction point in production.

    Image from @StabilityAI's post
  4. Lewis Tunstall @ COLM 🌉XAI score30

    Periodic Labs advances toward cracking condensed matter physics superconductor problem

    AIPeriodic Labs, the team behind high-throughput materials labs in Menlo Park, reports progress on one of condensed matter physics' hardest problems. Its open-source model Neon, trained with mid-training and RL on 1,300 H200s plus months of lab data, surpasses GPT-6 Astra on the company's analysis benchmark. The work targets materials science challenges including superconductors, magnets, and semiconductors.

  5. Sundar PichaiXAI score42

    Google outlines AI for science, weather, languages, and economic research

    AIGoogle says it is focusing AI efforts on health, disaster and weather resilience, learning, and economic opportunity. Recent examples include AlphaGenome Atlas, which maps all 9B possible single-letter genetic changes across the human genome and is openly available to researchers, and WeatherNext 3, described as its most accurate and capable global weather AI model to date. The post also cites AI & Economy ATLAS, an open-access look at global AI usage, and says its translation services now cover nearly 300 languages spoken by 7B people.

    Image from @sundarpichai's post
  6. Jason WeiXAI score40

    Jason Wei says wet-lab data lets a specialized model beat GPT-6 Astra

    AIJason Wei argues that specialized, often private wet-lab data can let a task-specific model outperform a general frontier model on scientific tasks. He cites Neon, an open-source model that Liam Fedus says was mid-trained and RL-tuned on experimental data using 1,300 H200s to surpass GPT-6 Astra on an analysis benchmark. The post frames this data as a potential moat as work moves toward the frontier of science.

  7. Dwarkesh PatelXAI score38

    Dwarkesh Patel on Materials Synthesis Search and Depth-First Experimentation

    AIDwarkesh Patel reports that materials synthesis experiment spaces are very wide but amenable to depth-first search, where each next experiment becomes better designed and more informative as data accumulates. The post is a lab visit reaction and does not name specific models, figures, or results.

  8. Google · Innovation & AIOfficialAI score44

    Google says its language tools now support over 300 languages used by 7 billion people

    AIGoogle says its technologies now support more than 300 languages spoken by 7 billion people, representing 86% of the global population. The company also released its AI & Economy ATLAS, which it describes as a look at how people are using AI globally. The post highlights recent AI science work, including AlphaGenome Atlas, WeatherNext 3, and a Planetary Prediction Engine.

  9. Baseten BlogOfficialAI score40

    LangChain uses Baseten Loops to train custom models for LangSmith Engine

    AILangChain is using Baseten Loops, a managed fine-tuning service, to train custom models for LangSmith Engine, its in-platform agent that debugs and improves AI agents. The article says LangChain fine-tunes large open-weight models on agent traces and trains smaller open-weight models such as Qwen for tasks like failure-mode categorization. Baseten Loops supports supervised fine-tuning, reinforcement learning, and long-context workloads, and lets checkpoints be evaluated and deployed directly to inference.

  10. Tencent HyOfficialAI score38

    EvolveScaler benchmarks AI on evolving world-state reasoning, frontier models struggle

    AITencent Hunyuan introduced EvolveScaler, a benchmark that builds worlds as executable state machines and renders them into natural language with 117 prototypes, 159 question operators, and five difficulty tiers. On the hardest tier, 14 frontier models' median avg@5 falls to 11.3. Training on EvolveScaler data yields a +5.25 average gain across 8 out-of-distribution benchmarks.

    Image from @TencentHunyuan's post

Sep 14

Sep 14Mon
  1. vLLM BlogOfficialAI score62

    How vLLM Speculators trained a DSpark draft model for Kimi K3 on GB300 NVL72

    AIThe vLLM team trained a DSpark speculative decoding draft model for Kimi K3, a 2.8T-parameter model, using the Speculators library on GB300 NVL72 hardware. They added a MooncakeHiddenStatesConnector to stream hidden states from disaggregated vLLM inference nodes to training nodes across multiple machines. The released speculator raises single-stream interactivity from about 110 to about 435 tokens per second per user on math reasoning, with up to about 3.5x higher output throughput under concurrent load.

    Why it matters: The post shows how hidden-state extraction and Mooncake transfers let a 2.8T-parameter model's speculator be trained across multiple nodes, a reusable pattern for similar setups.

  2. TinkerOfficialAI score44

    RLVR trains models to design power transformers with physics-based verifiers

    AITinker says engineers are using physics-based verifiers and Tinker to train models that design power transformers meeting specifications at low cost. Background from @gentrajectory says an RL-trained Kimi base model met 93% of unseen transformer specs, compressing multi-week engineering work into minutes of inference.

  3. LlamaIndex 🦙OfficialAI score29

    LlamaIndex proposes two-pass just-in-time OCR for agent document pipelines

    AILlamaIndex proposes a two-pass just-in-time OCR pattern for agents working through document collections, avoiding parsing every page upfront. LiteParse, an open-source Rust tool supporting 50+ formats, performs a fast layout-aware first pass with bounding boxes, headings, tables, and a per-page complexity flag, processing a full data room in 32 seconds. LlamaParse then parses only the pages needing deeper analysis, returning cell-level tables, bounding boxes, and confidence scores.

    Image from @llama_index's post
  4. Intern Large ModelsOfficialAI score62

    Intern-S2-397B released in BF16 and FP8 under Apache 2.0

    AIShanghai AI Laboratory's Intern Large Models announced Intern-S2-397B, available in BF16 and FP8 under Apache 2.0. The post reports 87.0 on FrontierScience-Olympiad and 84.0 on SWE-bench Multilingual, leading the reported comparison on both, and says it was jointly trained across 20+ scientific domains with long-horizon agent RL.

    Why it matters: The post names the benchmark scores and training scope behind Intern-S2-397B, letting readers compare its scientific and agentic claims against the table.

  5. Tencent · new models on Hugging FaceOfficialAI score44

    Tencent Releases SAS Sparse-Attention Gate Checkpoints for Qwen3 Models on Hugging Face

    AITencent released Simple-Attention-Sparsification (SAS) gate checkpoints for Qwen3-4B, Qwen3-8B, and Qwen3-14B, which learn to rank and select KV blocks using continuous gates optimized with the language-modeling loss. The router-only packages, 64 MiB to 81 MiB each with 33.0M to 42.0M gate parameters, require the frozen Qwen3 base model and the seer_attn backend in a forked sglang-blocksparse build. The default sparse decode budget is 2,048 tokens, and the checkpoints can be evaluated at 1,024, 2,048, or 4,096 budgets without retraining.

Sep 13

Sep 13Sun
  1. Ian Johnson 🔬🤖XAI score34

    Flying through 30 million embeddings as a video game

    AIIan Johnson visualizes 30 million jina-v5-nano embeddings from 12 billion tokens across multilingual FineWeb, StarCoder, The Pile, and RedPajama as a flyable video game. He frames the project as making data exploration engaging rather than a chore.

    Video from @enjalot's post

Sep 11

Sep 11Fri
  1. Dwarkesh PatelXAI score18

    Dwarkesh Patel on why Sonnet 5 and Opus 5 trail GLM 5.3

    AIDwarkesh Patel said a discussion with John, Beren, and Charlie questioned why Sonnet 5 and Opus 5 feel weaker than GLM 5.3, even though Anthropic could use raw logit distillation from Fable and train on Fable's environments. The discussion raised questions about the value of distillation, what makes it effective, and which model behaviors are hard to extract through it.

    Video from @dwarkesh_sp's post
  2. Dwarkesh PatelXAI score42

    Dwarkesh Patel releases podcast with AI researchers on frontier progress

    AIDwarkesh Patel announced a new episode featuring John Schulman, Chris O'Neill, and Beren Millidge, three AI researchers from openish companies. The discussion covers the case against recursive self-improvement, drivers of Chinese labs' progress, training of automated AI researchers, long-horizon RL, the sim-to-real gap, and the role of data and RL in recent progress.

    Video from @dwarkesh_sp's post

Sep 10

Sep 10Thu
  1. Together AI BlogOfficialAI score52

    Together AI expands Fine-Tuning with live metrics, expert LoRA, and early stopping

    AITogether AI expanded its Fine-Tuning service with support for newer open-weight models, live metrics tracking, and finer training controls. Expert LoRA adapters can be applied to Mixture-of-Experts expert layers, and early stopping keeps the checkpoint with the best validation loss. Dataset previews, sample weights, pre-flight validation, and lower prices on selected models are also included.

  2. Google Developers BlogOfficialAI score55

    Google details autonomous LLM post-training loops using Tunix on TPUs

    AIGoogle Developers Blog describes autofinetune, a project applying autonomous agent loops to LLM post-training with Tunix, Gemma, and Cloud TPUs. In an SFT case study on FunctionGemma, an agent ran 20 automated experiments on a Cloud TPU v5e-1 to adjust LoRA settings, optimizers, and learning rates. In a GRPO case study on Gemma 3 1B for GSM8K math reasoning, the agent ran 40 experiments on a Cloud TPU v6e-1 and improved total reward by about 10%.

  3. Amazon ScienceOfficialAI score40

    Amazon research explains why ML research agents don't overfit benchmarks

    AIAmazon Science researchers propose that machine learning research agents avoid overfitting benchmarks despite years of iteration against the same tests. They attribute this to generalizable strategies being expressed compactly, leaving no room for memorization, while overfitting strategies fail to survive a compression bottleneck.

  4. Sebastian RaschkaXAI score19

    Single-GPU mixture-of-experts LLM trained from scratch in 8 days

    AIGiles Thomas extended the GPT-2-style code from Sebastian Raschka's "Build a Large Language Model (from Scratch)" into a 6-expert, 2-active mixture-of-experts model and trained it from scratch over 8 days. Raschka praised the project as interesting LLM work done on a single GPU.