Skip to content

Areas · Latest news

Data & training

Dataset construction, synthetic data, pretraining and post-training methods, compute, and training cost.

85 top picks all-time · 51 in the past 30 days · chosen from 730 items collected all-time

Latest pick

Top picks archive · Page 5

Top picks 81–85 of 85

Feb 4

Feb 4Wed
  1. Intern Large ModelsOfficialAI score60

    Intern-S1-Pro: 1T MoE open-source multimodal scientific reasoning model released

    AIIntern Large Models introduces Intern-S1-Pro, a 1T-parameter MoE open-source multimodal scientific reasoning model with 1T-A22B active configuration. The post claims competitive scientific reasoning against leading closed-source models and supports vLLM and SGLang, with weights on Hugging Face and code on GitHub.

    Why it matters: The post pairs a 1T MoE open-source scientific reasoning model with benchmark tables against closed models, letting readers compare claimed strengths directly.

    Image from @intern_lm's post

Jan 29

Jan 29Thu
  1. Z.ai (GLM) · new models on Hugging FaceOfficialAI score60

    Z.ai releases open-source GLM-OCR multimodal document model

    AIZ.ai has released GLM-OCR, a 0.9B-parameter multimodal OCR model for complex document understanding, under the MIT License. The model scores 94.62 on OmniDocBench V1.5 and supports deployment through vLLM, SGLang, and Ollama, with an official SDK for document parsing.

    Why it matters: The page gives benchmark scores, a 0.9B parameter size, and supported serving frameworks, which help readers weigh OCR deployment options against heavier alternatives.

Jan 27

Jan 27Tue
  1. Tim DettmersBlogAI score72

    Tim Dettmers describes how SERA, an open coding agent, was built

    AITim Dettmers describes building SERA, Ai2's first Open Coding Agents release, using 32 GPUs and synthetic data. The method uses soft verification, which accepts generated patches that overlap at least 50% with the target patch, and fine-tunes a 32B model on a private codebase in about 19 GPU days. The post says the resulting model can match its teacher, GLM 4.5-Air, on that private data.

    Why it matters: The post explains how a small team built an open coding agent with cheap synthetic data and soft verification, a reusable recipe for specializing models on private code.

Oct 26, 2025

Oct 26, 2025Sun
  1. Thinking Machines LabOfficialAI score70

    Thinking Machines Lab explains on-policy distillation for cheaper LLM post-training

    AIThinking Machines Lab describes on-policy distillation, which samples rollouts from a student model and has a teacher grade each token with reverse KL. The authors report that this matches Qwen3-style reasoning results at a fraction of RL's cost, with AIME'24 reaching 70% in about 150 steps from a 400k SFT checkpoint. The method also helps recover instruction-following behavior lost during fine-tuning on internal documents.

    Why it matters: The post explains why on-policy distillation gives dense per-token feedback, letting a small model match RL results at much lower compute cost.

Nov 30, 2024

Nov 30, 2024Sat
  1. Liquid AI BlogOfficialAI score60

    Liquid AI's STAR uses evolutionary search to synthesize tailored model architectures

    AILiquid AI reports STAR, an evolutionary algorithm that synthesizes tailored neural network architectures from a numerical genome representation. The authors say it produced hundreds of designs that outperform Transformer and hybrid architectures in quality, with smaller caches and parameter counts, and can optimize for latency on target hardware. The full method is described in the arXiv technical report 2411.17800.

    Why it matters: The post explains how evolutionary search over a new architecture design space produced designs beating Transformers and hybrids, giving a concrete method for quality versus latency and memory trade-offs.