Skip to contentSkip to stories

Updated

#Data/Training

Jul 29

Jul 29Wed
  1. Fireworks AI BlogAI score54

    Fireworks tests whether LoRA or full fine-tuning gaps come from data, learning rate, or rank

    AIFireworks AI ran controlled SFT experiments on Qwen3.5-9B comparing LoRA with full parameter fine-tuning across three synthetic verifiable tasks. The post argues that a FullFT advantage can come from data coverage, learning-rate tuning, or adapter rank, and it recommends testing these in that order before switching methods. Under a fixed multi-task budget, FullFT kept a 4.29-point lead over the best LoRA recipe tested, while matched data exposure favored LoRA.

  2. Liquid AI NewsletterAI score46

    Liquid AI Expands LFM2 Tokenizer to 128K, Speeding On-Device Thai, Vietnamese, and Hindi

    AILiquid AI doubled the LFM2 tokenizer's vocabulary from 65K to 128K without retraining from scratch, extending the original BPE merges and initializing new embeddings as the mean of their sub-tokens. The expanded tokenizer needs 4.0× fewer tokens for Thai, 2.6× fewer for Vietnamese, and 2.4× fewer for Hindi, which the source says yields roughly 2.2–3.7× faster on-device decoding for these languages with no reported quality loss on previously supported languages. LFM2.5-8B-A1B and the expanded tokenizer are available on Hugging Face with open weights.

  3. Berkeley AI ResearchAI score44

    K-Search Adapts CUDA Kernel Expertise to Apple Silicon MLX Backend

    AIBerkeley AI Research extended the K-Search evolutionary kernel framework with an MLX backend and a CUDA-to-MLX translation layer, letting it adapt existing CUDA kernels for Apple Silicon. The team reports a 0.97x speedup relative to the native MLX Attention kernel and up to a 20x prefill speedup over the community mlx-lm implementation on the Mamba SSM kernel. The method uses Gemini 3.5 Pro Preview to both reason about optimizations and write candidate kernels.

Jul 28

Jul 28Tue
  1. Fireworks AI BlogAI score46

    Fireworks AI Shows Low-Cost Fine-Tuning Lifts Domain Embedding Retrieval

    AIFireworks AI describes fine-tuning Qwen3-Embedding-8B on private (query, positive) pairs using bidirectional InfoNCE loss through its Training SDK, then serving the model via an OpenAI-compatible embeddings endpoint. The post reports that around 150 training steps was enough, that rank-32 LoRA landed within about one point of full-parameter fine-tuning, and that gains were largest where the base model struggled, while tasks like CoSQA and FiQA2018 showed flat results.

Jul 26

Jul 26Sun
  1. Berkeley AI ResearchAI score44

    Berkeley AI Research Trains LLMs to Update Beliefs for Long Tasks

    AIBerkeley AI Research introduces ABBEL, a framework that replaces full interaction histories with natural-language belief states that models update as new observations arrive. On CollabBench collaborative coding, belief grading closes about half the performance gap to full-context models while using fewer peak tokens and training in 50 steps instead of 100.

Jul 25

Jul 25Sat

Jul 21

Jul 21Tue
  1. Meta AI BlogAI score44

    Meta's SAM 3 and DINOv3 Power SYNAPS-I's Genesis Mission Imaging Pipeline

    AISYNAPS-I, a multi-lab Genesis Mission project led by Lawrence Berkeley National Laboratory, uses Meta's open-source SAM 3 and DINOv3 models to segment X-ray and micro-CT scientific imagery. The fine-tuned pipeline, run on 300 A100 GPUs, reduced a grapevine xylem analysis from a month of expert annotation per time step to about 15 minutes. The team can deploy the open models inside secure national lab infrastructure, where research data must remain.

Jul 15

Jul 15Wed
  1. Liquid AI NewsletterAI score38

    Liquid AI Releases Antidoom and IFStruct to Fix Reasoning Loops and Schema Errors

    AILiquid AI released Antidoom, an open-source method that retrains a single overtrained token to eliminate "doom loops" in small reasoning models. On LFM2.5-2.6B and Qwen3.5-4B, loop rates fell from 10.2% to 1.4% and from 22.9% to 1%, respectively. The company also released IFStruct, an open-source benchmark measuring whether model outputs satisfy a schema, where LFM2.5-350M rose from 21.10% to 44.90% after training.

Jul 12

Jul 12Sun
  1. ByteDance · new models on Hugging FaceAI score41

    ByteDance releases UniVR-34B-Planning for visual-space reasoning and planning

    AIByteDance's UniVR-34B-Planning, built on Emu3.5 at 34B parameters, learns visual reasoning, physical dynamics, and long-term planning from visual demonstrations using a next-token objective and two-stage training on the VR-X dataset with VR-GRPO reinforcement learning. On the VR-X benchmark it scores 58.2 overall, up 18.4 points from the Emu3.5 34B baseline of 39.8. The Planning checkpoint is available on Hugging Face under CC BY 4.0, alongside a General checkpoint.

Jul 9

Jul 9Thu
  1. Thinking Machines LabAI score44

    Thinking Machines Argues the Future Worth Building Keeps Humans Central to AI Decisions

    AIThinking Machines Lab says AI should extend human will and judgment, with people shaping its goals through continuous feedback rather than relying on models trained once and frozen. The company outlines three technical directions: training strong models, building tools for customization including training model weights, and developing interfaces that let personal judgment influence AI work. It also says it will publish research for the scientific community.

Jul 8

Jul 8Wed
  1. Cognition Blog (Devin, Windsurf)AI score62

    Cognition releases SWE-1.7, a coding model trained with long-horizon RL

    AICognition launched SWE-1.7, which it says reaches frontier-level coding performance at lower cost, trained from a Kimi K2.7 base. The post describes RL methods including top-p sampling replay to preserve entropy, compressed weight deltas across multi-cluster training, and self-compaction for rollouts up to six hours. SWE-1.7 is available in Devin via Cerebras at 1000 TPS.

    Why it matters: The post details entropy preservation, multi-cluster weight sync, and self-compaction, offering concrete RL training techniques for long-horizon coding agents to compare against one's own pipeline.

Jul 7

Jul 7Tue
  1. Berkeley AI ResearchAI score62

    Berkeley researchers outline how data systems must change as agents take over knowledge work

    AIBerkeley AI Research authors argue that near-free inference will make agents the dominant workload for data systems, requiring redesign for agentic speculation, agent-run state and coordination, and agent-synthesized systems. The post cites inference prices falling 9x to 900x per year with a median near 50x, and reports that about 80-90% of sub-queries in a text-to-SQL benchmark were duplicates. It frames the three directions as data systems for, of, and by agents.

    Why it matters: The piece maps three concrete data-system challenges posed by near-free inference, useful for anyone designing infrastructure for agent workloads and memory.

Jun 29

Jun 29Mon
  1. Meta AI BlogAI score68

    Meta's Brain2Qwerty v2 decodes sentences from non-invasive brain recordings

    AIMeta released Brain2Qwerty v2, an end-to-end deep learning pipeline that decodes sentences in real time from non-invasive brain recordings. The model reached 61% word accuracy across participants, compared with 8% for other non-invasive methods, and 78% for the best participant. Meta also released the v1 and v2 training code, and partner BCBL released the v1 dataset.

    Why it matters: The source reports word accuracy and data-scaling results for non-invasive decoding, offering a benchmark against surgical brain-computer interfaces and prior non-invasive methods.

Jun 26

Jun 26Fri
  1. Qwen · new models on Hugging FaceAI score44

    Qwen3-ForcedAligner-0.6B-hf Adds Timestamp Alignment for Speech Transcripts

    AIQwen released Qwen3-ForcedAligner-0.6B-hf, a Transformers-format forced aligner that predicts timestamps for arbitrary units within up to 5 minutes of speech in 11 languages. The model accepts transcripts from any ASR system, and the documentation shows it paired with Qwen3-ASR-0.6B and NVIDIA Parakeet CTC. Until it ships in an official Transformers release, users must install Transformers from source.

Jun 18

Jun 18Thu
  1. OpenAI Alignment Research BlogAI score62

    OpenAI study finds beneficial-trait RL improves alignment across untrained domains

    AIOpenAI reports that reinforcement learning on realistic conversations targeting traits such as honesty, epistemic humility, and corrigibility improved a model across 44 out-of-distribution alignment evaluations. Gains included reward hacking, deception, and health benchmarks, and training only on health conversations still improved non-health alignment scores. The trained model was also harder to steer toward harmful behavior with adversarial persona prompts or harmful fine-tuning.

    Why it matters: The post tests whether reinforcement learning on beneficial traits in one domain transfers to unrelated alignment benchmarks and holds up under adversarial steering.

Jun 16

Jun 16Tue
  1. OpenAI Alignment Research BlogAI score60

    WildChat-based simulation predicts OpenAI production misalignment rates within roughly 3x

    AIOpenAI's alignment team found that re-generating 100,000 WildChat conversations with five recent OpenAI models predicted production failure rates across four orders of magnitude, with 95% of predictions within 1.04 orders of magnitude. The approach was weaker for agentic misalignment categories, where errors were about 37 times larger, and it still held roughly without access to chain-of-thought reasoning, with mean multiplicative error rising from 3.6x to 4.0x.

    Why it matters: The post tests whether public chat data can predict real production failure rates, and where that prediction breaks down for agentic behavior.

Jun 10

Jun 10Wed
  1. ByteDance · new models on Hugging FaceAI score52

    ByteDance open-sources Bernini-Diffusers for semantic video generation and editing

    AIByteDance open-sourced inference code and model weights for Bernini-Diffusers, a full video generation and editing pipeline with an MLLM-based semantic planner and a DiT-based renderer. The release bundles a Qwen2.5-VL planner and Wan2.2 diffusion components in one self-contained directory, and the source recommends it over the renderer-only Bernini-R for complex instruction following.

May 25

May 25Mon
  1. MiniMax BlogAI score67

    MiniMax Explains Why Its Model Failed to Output Certain Rare Chinese Tokens

    AIMiniMax says the M2 series could not generate the rare token "嘉祺" in names like Ma Jiaqi, and its investigation traced the cause to post-training. The company found the token was learned in pretraining, but low coverage of rare tokens in post-training data caused lm_head vectors to drift. Adding synthetic full-vocabulary repetition data restored generation for these tokens and reduced Japanese-to-Russian mixing from 47% to 1%.

    Why it matters: The post traces a specific token failure through tokenizer, embedding, and lm_head checks, showing a reusable way to diagnose post-training generation problems.

May 20

May 20Wed
  1. Stability AIAI score62

    Stability AI releases Stable Audio 3.0 model family with open-weight music models

    AIStability AI released Stable Audio 3.0, a family of four audio models trained on fully licensed data. Three of them, Small SFX, Small and Medium, have open weights on Hugging Face, while Large is available through the Stability AI API and enterprise self-hosting. Outputs can be distributed and commercialized under the Stability AI Community License, and organizations with more than $1M in annual revenue can use the Enterprise License.

    Why it matters: The source specifies which models are open-weight, their licensing terms, and clip-length limits, which matters for anyone deciding whether to build on them.

Apr 23

Apr 23Thu
  1. Apple · new models on Hugging FaceAI score40

    Apple releases CADD-Base-7B, a masked diffusion model for code generation

    AIApple has released CADD-Base-7B on Hugging Face, a 7B masked diffusion language model for code generation that uses Continuously Augmented Discrete Diffusion (CADD) to guide discrete denoising with a continuous flow-matching signal. The model loads through Transformers with trust_remote_code, and its diffusion_generate method supports CADD sampling modes "weighted" and "argmax" with alg options such as "entropy" and "maskgit_plus". The release builds on DiffuCoder and reuses Dream's modeling architecture and generation utilities.

Apr 20

Apr 20Mon
  1. Berkeley AI ResearchAI score44

    GRASP: A Gradient-Based Planner for Long-Horizon World Model Planning

    AIBerkeley AI Research introduces GRASP, a gradient-based planner for learned world models that aims to make long-horizon planning more robust. GRASP lifts trajectories into virtual states for parallel optimization across time, adds stochasticity to state iterates for exploration, and reshapes gradients to avoid brittle state-input gradients through high-dimensional vision models. The post identifies ill-conditioned gradients and non-greedy loss landscapes as core failure modes of standard rollout-based planning.

Apr 17

Apr 17Fri
  1. OpenAI · new models on Hugging FaceAI score41

    OpenAI Releases Privacy Filter, an Open-Weight PII Detection Model on Hugging Face

    AIOpenAI released Privacy Filter, a bidirectional token-classification model that detects and masks personally identifiable information in text under the Apache 2.0 license. The model has 1.5B total parameters with 50M active, supports a 128,000-token context window, and can run in a web browser or on a laptop. Users can fine-tune it and adjust precision/recall tradeoffs through preset operating points.

Apr 13

Apr 13Mon
  1. ARC PrizeAI score58

    ARC Prize Releases Human Performance Dataset for ARC-AGI-3 Benchmark

    AIARC Prize Foundation released an open-source human dataset for ARC-AGI-3, covering 342 step-by-step replays across 25 public environments from a study of 458 participants. The source reports that every environment was solved by at least two humans, and it updates scoring by moving the per-level baseline to the median human player and raising the per-level cap from 100% to 115%.

  2. Cognition Blog (Devin, Windsurf)AI score62

    Cognition introduces SWE-check, a fast RL-trained bug detection model for Windsurf

    AICognition and Applied Compute RL-trained SWE-check, a specialized bug detection model for the Windsurf IDE. It matches frontier performance on in-distribution evals and is an order of magnitude faster with cheaper inference, though it trails frontier models on out-of-distribution evals (delta F1 0.29 versus 0.49 before training). A preview is available in Windsurf Next, with a mainstream release planned.

    Why it matters: The post explains how production environment replication, reward linearization, and two-phase post-training trade bug-detection quality against latency for an IDE specialist model.

Mar 22

Mar 22Sun
  1. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceAI score32

    PrismAudio Adds Reinforcement Learning to Video-to-Audio Generation with Chain-of-Thought Planning

    AIPrismAudio is a framework that integrates reinforcement learning into video-to-audio generation, using a Chain-of-Thought planning mechanism. It builds on ThinkSound by splitting single-step reasoning into four CoT modules for semantic, temporal, aesthetic, and spatial dimensions, each with targeted reward functions. Code, model weights, and datasets are released for research and educational use under the MIT License, and commercial use requires explicit author authorization.

Mar 17

Mar 17Tue
  1. Apple · new models on Hugging FaceAI score43

    Apple releases SimpleSD-4B-thinking, a self-distilled Qwen model for code generation

    AIApple has published SimpleSD-4B-thinking on Hugging Face, a research checkpoint built on Qwen that improves code generation through Simple Self-Distillation without rewards, verifiers, teacher models, or reinforcement learning. On LiveCodeBench, it lifts Qwen3-4B-Thinking-2507 from 54.5% to 57.8% pass@1 on LCBv6 and from 59.6% to 63.1% pass@1 on LCBv5. The model is released as a reproducibility checkpoint under the Apple Machine Learning Research Model License, not as an optimized Qwen release.

  2. Apple · new models on Hugging FaceAI score46

    Apple releases SimpleSD-4B-instruct, a self-distilled Qwen code model

    AIApple has released SimpleSD-4B-instruct on Hugging Face, a research checkpoint fine-tuned from Qwen3-4B-Instruct-2507 on its own sampled outputs to improve code generation. On LiveCodeBench, the model scores 41.5% pass@1 on LCBv6, up from the base model's 34.0%, and 45.7% pass@1 on LCBv5, up from 34.3%. The model is released under the Apple Machine Learning Research Model License and is intended for reproducibility rather than as an optimized Qwen release.

Mar 13

Mar 13Fri
  1. Berkeley AI ResearchAI score34

    SPEX and ProxySPEX Identify Influential LLM Interactions at Scale with Fewer Ablations

    AIBerkeley AI Research introduces SPEX, a signal-processing framework that identifies influential interactions in LLMs using far fewer ablations than exhaustive analysis. A hierarchy-based extension, ProxySPEX, matches SPEX performance with around 10x fewer ablations. The methods apply to feature, data, and model component attribution.

  2. FunAudioLLM (Alibaba Tongyi) · new models on Hugging FaceAI score44

    Fun-CineForge Releases Open-Source Dubbing Pipeline, Model, and CineDub-CN Dataset

    AIFun-CineForge, from FunAudioLLM, is an open-source toolkit with an end-to-end dataset pipeline and an MLLM-based model for zero-shot movie dubbing across diverse cinematic scenes. The team built CineDub-CN, described as the first large-scale Chinese television dubbing dataset, and reports that its model outperforms state-of-the-art methods on audio quality, lip-sync, timbre transition, and instruction following. Inference code and checkpoints were released on March 16, 2026, and the model runs on a consumer-grade GPU.

Feb 13

Feb 13Fri
  1. MiniMax BlogAI score62

    MiniMax details Forge, a scalable agent RL framework behind M2.5

    AIMiniMax describes Forge, its internal reinforcement learning framework for training real-world agents, which was used during the development of MiniMax M2.5. The post explains a Windowed FIFO scheduler, prefix tree merging that the post says yields a 40x training speedup, and CISPO-based training across more than one hundred thousand agent scaffolds and environments.

    Why it matters: The post details how the Forge framework balances throughput, stability, and agent flexibility, with concrete scheduling and prefix-merging methods for training agent RL at scale.

Feb 4

Feb 4Wed
  1. Anthropic EngineeringAI score72

    Anthropic finds container resource limits can shift agentic coding eval scores

    AIAnthropic reports that resource configuration alone can move Terminal-Bench 2.0 scores by up to 6 percentage points, with infra error rates falling from 5.8% under strict enforcement to 0.5% when uncapped. Above about 3x the per-task specs, extra headroom starts letting agents solve tasks they previously could not, so limits can change what the eval measures.

    Why it matters: The source shows how container resource limits shift agentic coding scores, which helps readers interpret small leaderboard gaps and set up evals more consistently.

Jan 29

Jan 29Thu
  1. Z.ai (GLM) · new models on Hugging FaceAI score60

    Z.ai releases open-source GLM-OCR multimodal document model

    AIZ.ai has released GLM-OCR, a 0.9B-parameter multimodal OCR model for complex document understanding, under the MIT License. The model scores 94.62 on OmniDocBench V1.5 and supports deployment through vLLM, SGLang, and Ollama, with an official SDK for document parsing.

    Why it matters: The page gives benchmark scores, a 0.9B parameter size, and supported serving frameworks, which help readers weigh OCR deployment options against heavier alternatives.

Jan 14

Jan 14Wed
  1. Black Forest Labs · new models on Hugging FaceAI score46

    FLUX.2 [klein] 9B Base Released on Hugging Face as Undistilled Open-Weight Model

    AIBlack Forest Labs has released FLUX.2 [klein] 9B Base, a 9 billion parameter undistilled rectified flow transformer with open weights for text-to-image generation and multi-reference editing. The model is intended for fine-tuning, LoRA training, and research, and fits in about 29GB VRAM on NVIDIA RTX 4090-class GPUs. A reference implementation is available on GitHub, and the model works with ComfyUI and Diffusers.

Jan 10

Jan 10Sat
  1. Berkeley AI ResearchAI score36

    Information-Driven Design Framework Evaluates Imaging Systems by Mutual Information

    AIBerkeley AI Research proposes an information-based framework that evaluates and optimizes imaging systems using mutual information estimated directly from noisy measurements. The team reports that the metric predicts decoder performance across color photography, radio astronomy, lensless imaging, and microscopy, and that optimized designs match end-to-end methods while requiring less memory and compute.

Dec 2, 2025

Dec 2, 2025Tue
  1. Apple · new models on Hugging FaceAI score36

    Apple releases CLaRa-7B-Instruct for compressed-document retrieval-augmented QA

    AIApple has published CLaRa-7B-Instruct on Hugging Face, an instruction-tuned unified RAG model with built-in semantic document compression at 16× and 128× ratios. The model answers instruction-following questions directly from compressed document representations, and its paper, GitHub repository, and transformers usage example are referenced in the release.

Oct 26, 2025

Oct 26, 2025Sun
  1. Thinking Machines LabAI score70

    Thinking Machines Lab explains on-policy distillation for cheaper LLM post-training

    AIThinking Machines Lab describes on-policy distillation, which samples rollouts from a student model and has a teacher grade each token with reverse KL. The authors report that this matches Qwen3-style reasoning results at a fraction of RL's cost, with AIME'24 reaching 70% in about 150 steps from a 400k SFT checkpoint. The method also helps recover instruction-following behavior lost during fine-tuning on internal documents.

    Why it matters: The post explains why on-policy distillation gives dense per-token feedback, letting a small model match RL results at much lower compute cost.

Sep 3, 2025

Sep 3, 2025Wed
  1. Cognition Blog (Devin, Windsurf)AI score38

    Eight Sleep Uses Devin AI as Data Analyst to Clear Ad-Hoc Requests

    AIEight Sleep integrated Cognition's Devin into its data workflows, letting staff tag Devin in Slack to query Snowflake, dbt, and Looker and check Amplitude. The company says it is now shipping 3x as many data features and investigations each week, with its ad-hoc data request queue near zero. Devin was used to trace a suspicious revenue spike to a better-than-expected email campaign.

Aug 27, 2025

Aug 27, 2025Wed

May 5, 2025

May 5, 2025Mon
  1. Cognition Blog (Devin, Windsurf)AI score39

    Kevin-32B Uses Multi-Turn Reinforcement Learning to Write Faster CUDA Kernels

    AIStanford and Cognition AI researchers introduced Kevin-32B, a 32B-parameter model trained with multi-turn reinforcement learning to write CUDA kernels. On KernelBench, it solves 89% of tasks at best@16 and achieves 65% average correctness over eight refinement steps, versus 53% for o4-mini and 51% for o3. Its best@16 speedup is 1.41x, and multi-turn training outperforms single-turn training as refinement steps increase.

Nov 30, 2024

Nov 30, 2024Sat
  1. Liquid AI BlogAI score60

    Liquid AI's STAR uses evolutionary search to synthesize tailored model architectures

    AILiquid AI reports STAR, an evolutionary algorithm that synthesizes tailored neural network architectures from a numerical genome representation. The authors say it produced hundreds of designs that outperform Transformer and hybrid architectures in quality, with smaller caches and parameter counts, and can optimize for latency on target hardware. The full method is described in the arXiv technical report 2411.17800.

    Why it matters: The post explains how evolutionary search over a new architecture design space produced designs beating Transformers and hybrids, giving a concrete method for quality versus latency and memory trade-offs.