Skip to contentSkip to stories

Updated

#Data/Training

Sep 29

Sep 29Tue
  1. OpenBMBAI score72

    One-Shot OPD: One Training Query Matches Most of Full-Data Distillation Gains

    AIResearchers from Tsinghua NLP and collaborators show that on-policy distillation with a single training query recovers 87% of full-data gains on math, reaching 68.5 versus 69.8 by step 300. The paper attributes the slow progress to how fast the student absorbs the teacher's signal rather than to dataset size. Code and the paper are publicly available on GitHub and Hugging Face.

    Why it matters: The paper isolates training data from the algorithm, showing one query nearly matches full-data on-policy distillation, which reframes where post-training gains come from.

  2. Ahead of AI (Sebastian Raschka)AI score43

    Language Models for Text Classification: From Bag-of-Words to Jev

    AISebastian Raschka traces text classification from bag-of-words models such as naive Bayes and logistic regression through pre-transformer neural networks, then sets up an analysis of the recently released Jev AI model. The article frames Jev as a general-purpose classifier that trades specialized accuracy for speed, cost, and breadth of tasks.

  3. Azure BlogAI score46

    Microsoft Fabric and Copilot Integration: New Data Foundation Features for Agents

    AIMicrosoft is bringing business context from Fabric IQ into Microsoft Copilot, with Fabric IQ in Copilot Chat and Cowork generally available and integration into the new Code experience coming soon through the Frontier program. Power BI is also gaining agentic app creation in Power BI Desktop, letting users generate applications from trusted semantic models and publish them to Microsoft Fabric.

Sep 28

Sep 28Mon
  1. Alexander DoriaAI score14

    Document parsing favors large models: Astra annotates, Gemma 4 31B finetunes

    AIAlexander Doria says high parameter capacity still matters for harder document processing, running Astra for initial annotation and Gemma 4 31B for finetuning. Yifei Hu reports that gpt-6-sol improved over last week's version on domain-specific document parsing but remains far behind gpt-6-astra, with the benchmark itself built using Astra.

  2. Microsoft ResearchAI score30

    Microsoft Research Asia – Singapore marks one year advancing AI research, partnerships and talent

    AIMicrosoft Research Asia – Singapore, opened July 24, 2025 as Microsoft's first Southeast Asian research lab, reports progress after its first year. The lab's work spans next-generation AI models and agentic systems, domain-specific AI for real-world impact, AI-native research practices, and ecosystem and talent development. Its healthcare collaborations on multimodal and agentic AI for clinical decision-making are being deployed through partnerships across Singapore's healthcare ecosystem.

  3. SemiAnalysisAI score43

    How GLM-5.3 Sparse Attention Affects HBM and Serving Costs on GB200, GB300, and MI355X

    AISparse attention cuts per-operation KV cache reads but does not reduce overall memory capacity, so top-k cache misses still depend on HBM. SemiAnalysis's InferenceX estimates GB200 at about $0.044 per million total tokens at 150 tokens per second, roughly 12% below MI355X running ATOM at $0.049. Neither system holds a uniform cost advantage across the tested 100, 125, and 150 tokens-per-second targets.

  4. LlamaIndexAI score30

    LlamaIndex says frontier VLMs still struggle parsing tax and W-series forms

    AILlamaIndex argues that frontier vision-language models still fail on real forms such as W-2s, 1040s, W-9s, and scanned W-4s, because forms require detecting every field, preserving section hierarchy, linking values to their exact boxes, and reading handwriting and checkmarks. The company's blog post details these failure modes and presents a custom cookbook for LlamaParse as a cheaper way to handle such forms.

  5. Epoch AI · The Epoch BriefAI score62

    Epoch AI finds AI cost per benchmark score falling 13× per year

    AIEpoch AI estimates that the cheapest cost of reaching a given benchmark score has fallen about 13× per year over the past five years, faster than DNA sequencing, compute, lithium batteries, or electricity. Its example: a 75% GPQA Diamond score that cost about 30 cents per question with o3 in January 2025 cost $0.0004 per question with GPT-5.6 Luna under 18 months later. The authors caution that benchmarks are imperfect proxies for market prices, and the decline rate slows over time.

    Why it matters: The source compares AI price declines with other transformative technologies using benchmark-based cost estimates, giving readers a measured sense of how fast cost per capability is falling.

  6. Rest of WorldAI score46

    SK Hynix and Samsung race for bigger roles in U.S. AI chip boom

    AITwo South Korean companies, SK Hynix and Samsung, control 83% of the global memory chip market and are competing to deepen ties with U.S. AI firms. Samsung has more than doubled its market share over the past year, while SK Hynix has held a long lead in high-bandwidth memory and is building a $4 billion advanced packaging facility in Indiana. Both remain exposed to China, where they face new U.S. licensing requirements and growing domestic competition.

  7. ModelScopeAI score43

    Jina-OCR-v1 parses full pages into Markdown at 2.57 pages per second

    AIJina-OCR-v1, a 3.4B-parameter MoE model that activates 570M parameters per token, converts entire document pages into structured Markdown at 2.57 pages per second. It scores 91.14 on OmniDocBench v1.6 and 83.4 on olmOCR-Bench, 7.4 points above DeepSeek-OCR on the latter, and delivers the highest throughput among 14 evaluated systems at concurrency 32. The model is released under CC BY-NC 4.0, so commercial use requires permission.

Sep 27

Sep 27Sun
  1. Alexander DoriaAI score18

    Doria Argues Local Language Models Ease Industrial Audit Demands

    AIAlexander Doria argues that embedding AI models in industrial processes requires easing audits and supporting local languages and professional registers. He says this need is real despite sarcasm about the idea. Background from Aleph Alpha notes that reasoning models think in English even for German prompts, and that small German reasoning data doses can hurt performance while large doses mostly recover it.

  2. Xiaomi MiMoAI score62

    Xiaomi MiMo Explains Fixing Tool-Call Repetition in MiMo-V2.6 Models

    AIXiaomi MiMo reports that tool-call repetition in MiMo-V2.6 reached over 0.05% of responses across agent harnesses, causing stalled agents and wasted context. The team traced the cause to an RL flooding penalty set at 32 calls per turn, which missed smaller excess behavior, and replaced the approach with a specialized teacher distilled via MOPD. Repetition rates for both Pro and Flash dropped substantially, at roughly $90,000 versus an estimated $2.31 million for the alternative fix.

    Why it matters: The post traces an agent failure to a reward blind spot and compares the costs of two fixes, offering a transferable debugging method for RL-trained tool-calling models.

Sep 26

Sep 26Sat
  1. DeedyAI score40

    Economics of neolabs: why GPU spend makes frontier-chasing hard

    AIA neolab is a startup of AI researchers that raises large pre-production funding to finance GPU compute, with 1000 GB300s (about 14 NVL72 racks) costing $125-150M over 3 years, roughly 2-2.5MW. That buys about 10^25 FLOPs per quarter, enough for a GPT-4-level model that is 1-2 OOMs behind the frontier for pretraining. Recouping $10M in training at 50% inference margin would take serving about 10T tokens at a $2/M blended price, so neolabs often pivot to a different model game, proprietary data, or high-revenue niches.

  2. SemiAnalysisAI score62

    Intel Panther Lake teardown reveals 18A RibbonFET and PowerVia design details

    AISemiAnalysis tore down Intel's Panther Lake chip, examining its 18A process with RibbonFET gate-all-around transistors and PowerVia backside power delivery. Measurements put 18A compute logic and TSMC N3E GPU logic at similar logic density, though 18A does not lead TSMC N3P, N2, or Samsung SF2 in peak density. The analysis also compares the compute, GPU, I/O, and Foveros-S packaging against Lunar Lake and Samsung's SF2 process.

  3. Sebastian RaschkaAI score30

    Raschka's Reasoning from Scratch Covers Log-Probability Scoring and Self-Refinement

    AISebastian Raschka's fifth Reasoning from Scratch video explains log-probability scoring and self-refinement for LLMs. It covers token probabilities, PyTorch implementation, numerical stability, and a self-refinement loop evaluated on MATH-500, with the log-probability concept linked to cross-entropy loss in pre-training and distillation.

  4. Alexander DoriaAI score38

    Xiaomi open-sources 989 RL environments used for a 9B MiMo model

    AIAlexander Doria reports that the released set is a smaller selection of 989 environments for RL training a 9B distilled model, not the full MiMo. Rewards are not self-contained: the general part requires setting up a judge, and webdev relies on its own grader service and VLM. The most important content is in the general/envs directory and Docker setup rather than the Hugging Face dataset, offering a solid mix of real and simulated documents.

  5. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score45

    Intern-Decision-4B: Multimodal structured decision model from Qwen3.5-4B

    AIShanghai AI Lab's InternLM released Intern-Decision-4B, a multimodal structured decision model fine-tuned from Qwen3.5-4B, which returns answer distributions for multiple questions in one forward pass. On its benchmark table it scores an average of 90.02 with a Brier score of 0.347 and an ECE of 0.065, and per-query latency averages 44.16 ms on a single RTX 4090. The model is available with a Python DecisionEngine inference interface.

  6. InternLM (Shanghai AI Lab) · new models on Hugging FaceAI score46

    Intern-Decision-0.8B: InternLM's structured decision model on Hugging Face

    AIInternLM released Intern-Decision-0.8B, a multimodal structured decision model fine-tuned from Qwen3.5-0.8B that scores answers to multiple questions in one forward pass. The model reports a 79.38 average score and a 33.98 ms mean latency on a single RTX 4090, with 0.8B, 2B, and 4B sizes available. It is accessed through a Python DecisionEngine API that returns calibrated probabilities rather than generating free-form text.

Sep 25

Sep 25Fri
  1. Google Cloud · AI & Machine LearningAI score43

    Google Cloud Introduces Managed Reinforcement Learning Fine-Tuning for Gemini Models

    AIGoogle Cloud has launched a managed reinforcement learning fine-tuning service (RLFT) that lets customers adapt Gemini models using a reward function they define instead of labeled answers. Users supply prompts and a reward function, while Google handles the RL infrastructure and proprietary model internals. The guide advises exhausting prompting and supervised fine-tuning first, and notes that RLFT suits tasks that are easy to score but hard to demonstrate.

  2. SemiAnalysisAI score59

    China Holds Over 24GW of Datacenter Capacity, Shifting Inland With AI Demand

    AISemiAnalysis's China Datacenter Model tracks over 1,000 facilities across 60+ operators and puts China's fleet above 24GW, larger than EMEA. The report attributes the buildout to the Eastern Data, Western Compute policy and hyperscale AI demand, which is moving capacity to western hubs such as Inner Mongolia at construction speeds it says the West cannot match.

Sep 24

Sep 24Thu
  1. Lewis TunstallAI score42

    Hugging Face releases over 5,000 RL environments for data science tasks

    AIHugging Face released SmolDataEnvs, more than 5,000 open-source RL environments aimed at real-world data science tasks. They target the gap between simple educational games and frontier-level benchmarks, especially for improving coding in models under 10B parameters. The environments are designed as a testbed for developing new RL methods such as GRPO or OPSD.

  2. LlamaIndexAI score17

    LlamaIndex Explains Using Confidence Scores to Control Document Extraction Automation

    AILlamaIndex argues that extraction confidence scores are useful only when they help decide what can be automated and what needs human review. Using ExtractBench, the post compares extraction systems after confidence filtering, reporting that LlamaParse Agentic Plus reached 66.48% recall on expected fields at a 97% precision target. The post covers confidence cutoffs, precision versus recall, score coverage, score granularity, and human review volume.

  3. Azure BlogAI score67

    Microsoft Foundry adds voice agents and continuous optimization for production agents

    AIMicrosoft Foundry expands its agent platform with voice agents in public preview, long-running resilience for hosted agents, and tools for evaluating production agents. The post also says GPT-6 Sol, GPT-6 Luna, and Claude Opus 5.5 are now available in Foundry. Agent optimizer, Insights, and Rubric evaluator are described as tools for continuous improvement, with some reaching general availability later this month.

    Why it matters: The post shows how Foundry combines model choice, voice agents, long-running resilience, and production evaluation into one agent workflow, with a customer example.

  4. Google ResearchAI score38

    Google's John Platt on AI for climate, disease forecasting, and science

    AIIn a Latent Space podcast episode, Google's John Platt discusses using AI to address climate change, including reducing airplane contrails that contribute about 1% of human-caused warming and detecting fires with FireSat satellites. He also describes Google's Empirical Research Assistance (ERA), which uses Gemini and Monte Carlo Tree Search and achieved top marks in recent CDC benchmarks for forecasting COVID and flu cases a week ahead.

  5. vLLMAI score34

    vLLM and RL-Kernel achieve bit-exact logprob match on AMD MI300X

    AIThe RLKernel team integrated RL-Align/RL-Kernel with vllm-project/vime, and a 200-step Qwen3-8B GRPO run on 8× AMD MI300X recorded zero logprob mismatches between Megatron training and vLLM rollout. The strict path aligns reduction order, intermediate precision, rounding points, and math primitives across both sides to achieve bit-for-bit matching on ROCm.