Skip to contentSkip to stories

Updated

#Data/Training

Showing low-relevance items too. Hide low-relevance items

Sep 26

Sep 26Sat
  1. InternLM (Shanghai AI Lab) · new models on Hugging FaceOfficialAI score46

    Intern-Decision-0.8B: InternLM's structured decision model on Hugging Face

    AIInternLM released Intern-Decision-0.8B, a multimodal structured decision model fine-tuned from Qwen3.5-0.8B that scores answers to multiple questions in one forward pass. The model reports a 79.38 average score and a 33.98 ms mean latency on a single RTX 4090, with 0.8B, 2B, and 4B sizes available. It is accessed through a Python DecisionEngine API that returns calibrated probabilities rather than generating free-form text.

Sep 25

Sep 25Fri
  1. Google Cloud · AI & Machine LearningOfficialAI score43

    Google Cloud Introduces Managed Reinforcement Learning Fine-Tuning for Gemini Models

    AIGoogle Cloud has launched a managed reinforcement learning fine-tuning service (RLFT) that lets customers adapt Gemini models using a reward function they define instead of labeled answers. Users supply prompts and a reward function, while Google handles the RL infrastructure and proprietary model internals. The guide advises exhausting prompting and supervised fine-tuning first, and notes that RLFT suits tasks that are easy to score but hard to demonstrate.

  2. LlamaIndex 🦙OfficialAI score18

    LlamaParse preserves complex Fed forecast tables for AI analysis

    AILlamaIndex tested LlamaParse on the Fed's September 2026 projections PDF, where a 2029 column's June comparison cell is blank. The company says all nine GDP median figures on page 2 matched the original PDF after parsing, with alignment kept in the returned HTML.

    Image from @llama_index's post

Sep 24

Sep 24Thu
  1. Sundar PichaiXAI score38

    Google's Project Suncatcher tests TPUs in space on a SpaceX mission

    AIGoogle's Project Suncatcher will fly a prototype satellite built with Planet aboard SpaceX's Transporter-18 mission to test whether TPUs can survive and operate in space. The test asks whether the chips can function in orbit, and the post frames it as a first step.

    Video from @sundarpichai's post
  2. Lewis Tunstall @ COLM 🌉XAI score42

    Hugging Face releases over 5,000 RL environments for data science tasks

    AIHugging Face released SmolDataEnvs, more than 5,000 open-source RL environments aimed at real-world data science tasks. They target the gap between simple educational games and frontier-level benchmarks, especially for improving coding in models under 10B parameters. The environments are designed as a testbed for developing new RL methods such as GRPO or OPSD.

  3. LlamaIndex 🦙OfficialAI score17

    LlamaIndex Explains Using Confidence Scores to Control Document Extraction Automation

    AILlamaIndex argues that extraction confidence scores are useful only when they help decide what can be automated and what needs human review. Using ExtractBench, the post compares extraction systems after confidence filtering, reporting that LlamaParse Agentic Plus reached 66.48% recall on expected fields at a 97% precision target. The post covers confidence cutoffs, precision versus recall, score coverage, score granularity, and human review volume.

    Video from @llama_index's post
  4. Google ResearchOfficialAI score38

    Google's John Platt on AI for climate, disease forecasting, and science

    AIIn a Latent Space podcast episode, Google's John Platt discusses using AI to address climate change, including reducing airplane contrails that contribute about 1% of human-caused warming and detecting fires with FireSat satellites. He also describes Google's Empirical Research Assistance (ERA), which uses Gemini and Monte Carlo Tree Search and achieved top marks in recent CDC benchmarks for forecasting COVID and flu cases a week ahead.

  5. vLLMOfficialAI score34

    vLLM and RL-Kernel achieve bit-exact logprob match on AMD MI300X

    AIThe RLKernel team integrated RL-Align/RL-Kernel with vllm-project/vime, and a 200-step Qwen3-8B GRPO run on 8× AMD MI300X recorded zero logprob mismatches between Megatron training and vLLM rollout. The strict path aligns reduction order, intermediate precision, rounding points, and math primitives across both sides to achieve bit-for-bit matching on ROCm.

  6. Liquid AI NewsletterOfficialAI score38

    Liquid AI optimizes its on-device context layer for Snapdragon processors and releases longevity models

    AILiquid AI says its Liquid Context on-device context layer is now optimized for Snapdragon processors using the Qualcomm Hexagon NPU, announced at Qualcomm's Snapdragon Summit. The company also released LFM2-1.2B-Longevity and LFM2-2.6B-Longevity, which it says match or outperform much larger frontier LLMs on longevity prediction, with LFM2-2.6B-Longevity more than 80% accurate on clinical age prediction. The open LongevityBench benchmark, with 17 tasks and 25,457 prompts, is available on Hugging Face.

  7. Google · Innovation & AIOfficialAI score62

    Google's Project Suncatcher will test TPUs in orbit on a prototype satellite

    AIGoogle's Project Suncatcher will launch a prototype satellite on the Transporter-18 rideshare mission with SpaceX to test how its TPUs handle spaceflight. Initial ground tests showed the Trillium TPUs survived vibration and a radiation dose greater than a five-year space mission would deliver. Google says cooling with heat pipes and radiators and laser links between satellites in 2027 remain open engineering challenges.

    Why it matters: The source reports concrete radiation, vibration, and cooling test results for TPUs, showing what space-based AI compute still has to solve.

  8. Goodfire ResearchOfficialAI score54

    LLM activations trace emotional story arcs through neural geometry over time

    AIGoodfire Research examines how language models track emotional dynamics across a story, sentence by sentence. The authors prompt Llama 3.1 8B to rate six emotions after each sentence, then harvest activations from each sentence's last token and fit a manifold to show stories tracing trajectories through it.

  9. Goodfire ResearchOfficialAI score57

    Goodfire finds sparse autoencoder features capture curved neural geometry in three ways

    AIGoodfire Research examines how sparse autoencoder directions relate to curved manifolds in neural representations, identifying shattering, compact capture, and dilution as three ways lines can represent them. The team trained an autoencoder on synthetic data containing shapes such as donuts, spheres, and Möbius strips, and reports that real features in Llama 3.1 8B show dilution. It also describes an unsupervised pipeline that clusters features by firing patterns to surface manifolds in that model.

  10. KrASIA · Big TechNewsAI score55

    Mind Lab launches Mint Recursive, a post-training platform for companies

    AIMind Lab unveiled Mint Recursive, a post-training and inference platform for industry use, alongside Macaron-V1.1, a model post-trained entirely on it. Macaron-V1.1 is a 752-billion-parameter model built from GLM-5.3 with four two-billion-parameter LoRA expert modules for chat, agents, coding, and generation. The platform is serverless and bills by token usage, and it collects feedback from models in use to support continued training.

  11. AI at MetaOfficialAI score22

    Meta distills 40-step video diffusion into a 2-step live streaming model

    AIMeta distilled a 40-step diffusion teacher using 3-way CFG, requiring 120 evaluations per video chunk, into an unguided 2-step causal student with a fixed-length KV cache. The student uses self-forcing to resist drift and keep near-teacher quality while needing 60x fewer evaluations, enabling instant responses in live video streaming.

    Image from @AIatMeta's post
  12. LangChain BlogOfficialAI score50

    LangSmith Fine-Tuning and smithtune Turn Agent Trajectories Into Custom Models

    AILangChain launched LangSmith Fine-Tuning and smithtune, a CLI that turns LangSmith agent trajectories into fine-tuned models through dataset creation, training with Fireworks or Baseten, and evaluation in LangSmith. smithtune currently supports supervised fine-tuning, training models on recorded examples of good agent behavior by updating model weights. The tool lets teams train specialized models without building the data pipeline by hand.

Sep 23

Sep 23Wed
  1. Tencent HyOfficialAI score38

    Tencent Hunyuan studies batch-size scaling for LLM reinforcement learning efficiency

    AITencent Hunyuan extends classical critical-batch-size theory to online LLM reinforcement learning, where models generate their own training data. Across GRPO and PPO, learning-rate retuning preserves learning per response over a bounded range of batch sizes. On fixed hardware, larger batches raise PPO generation-stage throughput by up to 2.29×, and the best measured GRPO setup reaches the same validation target in 29% less time.

  2. Google Developers BlogOfficialAI score62

    Google reproduces Olmo 3 7B pre-training in MaxText on TPUs

    AIGoogle Developers reproduced Ai2's Olmo 3 7B from scratch in MaxText on Google Cloud TPUs, covering both the stage-1 pre-training run and the stage-2 mid-training anneal. The match was checked on held-out C4 loss, an 8-task accuracy suite, multi-domain perplexity, and token-level KL, not just the training loss curve. The post also describes a data-loader bug that made training loss look better than the reference while held-out metrics did not move.

    Why it matters: The post documents how a faithful reproduction was verified on held-out metrics, including a data bug that training loss alone would have hidden.

  3. eric zakariassonXAI score67

    Cursor shares a prompt for reducing token cost in agent harnesses

    AICursor's Eric Zakariasson shared a prompt for improving an LLM agent harness to lower token cost per completed task without losing quality. The prompt covers the system prompt, tool definitions, cache layout, tool results, compaction, and subagents, and reports that one team's round of these changes cut overall token cost about 7%.

    Why it matters: The prompt gives a concrete checklist for cutting agent token cost per completed task, with tested figures on cache layout, tool offloading, and compaction.

  4. Dario AmodeiXAI score76

    Claude Helps Discover a Possible New Gene Editing Enzyme System

    AIAnthropic announced that Claude, working mostly on its own, identified a previously unknown enzyme system in bacteriophage DNA that may represent a new gene editing mechanism. Claude read literature and genome data, proposed experiments, and Anthropic's team carried them out. The function and biotechnological utility of the system remain unclear.

    Why it matters: The post pairs a Claude-led discovery with the lab workflow used to verify it, showing how AI and humans split the research work in biology.

  5. AnthropicOfficialAI score62

    Claude finds a previously unknown enzyme system in bacteriophage DNA

    AIClaude has identified a previously unknown enzyme system in bacteriophage DNA, located beside a long array of repeating DNA that somewhat resembles CRISPR. Anthropic says its function is not yet understood, but only a handful of known systems share its features, all of which can cut, copy, and paste DNA. The source notes that programmable systems like CRISPR have been important to medicine, but more work is needed to learn what this system does and whether it can be used similarly.

    Why it matters: The source reports a newly found enzyme system with CRISPR-like repeats, noting that its function is still unknown and needs further study.

  6. Google DeepMindOfficialAI score62

    Google DeepMind details server-side memory for Private AI Compute

    AIGoogle DeepMind describes a persistent memory layer for its Private AI Compute platform that stores user context encrypted in the cloud. The encryption keys are held on the user's devices, and data is decrypted only inside hardware-isolated secure enclaves before being re-encrypted. The company says it is publishing a tamper-proof public record of its server software and an independent audit.

    Why it matters: The post explains how persistent cloud memory can keep personal AI context encrypted under keys held on the user's device, a concrete privacy design.

  7. Baseten BlogOfficialAI score62

    Baseten launches NVIDIA Nemotron 3 Diarization with four latency profiles

    AIBaseten has made NVIDIA Nemotron 3 Diarization available as batch, streaming, and real-time diarized transcription presets. The single checkpoint serves four algorithmic latencies from 0.32 to 30.4 seconds, and the post reports DER of 9.8% on AISHELL-4 at the low profile versus 27.2% for Streaming Sortformer v2.1.

    Why it matters: The post shows one checkpoint serving four latency profiles with DER figures against named baselines, useful for judging real-time speaker labeling tradeoffs.

  8. ModelScopeOfficialAI score40

    TeleOCR: 1.2B vision-language model parses documents, tops OmniDocBench v1.6

    AITeleOCR, a lightweight 1.2B vision-language model released under Apache 2.0, parses digital PDFs and warped phone photos without a separate dewarping model. It scores 96.87 overall on OmniDocBench v1.6, the highest among listed specialized VLMs, and ranks #1 in the ICDAR 2026 Sci-ImageMiner Challenge. It supports structured parsing of text, tables, formulas, layouts, and reading order, with synchronous or asynchronous vLLM inference.

    Image from @ModelScope2022's post
  9. QwenOfficialAI score62

    Qwen-Audio-3.1 upgrades ASR, TTS and Realtime and adds two new models

    AIAlibaba's Qwen team released Qwen-Audio-3.1, upgrading its ASR, TTS and Realtime models and adding TTS-Next and ASR-Next. The post says TTS prices fell about 70%, Realtime about 85%, and ASR up to 95%. More APIs are coming soon.

    Why it matters: The post lists the new audio models and price changes together, which helps developers compare the upgraded lineup with their current speech workflows.

    Image from @Alibaba_Qwen's post
  10. ModelScopeOfficialAI score62

    Shanghai AI Lab and SJTU release open-weight 8.9B NCP-ArchPreview model under Apache 2.0

    AIShanghai AI Lab and SJTU's LUMIA Lab released NCP-ArchPreview, an 8.9B open-weight language model under Apache 2.0. The model reportedly reaches OLMo-3-7B's final Stage 1 loss using 51.3% of the tokens from the 5.73T Dolma 3 corpus, a 1.95× convergence gain. Its concept module jointly predicts tokens and concepts, and domain adaptation updates only its 17M parameters while the token backbone stays frozen.

    Why it matters: The post pairs an Apache 2.0 open-weight release with training-efficiency figures, showing how the concept module adapts to new domains with few trainable parameters.

    Image from @ModelScope2022's post
  11. Anthropic NewsroomOfficialAI score73

    Claude agents discover a novel CRISPR-like enzyme system in bacteriophages

    AIAnthropic's new life sciences group reports that Claude autonomously identified a previously uncharacterized enzyme system, called array-associated reverse transcriptase (ART), in bacteriophages. Claude agents searched over 200,000 reverse transcriptases, narrowed 3,500 candidates to 20, and one agent flagged a CRISPR-like repeat array after about 21 hours. Human scientists then validated the finding in the lab, and the function of ART remains unknown.

    Why it matters: The post shows how Claude agents surveyed DNA sequence data, flagged a candidate, and then led to lab validation, which is a concrete workflow for AI-assisted biology research.

  12. Prime Intellect BlogOfficialAI score60

    Prime Intellect makes Prime Sandboxes generally available as microVMs for agentic RL

    AIPrime Intellect has made Prime Sandboxes generally available, offering each sandbox as a full Linux virtual machine with its own kernel and support for Docker Compose. The product is available through its CLI/SDK and RL suite, with accounts starting at 1,024 concurrent sandboxes, and pricing listed at $0.02 per vCPU-hour, $0.0125 per GiB-hour of memory, and $0.0002 per GiB-hour of disk, valid through December 22. The company says GPU microVMs, snapshotting, sandbox forking, and persistent workspaces are planned next.

    Why it matters: The post explains why full VMs rather than gVisor containers matter for agentic RL, since silent environment differences can reward behaviors that fail to transfer.

Sep 22

Sep 22Tue
  1. TinkerOfficialAI score25

    Tinker fine-tunes Qwen3.6 for Jev-style probability prompts in 10 minutes

    AITinker says an open LLM can serve a Jev-like interface that takes discrete options and returns fast probabilities, since next-token prediction is already a probabilistic classifier. A post by @ekzhang1 reports that a $5, 10-minute supervised fine-tuning run on Tinker improved Qwen3.6-35B-A3B's handling of Jev-style prompts, with +8% on GPQA Diamond and +12% on MMLU-Pro.

  2. Together AI BlogOfficialAI score38

    How to train your own Jev classifier for $17 with Together AI

    AIThe Together AI blog shows how to fine-tune a Qwen3.5 4B base model into a classification model using about 38,000 examples sampled from six Hugging Face datasets, at a training cost of roughly $17.0. The tutorial covers cloning the tev1 repository, normalizing data with provided scripts, launching a Together AI fine-tuning job that takes about 25 minutes, and deploying the result to a dedicated H100 endpoint.

  3. Fireworks AI BlogOfficialAI score46

    Fireworks ARCv3 cuts RL weight-update payloads nearly 50% for cross-region training

    AIFireworks released ARCv3, a lossless compressor for BF16 weight-update deltas sent from trainers to RL rollout machines. Across 1,000 production RL deltas, ARCv3 produced payloads nearly 50% smaller than ARCv2, averaging about 0.19% of the BF16 weight size versus 0.36%. ARCv3 is available through the Fireworks Training API as fireworks-delta-compression.

  4. ZyphraOfficialAI score20

    Zyphra's Beren Millidge on why multi-silicon AI infrastructure matters

    AIZyphra's Chief Scientist Beren Millidge, in an AI Infra Summit interview with vCluster Labs CEO Lukas Gentele, argued that a heterogeneous compute future is inevitable. The interview covers why Zyphra chose AMD over NVIDIA, along with topics such as kernel writing, surviving GPU failures mid-run, and routing. Zyphra says it is working to build a strong multi-silicon ecosystem.

  5. Tri DaoXAI score44

    Rigel: 2.3B hybrid Mamba-2 MoE nears Llama-3.2-3B with <1% FLOPs

    AIMayank's Rigel, a 2.3B-parameter MoE (360M active) hybrid Mamba-2 model, was pretrained across H100, A100, V100 GPUs and TPU v5p/v6e on one codebase. The model lands within a few points of Llama-3.2-3B while using under 1% of its pretraining FLOPs. Tri Dao praised the work's engineering effort and the model's strength for its small size.

  6. Greg BrockmanXAI score81

    OpenAI launches GPT-6 Sol and Luna with 50% lower API prices than GPT-5.6

    AIOpenAI introduced GPT-6 Sol and GPT-6 Luna, which it says bring much of the strength of GPT-6 Astra into faster and more affordable models. The company also reports more efficient caching and inference, with API prices 50% lower than GPT-5.6 promotional pricing.

    Why it matters: The quoted announcement names specific pricing and access changes for Sol and Luna, which matter for teams weighing cost against the Astra tier.