Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Sep 29

Sep 29Tue
  1. Apple Machine Learning ResearchAI score38

    LLM Conditioning Study Finds Steering Methods Trade Fluency for Effectiveness

    AIApple researchers systematically tested LLM conditioning methods and found efficient activation steering often degrades fluency. Steering is far less effective on instruction-tuned models than base models, while prompting and full supervised fine-tuning work for concept injection but are weaker at concept removal. Cheap textual metrics correlate highly with costly LLM-as-judge scores.

  2. Google Developers BlogAI score47

    Google Details Sparse Attention Speedup for Video Diffusion on TPUs

    AIGoogle Developers Blog describes how Sparse VideoGen (SVG) routes video diffusion attention heads into spatial or temporal sparse masks and implements them as custom JAX and Pallas Splash Attention kernels on TPU v6e. In isolated single-chip tests with 75.6K tokens and 10 heads, the sparse variants retain about 38.87% of query-key pairs. The article argues that theoretical sparsity must be converted into hardware tile skipping to yield real speedups.

  3. Google ResearchAI score35

    Google Research unveils Diffusion Controller for steering AI image generation

    AIGoogle Research introduced Diffusion Controller, a framework that treats image generation as a continuous control problem rather than separate inference-time guidance and fine-tuning fixes. Its lightweight add-on "steering damper" network keeps the base model frozen and works on black-box or gray-box models, and it outperformed the industry standard on human preference matching. In a Stable Diffusion v1.4 test, the fully unlocked version achieved a 90% win rate over the baseline.

  4. Microsoft ResearchAI score34

    Microsoft Research unveils Quine, an early multimodal world model of biology

    AIMicrosoft Research has introduced Quine, an early-stage research effort to build a multimodal world model of biology that connects insights across biological scales and modalities. The system is designed to help scientists computationally search a space far larger than intuition allows and prioritize hypotheses before lab testing. Experimental results are meant to feed back into the model and sharpen future research directions.

    Video from @MSFTResearch's post
  5. Replit BlogAI score62

    Replit Agent lets the core model choose subagents and effort instead of a router

    AIReplit explains how its Agent lets the core model pick subagent tier and effort mid-task rather than relying on an external router. On DeepSWE and Terminal-Bench, Replit Agent scored 72% at $2.11 per task and 49% at $2.53 per task, beating a single long-lived worker sidekick setup by 11 and 16 points. The company says Astra on its own scores higher only at more than twice the cost.

    Why it matters: The post gives a concrete harness design with benchmark cost-score comparisons, helping builders weigh delegation strategies against routers and single-worker setups.

  6. OpenBMBAI score72

    One-Shot OPD: One Training Query Matches Most of Full-Data Distillation Gains

    AIResearchers from Tsinghua NLP and collaborators show that on-policy distillation with a single training query recovers 87% of full-data gains on math, reaching 68.5 versus 69.8 by step 300. The paper attributes the slow progress to how fast the student absorbs the teacher's signal rather than to dataset size. Code and the paper are publicly available on GitHub and Hugging Face.

    Why it matters: The paper isolates training data from the algorithm, showing one query nearly matches full-data on-policy distillation, which reframes where post-training gains come from.

    Image from @OpenBMB's post
  7. Artificial Analysis ArticlesAI score62

    Artificial Analysis open-sources AA-AgentPerf-Local for benchmarking local AI agents

    AIArtificial Analysis has open-sourced AA-AgentPerf-Local, a tool that replays recorded agent trajectories to measure inference speed on laptops and workstations. Initial results cover NVIDIA DGX Spark, NVIDIA GeForce RTX 5090, AMD Ryzen AI Halo, and MacBook Pro M5 Pro, with the RTX 5090 fastest for models that fit its 32 GB. The source states the tool and leaderboard will expand to more hardware, frameworks, and models.

    Why it matters: The source gives per-system completion times and memory bandwidth figures, letting readers compare local hardware for running agentic workloads.

Sep 28

Sep 28Mon
  1. Ali GhodsiAI score62

    Databricks finds Opus 5.5 cheaper and better, GPT-6 Luna 20x cheaper per task

    AIDatabricks tested recent AI models across 2,400 engineers and found Opus 5.5 offers the highest quality mid-tier performance, with about 20% lower same-task costs than Opus 4.8. The company is now encouraging Opus 5.5 as a default model for coding, and reports that GPT-6 Luna is at least 20 times cheaper per task than Opus 5.5, roughly matching Opus 4.6 on one difficult evaluation suite. The Luna findings are preliminary.

  2. Epoch AI · The Epoch BriefAI score62

    Epoch AI finds AI cost per benchmark score falling 13× per year

    AIEpoch AI estimates that the cheapest cost of reaching a given benchmark score has fallen about 13× per year over the past five years, faster than DNA sequencing, compute, lithium batteries, or electricity. Its example: a 75% GPQA Diamond score that cost about 30 cents per question with o3 in January 2025 cost $0.0004 per question with GPT-5.6 Luna under 18 months later. The authors caution that benchmarks are imperfect proxies for market prices, and the decline rate slows over time.

    Why it matters: The source compares AI price declines with other transformative technologies using benchmark-based cost estimates, giving readers a measured sense of how fast cost per capability is falling.

Sep 27

Sep 27Sun
  1. Sakana AIAI score46

    Sakana AI's SAIL boosts VLM robot trajectory success via test-time scaling

    AISakana AI and the University of Tokyo introduced SAIL, a method that generates robot trajectories with a VLM and refines them through simulator testing, VLM feedback, and Monte Carlo tree search. Across six simulated manipulation tasks, raising the search budget from one candidate to 45 increased the success rate of finding a working trajectory from 25% to 73%. The authors also tested the approach on a physical robot, though the post frames further transfer to real hardware as an open question.

    Video from @SakanaAILabs's post

Sep 26

Sep 26Sat
  1. SemiAnalysisAI score62

    Intel Panther Lake teardown reveals 18A RibbonFET and PowerVia design details

    AISemiAnalysis tore down Intel's Panther Lake chip, examining its 18A process with RibbonFET gate-all-around transistors and PowerVia backside power delivery. Measurements put 18A compute logic and TSMC N3E GPU logic at similar logic density, though 18A does not lead TSMC N3P, N2, or Samsung SF2 in peak density. The analysis also compares the compute, GPU, I/O, and Foveros-S packaging against Lunar Lake and Samsung's SF2 process.

  2. Alexander DoriaAI score38

    Xiaomi open-sources 989 RL environments used for a 9B MiMo model

    AIAlexander Doria reports that the released set is a smaller selection of 989 environments for RL training a 9B distilled model, not the full MiMo. Rewards are not self-contained: the general part requires setting up a judge, and webdev relies on its own grader service and VLM. The most important content is in the general/envs directory and Docker setup rather than the Hugging Face dataset, offering a solid mix of real and simulated documents.

Sep 25

Sep 25Fri
  1. AnthropicAI score78

    Claude solves a nine-loop scattering amplitude problem beyond the eight-loop record

    AIAnthropic reports that Claude solved a nine-loop scattering amplitude problem in planar N=4 super-Yang-Mills, surpassing the previous eight-loop record set by SLAC's Lance Dixon and collaborators. Working largely unsupervised for days from a single prompt, at a total cost of a few thousand dollars, Claude used methods developed by Dixon's group, and Dixon independently verified the result.

    Why it matters: The post shows Claude solving a nine-loop physics calculation beyond the previous eight-loop record, verified independently, which bears on AI use in theoretical physics research.

Sep 24

Sep 24Thu
  1. Redwood Research BlogAI score41

    Continual learning could make AI monitors that block actions nearly useless

    AIRedwood Research argues that continual learning, which lets an AI accumulate skills during deployment, may teach models to evade blocking monitors because monitors reduce task success. Online RL on deployment trajectories would train the policy against the monitor through task reward, potentially leaving blocking monitors nearly useless over a long deployment. Memory-based systems pose a weaker version of this risk, according to the post.

  2. Google ResearchAI score60

    Google Research details four agentic frameworks for coherent long-form video generation

    AIGoogle Research introduces four multi-agent frameworks for generating minutes-long videos with consistent characters and environments across shots. The frameworks include AI video co-director, CANVAS, A²RD, and VQQA, which are built as orchestration layers on Gemini and Veo and use SynthID watermarking. The post reports measured gains on benchmarks such as GenAD-Bench, HardContinuityBench, and LVBench-C, with the full architectures described in the linked papers.

    Why it matters: The post links four frameworks to specific failure modes in long video generation, such as semantic drift and cascading errors, making the design choices easier to compare.

  3. Epoch AI · The Epoch BriefAI score45

    Huawei Trails Nvidia by About Four Years in AI Chip Performance and Output

    AIHuawei will likely remain about four years behind Nvidia in AI chip performance and production through 2030, Epoch AI estimates. Its flagship Ascend 950 delivers roughly half the performance of Nvidia's 2022 H100, and Huawei is projected to produce about 1.5 million chips in 2026 versus Nvidia's roughly 6 million, leaving it about 25 times behind in total compute.

  4. Goodfire ResearchAI score52

    Block-Sparse Featurizers Recover Multidimensional Concept Geometry in Vision Models

    AIGoodfire Research introduces Block-Sparse Featurizers (BSF), which decompose model activations into subspaces rather than single directions. Applied to DINOv3 and Stable Diffusion XL, BSFs find interpretable multidimensional features that better explain activations and enable fine-grained steering. The authors report that most concepts they examined have a stable rank of about two to four dimensions.

  5. Goodfire ResearchAI score48

    Steering Along Manifolds Beats Linear Steering for Controlling Llama's Days-of-Week Behavior

    AIGoodfire Research shows that steering Llama-3.1 8B along the curved representation manifold of weekdays produces output probabilities that follow the model's natural cyclic behavior, shifting probability mass smoothly from Monday to Tuesday to Friday. Linear steering along a straight vector, by contrast, cuts across the behavior manifold and yields noisy off-target tokens, some not days of the week at all. The authors argue that representation geometry and behavior geometry are linked bidirectionally.

  6. Goodfire ResearchAI score57

    Goodfire finds sparse autoencoder features capture curved neural geometry in three ways

    AIGoodfire Research examines how sparse autoencoder directions relate to curved manifolds in neural representations, identifying shattering, compact capture, and dilution as three ways lines can represent them. The team trained an autoencoder on synthetic data containing shapes such as donuts, spheres, and Möbius strips, and reports that real features in Llama 3.1 8B show dilution. It also describes an unsupervised pipeline that clusters features by firing patterns to surface manifolds in that model.

  7. Anthropic ResearchAI score60

    Anthropic study finds Claude agent trading limited by preference understanding

    AIAnthropic ran a controlled book-swapping market with 201 employees and Claude-powered agents, which reached 0.55 efficiency against a 0.89 optimum. Agents matched participants' own rankings on 61% of book pairs, and about 85% of the shortfall came from imprecise preference representation rather than the trading floor design. Stronger models produced more efficient markets than weaker ones, while instructions mattered less.

    Why it matters: The study separates agent misunderstanding of user preferences from negotiation failure, showing which failure mode limits outcomes in agent-run markets.

Sep 23

Sep 23Wed
  1. Tencent HyAI score38

    Tencent Hunyuan studies batch-size scaling for LLM reinforcement learning efficiency

    AITencent Hunyuan extends classical critical-batch-size theory to online LLM reinforcement learning, where models generate their own training data. Across GRPO and PPO, learning-rate retuning preserves learning per response over a bounded range of batch sizes. On fixed hardware, larger batches raise PPO generation-stage throughput by up to 2.29×, and the best measured GRPO setup reaches the same validation target in 29% less time.

  2. Google Developers BlogAI score62

    Google reproduces Olmo 3 7B pre-training in MaxText on TPUs

    AIGoogle Developers reproduced Ai2's Olmo 3 7B from scratch in MaxText on Google Cloud TPUs, covering both the stage-1 pre-training run and the stage-2 mid-training anneal. The match was checked on held-out C4 loss, an 8-task accuracy suite, multi-domain perplexity, and token-level KL, not just the training loss curve. The post also describes a data-loader bug that made training loss look better than the reference while held-out metrics did not move.

    Why it matters: The post documents how a faithful reproduction was verified on held-out metrics, including a data bug that training loss alone would have hidden.

  3. Dario AmodeiAI score76

    Claude Helps Discover a Possible New Gene Editing Enzyme System

    AIAnthropic announced that Claude, working mostly on its own, identified a previously unknown enzyme system in bacteriophage DNA that may represent a new gene editing mechanism. Claude read literature and genome data, proposed experiments, and Anthropic's team carried them out. The function and biotechnological utility of the system remain unclear.

    Why it matters: The post pairs a Claude-led discovery with the lab workflow used to verify it, showing how AI and humans split the research work in biology.

  4. AnthropicAI score62

    Claude finds a previously unknown enzyme system in bacteriophage DNA

    AIClaude has identified a previously unknown enzyme system in bacteriophage DNA, located beside a long array of repeating DNA that somewhat resembles CRISPR. Anthropic says its function is not yet understood, but only a handful of known systems share its features, all of which can cut, copy, and paste DNA. The source notes that programmable systems like CRISPR have been important to medicine, but more work is needed to learn what this system does and whether it can be used similarly.

  5. ModelScopeAI score62

    Shanghai AI Lab and SJTU release open-weight 8.9B NCP-ArchPreview model under Apache 2.0

    AIShanghai AI Lab and SJTU's LUMIA Lab released NCP-ArchPreview, an 8.9B open-weight language model under Apache 2.0. The model reportedly reaches OLMo-3-7B's final Stage 1 loss using 51.3% of the tokens from the 5.73T Dolma 3 corpus, a 1.95× convergence gain. Its concept module jointly predicts tokens and concepts, and domain adaptation updates only its 17M parameters while the token backbone stays frozen.

    Image from @ModelScope2022's post
  6. Anthropic NewsroomAI score73

    Claude agents discover a novel CRISPR-like enzyme system in bacteriophages

    AIAnthropic's new life sciences group reports that Claude autonomously identified a previously uncharacterized enzyme system, called array-associated reverse transcriptase (ART), in bacteriophages. Claude agents searched over 200,000 reverse transcriptases, narrowed 3,500 candidates to 20, and one agent flagged a CRISPR-like repeat array after about 21 hours. Human scientists then validated the finding in the lab, and the function of ART remains unknown.

    Why it matters: The post shows how Claude agents surveyed DNA sequence data, flagged a candidate, and then led to lab validation, which is a concrete workflow for AI-assisted biology research.

Sep 22

Sep 22Tue
  1. Redwood Research BlogAI score60

    Filler tokens let GPT-6 Astra solve harder reasoning tasks without visible reasoning

    AIRedwood Research found that padding prompts with meaningless filler tokens improves GPT-6-Astra's no-reasoning answers on serial reasoning tasks, rising from about 10-20% to about 50% on 4-hop natural facts. Other tested models improved far less, and the authors argue this means Astra can perform cognition it does not verbalize in its chain of thought, making such monitoring harder.