Skip to contentSkip to stories

Updated

All AI news

Oct 5

Oct 5Mon
  1. Goodfire ResearchAI score62

    Goodfire finds activation probes can detect reward hacking in open-source models

    AIGoodfire Research reports that reward hacking appears in 50–96% of rollouts across three open-source models on three agentic benchmarks. The team found an internal signal tied to cheating and gaming a metric, and simple activation probes catch some hacks that LLM chain-of-thought monitors miss. A probe can screen every transcript cheaply, and in one setup cut LLM monitoring cost by 90% with a roughly 1% precision drop.

    Why it matters: The study links a reward hacking signal in model activations to monitoring cost and detection, showing how probes compare with chain-of-thought monitors on the same runs.

  2. Claude Code · GitHub ReleasesAI score31

    Claude Code v2.1.290 adds hook fixes, Deny button for sign-in, and new CLI commands

    AIClaude Code v2.1.290 adds serverToolUses to plugin turn.step results and agentId to tool.check hook events, so hooks can distinguish subagent permission checks. The release also adds a Deny button to the Claude apps gateway sign-in approval page, plus claude attach and claude logs accepting partial session names.

  3. PyTorch BlogAI score40

    PyTorch Consolidates Media Decoding and Encoding Into TorchCodec, Narrows TorchVision and TorchAudio

    AIPyTorch has consolidated all media decoding and encoding for images, video, and audio into TorchCodec, which now runs on CPU and CUDA. TorchVision and TorchAudio are narrowed to focus on their transforms, with models, datasets, and pipelines no longer under active development. All three libraries are now ABI stable and no longer need rebuilding for each PyTorch release.

  4. Liquid AI · new models on Hugging FaceAI score44

    LiquidAI releases d1-omni-600M, a 600M decision model for text, image and audio

    AILiquidAI has released d1-omni-600M on Hugging Face, a 587M-parameter model that answers named yes/no, choice and score questions over text, images or up to 30 seconds of speech in a single forward pass. It returns typed answers with zero output tokens by reading the model's distribution over options, and is built on LFM2.5-Encoder-350M with a 16,384-token context length. The model is not a chat model and does not generate text.

  5. GitHub Blog · AI & MLAI score63

    GitHub releases ReviewBench, an open benchmark for AI code review agents

    AIGitHub has released ReviewBench, an open benchmark for evaluating AI code review agents on 219 public pull requests across 19 languages. The benchmark reports grounded and augmented precision, recall, and F1 metrics, and its dataset, rubric, and judge are publicly available. GitHub says ReviewBench predicted the direction of a Copilot code review ensemble experiment's production results before A/B testing.

    Why it matters: The post explains how ReviewBench was built and validated, and reports an offline-to-production comparison that shows how well a benchmark predicts real experiment outcomes.

  6. PyTorch BlogAI score24

    PyTorch's Accelerator Working Group Standardizes Hardware Backend Integration in H1 2026

    AIThe PyTorch Accelerator Integration Working Group released updates on its H1 2026 progress toward standardizing how new hardware connects to the framework. Key workstreams include the Cross-Repository CI Relay (CRCR), which automatically reports downstream backend test results to a shared dashboard, and refactored test suites that decouple PyTorch's 600,000-plus tests from specific accelerators.

  7. NVIDIA BlogAI score41

    AI Tools From NVIDIA Inception Startups Target Breast Cancer Screening, Diagnosis and Treatment Gaps

    AIiSono Health's FDA-cleared ATUSA wearable 3D ultrasound captures a breast volume in about two minutes per breast, compared with up to 45 minutes for handheld ultrasound, and is commercially available through partner clinics in several U.S. states. Whiterabbit.ai's FDA-cleared WRDensity software automatically assesses breast density from mammograms, while Ataraxis AI is building models that predict treatment response from digital pathology slides.

  8. Cloudflare Blog · AIAI score40

    Cloudflare Birthday Week 2026 unveils cf CLI, EmDash CMS, and post-quantum tools

    AICloudflare announced 46 products and updates during Birthday Week 2026, including the cf CLI for the entire Cloudflare API and EmDash, an open-source Astro-based serverless CMS whose plugins run in isolated Worker sandboxes. The company also said it plans to become a public certificate authority that issues free Merkle Tree Certificates for post-quantum authentication.

  9. Microsoft AI BlogAI score23

    Microsoft and NVIDIA release Sovereign AI white paper on control and choice

    AIMicrosoft and NVIDIA have co-developed a Sovereign AI white paper offering a framework built on control, choice, flexibility, and resilience for AI workloads. Microsoft defines sovereign AI as designing, deploying, and operating AI workloads under defined controls for data, access, governance, infrastructure, and operations. The framework is intended to help leaders decide the level of control each workload needs.

  10. ElevenLabs BlogAI score40

    How audio transcription with timestamps and event tagging works in Scribe

    AIA native word-level transcription model outputs structured, timestamped arrays of word, spacing, and audio_event tokens directly from audio input, without a secondary forced-alignment pass. Audio events such as laughter or applause are tagged separately, which the source says helps with captioning, searchable archives, and highlight identification. The source notes Scribe's word-level transcription supports up to 5 independently transcribed channels.

  11. Liquid AI · new models on Hugging FaceAI score67

    Liquid AI releases d1-3B, a 3B multimodal decision model for edge deployment

    AILiquid AI has released d1-3B, a 3B parameter multimodal model post-trained to return calibrated, typed answers to yes/no, choice, and score questions in one forward pass. The source reports a Decision Index 0.2.1 score of 48.57, the highest among models under 10B in its table, and 8 ms per decision on an NVIDIA RTX 4090.

    Why it matters: The source gives benchmark scores against named peer models and edge latency figures across several hardware targets, helping readers judge fit for on-device decision pipelines.

Oct 4

Oct 4Sun
  1. PromptArmor Threat IntelligenceAI score47

    Databricks Genie Code Malicious Skill Enables Phishing and Data Exfiltration

    AIPromptArmor reports that a malicious Skill can make Databricks Genie Code display a phishing modal and exfiltrate tenant data without human approval. The attack exploits Skills loaded from users' personal workspaces and a display interface that lacks egress controls, and Databricks, after disclosure on August 16, 2026, said users are responsible for ensuring uploaded Skills contain no malicious content.

  2. Apple Machine Learning ResearchAI score22

    Apple Study Examines How Users Negotiate Ontological Boundaries in Personal Sensing Systems

    AIApple and Stanford researchers built two open-ended probes using a Wizard of Oz technique so participants could train personalized machine learning systems on phenomena they defined themselves. In a week-long exploratory study, participants identified four sites where ontological boundaries were negotiated: the boundaries of a phenomenon, the subject as part of relations, signal versus noise, and the objectivity of data. The paper offers starting points for supporting boundary negotiation through design.

  3. Liquid AI BlogAI score70

    Liquid AI releases d1 decision model with image input support

    AILiquid AI introduces d1, its first decision model, now accepting both text and images. The company says d1 matches or beats GPT-6.1 Sol on four of six tested applications, at 19x to 200x lower cost and with faster answers on every task. d1 is available on the Liquid AI API and through Vercel and OpenRouter, with text-only support on those two platforms for now.

    Why it matters: The post gives benchmark comparisons against named models along with per-token pricing and latency figures, which makes the cost and speed tradeoff checkable.

  4. OpenRouter BlogAI score44

    Server-Side Code Execution Tools for AI Agents, Compared

    AIOpenRouter's shell and bash tools, along with those from OpenAI and Anthropic, run an agent's commands in provider-managed sandboxes during the same API request, so developers don't provision or patch containers. OpenRouter's tools are in beta, with sandbox time billed at $0.0001 per second and a 30-second minimum for a new or sleeping container. The article compares the four providers and notes that self-run sandboxes remain better for custom base images, GPU work, or multi-hour sessions.

  5. Epoch AIAI score62

    OpenAI researchers' coding-agent usage is doubling about monthly, Epoch AI reports

    AIOpenAI researchers' daily coding-agent usage, valued at API prices, rose from under $1 in January 2026 to $601 for the median researcher by mid-August. The 90th-percentile researcher reached over $7,000 per day, and both groups show doubling times of roughly one month. Epoch notes these are API-list values, not OpenAI's internal costs.

    Why it matters: The figures show internal coding-agent usage growing fast enough to matter for research cost, though they measure API-list value rather than OpenAI's actual spending.

Oct 3

Oct 3Sat
  1. Claude Code · GitHub ReleasesAI score7

    Claude Code v2.1.289 fixes plugin, sandbox, and terminal rendering bugs

    AIClaude Code v2.1.289 fixes a series of bugs, including deny and ask rules being bypassed on nested parts of compound shell commands on managed machines. It also fixes terminal freezes on short code blocks with unclosed tags, Read deny rules not applying to files reached through symlinks in the IDE, and plugin panes that drew nothing for certain link formats. A change to claude auth status that may have increased sign-outs in VSCode was reverted.

  2. Hugging Face BlogAI score67

    Microsoft ThinkingBox grades AI agents on database state across 20 repeated runs

    AIMicrosoft and Hugging Face released ThinkingBox, a benchmark that grades AI agents on the terminal backend state and side effects they leave behind rather than their final responses. Each of 507 stateful business tasks runs 20 times from a clean backend, and the post reports pass@1, pass@20, and observed 20/20 counts, plus cost per successful and per dependable task across 18 models. The harness and dataset are available on Hugging Face, with the OpenEnv interface for running evaluations.

    Why it matters: The post shows why checking the database state, not tool calls or final replies, exposes agent failures, and gives a repeat-run method for judging reliability.

  3. IndexTeam (Bilibili) · new models on Hugging FaceAI score22

    Index-Echo-S2ST-9B-FP4 released as NVFP4 quantized speech translation model

    AIIndexTeam released Index-Echo-S2ST-9B-FP4, an NVFP4 (W4A4) quantization of the Index-Echo-S2ST-9B speech-to-speech translation model, with only its text LLM backbone quantized. Perplexity rose from 3.8218 to 3.9650 (+3.75%) on a fixed corpus, while zh→en and en→zh outputs were semantically equivalent, and full FP4 speedup requires an NVIDIA Blackwell GPU.

  4. IndexTeam (Bilibili) · new models on Hugging FaceAI score27

    Index-Echo-S2ST-2B FP4 Quantized Speech-to-Speech Translation Model Released on Hugging Face

    AIIndexTeam released Index-Echo-S2ST-2B-FP4, an NVFP4 (W4A4) quantized version of the Index-Echo-S2ST-2B speech-to-speech translation model, with only the text LLM backbone quantized and the audio components kept in BF16. On a fixed corpus, perplexity rose from 5.9332 to 6.4980 (+9.52%), while zh->en and en->zh generations matched the original. Full FP4 acceleration requires an NVIDIA Blackwell GPU, and the model loads via compressed-tensors in vLLM or transformers.

  5. IndexTeam (Bilibili) · new models on Hugging FaceAI score20

    IndexTeam releases NVFP4 quantized Index-Echo-S2TT-9B speech translation model

    AIIndexTeam published an NVFP4 (W4A4) quantized version of its Index-Echo-S2TT-9B speech-to-text translation model, quantizing only the text LLM backbone while keeping the audio tower and other components in BF16. On an NVIDIA A100, perplexity rose from 3.4155 to 3.5113 (+2.81%), with zh->en and en->zh outputs semantically equivalent under greedy decoding. Full FP4 speedup requires an NVIDIA Blackwell GPU, while older GPUs get only memory reduction.

  6. IndexTeam (Bilibili) · new models on Hugging FaceAI score20

    IndexTeam releases NVFP4 quantized Index-Echo-S2TT-2B speech translation model

    AIIndexTeam has published an official NVFP4 (W4A4) quantized version of its Index-Echo-S2TT-2B speech-to-text translation model on Hugging Face. Only the text LLM backbone is quantized, while the audio tower, connector, and speech-synthesis components remain in BF16. Perplexity rises 5.80%, from 4.8772 to 5.1599, on a fixed corpus, and full FP4 speedup requires an NVIDIA Blackwell GPU.

  7. IndexTeam (Bilibili) · new models on Hugging FaceAI score22

    Index-Nailong-9B-FP4 NVFP4 quantized translation model released on Hugging Face

    AIIndexTeam released Index-Nailong-9B-FP4, an official NVFP4 (W4A4) quantization of the Index-Nailong-9B multilingual translation model, which covers 150 languages. In a validation on an NVIDIA A100 against the BF16 checkpoint, perplexity rose 3.10% (2.4339 to 2.5094), and zh-en and en-zh outputs were semantically equivalent. Full FP4 compute acceleration requires an NVIDIA Blackwell GPU, while older GPUs get memory savings only; the FP8 build is recommended for Hopper and Ampere.

  8. IndexTeam (Bilibili) · new models on Hugging FaceAI score29

    Index-Nailong-2B-FP4 Released as NVFP4 Quantized Translation Model

    AIIndexTeam has released Index-Nailong-2B-FP4, an official NVFP4 (W4A4) quantization of its Index-Nailong-2B multilingual translation model, which supports 150 languages. The checkpoint keeps lm_head, embeddings, and MoE router gates in BF16, and a perplexity test on a fixed corpus rose from 3.2806 to 3.4998 (+6.68%), while zh->en and en->zh outputs matched BF16 semantically. Full FP4 acceleration requires an NVIDIA Blackwell GPU; on Hopper or Ampere, vLLM provides only memory savings, so the FP8 build is recommended.

  9. IndexTeam (Bilibili) · new models on Hugging FaceAI score23

    Index-Homura-9B-FP4 released with NVFP4 quantization for translation model

    AIIndexTeam released Index-Homura-9B-FP4, an official NVFP4 (W4A4) quantization of the Index-Homura-9B translation model from the Index-Translate family. On a fixed corpus, perplexity rose from 2.5386 in BF16 to 2.6245, a 3.38% increase, and zh->en generations matched the original. Full FP4 compute acceleration requires an NVIDIA Blackwell GPU, while older GPUs get only weight-only memory savings and the FP8 build is recommended for them.

  10. IndexTeam (Bilibili) · new models on Hugging FaceAI score29

    Index-Homura-2B-FP4 released as NVFP4 quantized translation model

    AIIndexTeam released Index-Homura-2B-FP4, an official NVFP4 (W4A4) quantization of its Index-Homura-2B multilingual translation model, which supports 150 languages. The quantized checkpoint shows a 5.73% perplexity increase over the BF16 original (3.5011 to 3.7017) on a fixed corpus, and its zh-en and en-zh outputs are semantically equivalent under greedy decoding. Full FP4 acceleration requires an NVIDIA Blackwell GPU, while the source recommends the FP8 build for Hopper and Ampere hardware.

Oct 2

Oct 2Fri
  1. Baseten BlogAI score70

    Baseten's agent-built VibeQwen engine beats vLLM on Qwen-3.6 decode speed

    AIBaseten tested the MetaInfer skills-only approach by having Claude Code build an inference engine, VibeQwen, for Qwen-3.6-35B-A3B in NVFP4 on a single B200. On single-stream text, VibeQwen decoded 90% faster than a tuned vLLM 0.25.1 deployment (1,792 vs. 943 TPS) and cut time to first token from 28 ms to 12 ms, with a 71% throughput gain at concurrency 32. The author notes this was an outcome-focused run that allowed some numerically different outputs as long as accuracy stayed at or above the BF16 baseline.

    Why it matters: The post tests a skills-only inference engine method on a real model and states the speed and accuracy constraints used, helping readers judge how far such automated optimization can be trusted.

  2. Claude Code · GitHub ReleasesAI score38

    Claude Code v2.1.288 is released with fixes and new controls

    AIAnthropic released Claude Code v2.1.288, adding $.ui.selection() for mods, a built-in gh api for cloud sessions without the GitHub CLI, and --max-findings for /code-review. The release also fixes many issues, including mid-response API timeouts, resume and compaction bugs, and auto mode denials and model switching on Bedrock and Mantle.

  3. PyTorch BlogAI score47

    Helion Linear Backend Boosts vLLM Hopper GPU Inference Throughput Over CUTLASS and DeepGEMM

    AIThe vLLM team integrated Helion, a PyTorch-native kernel DSL, into vLLM's linear backend, using per-shape autotuning to select among Standard GEMM, Split-K, and Swap-AB variants. On NVIDIA Hopper GPUs, the Helion backend outperformed the default CUTLASS and DeepGEMM backends across the evaluated models, with more than 10% throughput gains for some workloads. The work focuses on FP8 and INT8 quantized GEMM.

  4. PyTorch BlogAI score24

    PyTorch Certified Associate Gets New Four-Module Certification Pathway

    AIThe Linux Foundation Education has launched a PyTorch Certified Associate (PTCA) Certification Pathway that combines four self-paced learning modules with the PTCA exam. The pathway includes 15–17 hours of self-paced learning and hands-on labs covering tensors, data handling, model development, and performance optimization. The source recommends additional hands-on practice before taking the exam.

  5. MIT News · AIAI score14

    MIT's Cathy Wu Uses Reinforcement Learning to Tackle Transportation Challenges

    AIMIT associate professor Cathy Wu is applying machine learning and reinforcement learning (RL) to design safer, more efficient transportation systems. Her team found RL can train effectively on about 10 percent of related problems, and a selection algorithm improved training efficiency by up to 30 times. Her recent work estimates eco-driving measures could cut vehicle emissions by 11 to 22 percent.

  6. GitHub Copilot ChangelogAI score34

    Copilot code review gains API access and Balanced default effort level

    AIGitHub Copilot code review can now be requested through the REST and GraphQL APIs, with an optional review effort level set per request. Balanced became the default review effort level for new and existing repositories and organizations as of September 28, 2026, while users who explicitly selected Lite keep that setting. The changes are generally available to Copilot Pro, Pro+, Max, Business, and Enterprise plans.

  7. Epoch AI · The Epoch BriefAI score62

    Epoch AI estimates 2026 compute could run hundreds of millions of AI agents

    AIEpoch AI estimates that compute built from projected 2025 to 2027 high-bandwidth memory shipments could support tens to hundreds of millions of frontier AI agents, or billions of cheaper ones. Running nonstop, the top-tier agents would match the working hours of 140 million to 700 million full-time employees, and the central DeepSeek V4 Pro estimate of about 1.9 billion agents would match 8 billion workers.

    Why it matters: The estimate converts memory shipments into agent capacity and revenue ranges, showing how hardware supply could translate into labor and sales if demand keeps up.

  8. MIT News · AIAI score29

    Tech Worker Movement Against Industry Power Faces Backlash, New Book Chronicles

    AIFormer tech workers JS Tan and Clarissa Redwine have published "Against Tech Oligarchy: Worker Resistance in the World's Most Powerful Industry" (Haymarket Books, 2026), chronicling how tech employees organized over the past decade. The book traces early successes, including Google's 2018 decision not to renew its Project Maven Pentagon contract after employee protests. It also argues that rising interest rates, job-security fears, and agentic AI coding tools have weakened worker leverage.

  9. Hugging Face BlogAI score70

    Ai2 open-sources AstaBrief 8B, a fast model for generating cited research reports

    AIAi2 released AstaBrief 8B, an open-weights model that turns a research question and retrieved literature excerpts into a cited report, along with its training data. The model runs as Fast mode in Asta, averaging 51.1 seconds per report versus 178.5 seconds for Thinking mode, about 3.5x faster. The post also describes filtering synthetic training data by citation density and building DPO pairs judged by two models that agreed.

    Why it matters: The post explains how supervised fine-tuning, preference data, and citation-density filtering were used to build a cited-report model, which is useful for teams training their own models.

  10. GitHub Blog · AI & MLAI score23

    Three Skills Developers Need as AI Changes Their Work

    AIAI is changing developer work, and the article recommends three skills: directing AI agents, reviewing AI output instead of trusting the first answer, and using saved time for judgment-heavy problems such as customer needs and tradeoffs. It cites GitHub Copilot's built-in Rubber Duck agent, which uses a second model to critique plans, code, and tests. The author argues that developers remain responsible for outcomes while AI handles more implementation.