Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 9

TodayOct 9Fri
  1. The DecoderNewsAI score62

    Anthropic launches a free AI scanner for open-source projects

    AIAnthropic has launched Cyber Mission, a long-term program to protect critical infrastructure and open-source software from cyberattacks. A free OSS AI scanner will regularly check open-source projects, flag and explain vulnerabilities, and suggest patches. Anthropic expects over 90 percent accuracy, but reports ship without human review and may contain errors.

  2. The Robot ReportNewsAI score36

    Boston Dynamics details its redesigned four-fingered humanoid robot hand

    AIBoston Dynamics says its new humanoid hand has four fingers, 13 degrees of freedom, and direct actuation, and it is designed for mass manufacturing. The company dropped the pinky finger and reduced the gripper size, and built the hand for high-fidelity simulation to enable sim-to-real reinforcement learning. Alberto Rodriguez, director of robot behavior for Atlas, says the earlier hands could already lift more than 100 lb., so the new design focuses on reliability and tool use.

  3. Baseten BlogOfficialAI score38

    Baseten launches Project Beacon with Goodfire AI for inline safety controls on models

    AIBaseten announces Project Beacon with Goodfire AI, adding inline safety controls to model inference. Goodfire's activation-based monitors read a model's internal activations during generation, so policies can flag unsafe events before output reaches a user or tool. Baseten plans to release the capabilities over the next several months with selected models and early partners.

  4. Mistral AI · new models on Hugging FaceOfficialAI score47

    Mistral releases Voxtral Mini 4B Realtime Arabic speech-to-text model

    AIMistral releases Voxtral Mini 4B Realtime Arabic, a streaming speech-to-text model for Arabic dialects and Modern Standard Arabic under the Apache 2.0 License. The model has about 4.4 billion parameters, is fine-tuned from Voxtral-Mini-4B-Realtime-2602, and reaches an average 8.82% Character Error Rate across seven Arabic benchmarks at a 480 ms transcription delay. It can be run with vLLM or Transformers 5.2.0 or later.

  5. Mistral AI · new models on Hugging FaceOfficialAI score32

    Mistral releases LIDstral-Arabic, a language and dialect classifier for Arabic-script text

    AIMistral AI has released LIDstral-Arabic, a fast classifier that identifies Modern Standard Arabic, Arabic dialects, and non-Arabic languages written in Arabic script across 51 classes. On Moroccan Darija, it scores 88.67% F1, versus 72.65% for LahjatBERT ALDi CL and 71.28% for GlotLID v3, across 84,870 evaluation examples. The model runs on CPU and is available from a private Hugging Face repository under Apache 2.0.

  6. elvisXAI score62

    StepFun's Step 5 Preview targets long coding agent runs

    AIElvis Saravia says he has tested StepFun's Step 5 Preview as a coding agent since early access and found that it checks its own work and stops when tasks are done. The post says the model is built for engineering tasks such as bug fixing, multi-file features, and refactoring, plus frontend generation and financial report output.

    Image from @omarsar0's post
  7. Microsoft CopilotOfficialAI score30

    Microsoft unveils new Copilot Home combining Chat and Cowork

    AIMicrosoft says the new Copilot Home brings Chat and Cowork into one experience, so users can move from thinking to doing without losing context. Microsoft Copilot EVP Jacob Andreou explains why Home is the new starting point for work.

    Video from @MSFTCopilot's post
  8. LangChain BlogOfficialAI score40

    LangChain adds emoji reactions to Managed Deep Agents Slack channels

    AILangChain's Managed Deep Agents v0.9 adds a reactions attribute for Slack channels that accepts either an emoji string or a callable returning one. The article shows a function that returns a bug emoji when a message contains "broken" and eyes otherwise. It also shows a TypeSafe Classifier that picks from a seven-emoji vocabulary and falls back to eyes below 25% confidence.

  9. a16z NewsBlogAI score40

    a16z leads investment in TypeSafe AI, maker of Jev System One model

    AIa16z says it is leading an investment in TypeSafe AI, whose Jev model hands decisions to code as typed values and reached 1 trillion tokens generated three days after launch. The company says Jev costs roughly 1/100 to 1/500 of frontier models and runs 100x faster on classification tasks at comparable accuracy. TypeSafe says 25% of the Fortune 500 have integrated Jev.

  10. a16z NewsBlogAI score33

    Prediction markets show no partisan bias in election pricing, NBER study finds

    AIA preliminary NBER working paper by Prof. Zitzewitz, covering over 100 years of prediction markets, finds no statistically significant bias by political affiliation, gender, race, or age. The only exception is non-US elections, where markets appear to overrate right-leaning candidates, but that result is not statistically significant. Separately, prediction markets had Flávio Bolsonaro's Brazilian presidential rise about three weeks before his first-round win.

  11. 🚨 AI News | TestingCatalogXAI score62

    Anthropic moves dynamic workflows in Claude Managed Agents into public beta

    AIAnthropic has expanded dynamic workflows in Claude Managed Agents into a public beta, according to Testing Catalog. Users can configure their agents for multiagent orchestration, with Claude planning and operating a fleet of agents to achieve a goal. The post also links a video from Anthropic's ClaudeDevs account, which the author describes as a new SWE norm.

    Video from @testingcatalog's post
  12. merveXAI score28

    Hugging Face lets agents train Qwen3.8-27B on Nebius GPUs

    AIHugging Face launches an arena where users bring their own agent, which gets Nebius GPUs to build RL environments that improve Qwen3.8-27B across eight domains. The arena runs on PostTrainArena from BenchFlow, with compute from Nebius. Setup requires only a few steps through the linked OpenEnv Arena space.

    Video from @mervenoyann's post
  13. ClaudeDevsOfficialAI score60

    Claude Managed Agents adds dynamic workflows in public beta

    AIAnthropic's ClaudeDevs account announces that dynamic workflows for Claude Managed Agents are now available in public beta. The feature is a new type of multiagent orchestration in which a lead agent writes a plan that runs across many agents in phases, then combines their results at the end.

    Why it matters: The post describes how a lead agent plans work across many agents in phases and merges their results, a structure useful for understanding complex agent orchestration.

    Video from @ClaudeDevs's post
  14. Perplexity DevelopersOfficialAI score34

    Perplexity releases cookbook for a browser agent using the Decisions API

    AIPerplexity Developers says its new cookbook builds a browser agent that sends a screenshot and questions to pplx-decider-v1.1-27b through the Decisions API, which accepts text and image inputs. The developer's code converts the returned probabilities into clicks, scrolls, and stops.

  15. Replit ⠕OfficialAI score22

    TikTok Ads MCP lets users run TikTok Ads from Replit

    AIReplit's X account shares a showcase of TikTok Ads MCP, which runs TikTok Ads from Replit. The post is a broadcast link with no further details on features, pricing, or availability.

  16. elvisXAI score60

    Meta researchers propose agent plasticity to measure self-improvement efficiency

    AIResearchers from UC Berkeley, Meta Superintelligence Labs, and other institutions introduce agent plasticity, the gain on held-out tasks per dollar of learning cost, with model weights frozen. The paper reports that in chess, Go, and Hex, Claude Fable 5 reaches the highest final score while GPT-5.6 Sol gains the most per dollar, and in NetHack only Claude Opus 5.5 improves significantly.

    Image from @omarsar0's post
  17. Ai2OfficialAI score22

    Ai2 replaces its GPU scheduler after idle jobs hoarded capacity

    AIAi2 says its old scheduler made every scheduled workload eventually run at HIGH priority. Researchers kept idle jobs running to reserve GPUs for experiments, because the incentives rewarded holding capacity even with no active work.

  18. Ai2OfficialAI score25

    Ai2's new scheduler delivers 98% of owed GPU hours in 30-day test

    AIAi2 reports that over a 30-day test of its new scheduler, teams received 98% of the GPU hours they were owed, based on actual demand. Cluster occupancy stayed at 98%, and spare capacity went to interruptible work without drawing down team budgets.

  19. The Algorithmic BridgeBlogAI score40

    Meta's AI comeback follows heavy Anthropic Claude spending and a new Muse Spark model

    AIMeta spent heavily on Anthropic's Claude models, with internal use reaching up to 60,000 employees and a projected $10 billion yearly spend, according to The Algorithmic Bridge. The author says Meta then released Muse Spark, which scored 52 on the Artificial Analysis intelligence benchmark, on par with Claude Opus 4.6.

  20. Ai2 (Allen Institute for AI)OfficialAI score46

    Ai2 describes GPU time budgets that replaced its priority-based cluster scheduler

    AIAi2's AI Infrastructure team replaced its priority-based scheduler for GPU clusters with GPU time budgets, hierarchical fair-share allocation, and a time-slicing contract. The team says the change moved debates over how much GPU time each research project deserves from case-by-case operational decisions into a transparent budgeting process. The clusters range from 88 to 1024 GPUs across NVIDIA H100, B200, and B300 hardware, and serve about 150 internal researchers.