Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Oct 9

Oct 9Fri
  1. a16z NewsBlogAI score40

    a16z leads investment in TypeSafe AI, maker of Jev System One model

    AIa16z says it is leading an investment in TypeSafe AI, whose Jev model hands decisions to code as typed values and reached 1 trillion tokens generated three days after launch. The company says Jev costs roughly 1/100 to 1/500 of frontier models and runs 100x faster on classification tasks at comparable accuracy. TypeSafe says 25% of the Fortune 500 have integrated Jev.

  2. a16z NewsBlogAI score33

    Prediction markets show no partisan bias in election pricing, NBER study finds

    AIA preliminary NBER working paper by Prof. Zitzewitz, covering over 100 years of prediction markets, finds no statistically significant bias by political affiliation, gender, race, or age. The only exception is non-US elections, where markets appear to overrate right-leaning candidates, but that result is not statistically significant. Separately, prediction markets had Flávio Bolsonaro's Brazilian presidential rise about three weeks before his first-round win.

  3. O'Reilly RadarBlogAI score40

    US AI oversight debate, OpenAI Dots, and Gemini 4 Argon featured in This Week in AI

    AIThe Trump administration announced a voluntary agreement with major AI companies calling for internal safety monitoring, external audits, and independent board reviews, and the Federal Trade Commission launched an investigation into OpenAI, Anthropic, and other AI companies over potential consumer risks. OpenAI released Dots, a proactive assistant that retains context, works across applications, and acts without waiting for prompts. Google says Gemini 4 Argon can generate up to a million output tokens in a single response.

  4. 🚨 AI News | TestingCatalogXAI score41

    Pine AI launches Pine Computer, a cloud runtime for agentic tasks

    AIPine AI launched Pine Computer, a cloud computer, harness, and runtime layer built for agentic tasks. On the publisher's SaaS-Bench v1.1, it posts a 78.3% checkpoint score against 74.3% for Opus 5 with Claude Code, but completes fewer whole tasks, 27.4% against 31.1%. Instead of simulating clicks and screenshots, it reads web pages as structured data, and access is through a private beta waitlist.

    Image from @testingcatalog's post
  5. 🚨 AI News | TestingCatalogXAI score62

    Anthropic moves dynamic workflows in Claude Managed Agents into public beta

    AIAnthropic has expanded dynamic workflows in Claude Managed Agents into a public beta, according to Testing Catalog. Users can configure their agents for multiagent orchestration, with Claude planning and operating a fleet of agents to achieve a goal. The post also links a video from Anthropic's ClaudeDevs account, which the author describes as a new SWE norm.

    Video from @testingcatalog's post
  6. SantiagoXAI score44

    Pine launches agentic cloud computers with built-in AI agents

    AIPine has released a cloud computer service with a built-in AI agent that applications can control through its SDK. Developers give the agent a plain-English task, and it can use a browser, files, and a shell while the app receives notifications and final outputs. Pine's Stanley Wei says the computer is built for AI rather than humans.

    Video from @svpino's post
  7. Vaibhav (VB) SrivastavXAI score4

    OpenAI rolls out invites for DevDay Exchange Berlin, Paris, and London

    AIOpenAI says invites for its DevDay Exchange events in Berlin, Paris, and London are rolling out now. Recipients are asked to register as soon as possible to secure a spot. People still waiting can reply with their city, what they are building or exploring, and why they want to attend.

    Image from @reach_vb's post
  8. Alexandr WangXAI score12

    Alexandr Wang discusses AI safety in a conversation with Cleo Abram

    AIAlexandr Wang calls Cleo Abram brilliant and says their conversation on AI safety is one of his favorites. He also says he included a subtle dig at some well-known people during the interview. The interview covers what Muse can do, when it can be trusted, and how to make sure the technology goes right.

  9. LlamaIndex 🦙OfficialAI score12

    LlamaParse keeps nested tables and values intact in earnings decks

    AILlamaIndex says its LlamaParse keeps all 18 values from Micron's latest earnings deck under the correct headers, despite nested tables with two business units and repeated row names. The post presents this as a sample document rather than a benchmark result.

    Image from @llama_index's post
  10. Stanford HAIOfficialAI score16

    Stanford HAI hosts Queen Elizabeth Prize for Engineering discussion on AI

    AIFei-Fei Li, Stanford HAI founding director and QEPrize laureate, joined former Stanford president John Hennessy and Lord Vallance of Balham to discuss sustaining engineering breakthroughs in the age of AI. The event took place during the Queen Elizabeth Prize for Engineering's visit to Stanford this week.

    Image from @StanfordHAI's post
  11. merveXAI score28

    Hugging Face lets agents train Qwen3.8-27B on Nebius GPUs

    AIHugging Face launches an arena where users bring their own agent, which gets Nebius GPUs to build RL environments that improve Qwen3.8-27B across eight domains. The arena runs on PostTrainArena from BenchFlow, with compute from Nebius. Setup requires only a few steps through the linked OpenEnv Arena space.

    Video from @mervenoyann's post
  12. ClaudeDevsOfficialAI score60

    Claude Managed Agents adds dynamic workflows in public beta

    AIAnthropic's ClaudeDevs account announces that dynamic workflows for Claude Managed Agents are now available in public beta. The feature is a new type of multiagent orchestration in which a lead agent writes a plan that runs across many agents in phases, then combines their results at the end.

    Why it matters: The post describes how a lead agent plans work across many agents in phases and merges their results, a structure useful for understanding complex agent orchestration.

    Video from @ClaudeDevs's post
  13. Perplexity DevelopersOfficialAI score34

    Perplexity releases cookbook for a browser agent using the Decisions API

    AIPerplexity Developers says its new cookbook builds a browser agent that sends a screenshot and questions to pplx-decider-v1.1-27b through the Decisions API, which accepts text and image inputs. The developer's code converts the returned probabilities into clicks, scrolls, and stops.

  14. elvisXAI score34

    Elvis Saravia urges builders to focus on agent harnesses and environments

    AIElvis Saravia says AI models are already smart, but they need better harnesses and environments, with major cost implications. He recommends reading a report on how Pine Computer can help teams, and says he will test it himself and share more later. The quoted post from Stanley Wei argues that real-world AI tasks remain slow, expensive and unreliable because AI runs on computers built for humans, and announces Pine Computer.

    Image from @omarsar0's post
  15. Mike KnoopXAI score43

    Knoop says ARC-AGI-2 is much harder than ARC-AGI-1

    AIMike Knoop says ARC-AGI-2 is far harder than ARC-AGI-1, even amid rapid progress on math. He adds that the final open solutions will be useful artifacts to study, and notes that the Kaggle Grand Prize bonus threshold of 85% has been reached this year, the final year for ARC-AGI-2 on Kaggle.

  16. dexXAI score40

    Dex Horthy says small tasks should skip heavy planning workflows

    AIDex Horthy says the share of tasks that can be one-shot without strict process has grown, but alignment, grilling, and planning workflows still matter. He argues that heavy planning on small tasks makes developers feel slower, and predicts tools will add escape hatches so humans or models can decide to ship directly. He adds that as model capabilities improve, the "smart zone" has grown to roughly 200k–400k tokens, and HumanLayer is prototyping research-to-implement and research-to-short-design-to-implement workflows.

  17. LangChainOfficialAI score16

    LangChain opens Interrupt session recordings on-demand

    AILangChain says its Interrupt archive is available, with every session watchable on-demand at Context from the quoted post: LangChain simplified agent authentication, memory, and channels, and added web search as a pre-built tool for Managed Deep Agents.

  18. Replit ⠕OfficialAI score22

    TikTok Ads MCP lets users run TikTok Ads from Replit

    AIReplit's X account shares a showcase of TikTok Ads MCP, which runs TikTok Ads from Replit. The post is a broadcast link with no further details on features, pricing, or availability.

  19. SiliconANGLE · AINewsAI score40

    OpenAI reports $18 billion less revenue, hitting AI stocks Thursday

    AIOpenAI told investors it had $18 billion less revenue than the $68 billion it reported last month, and AI-linked stocks including Nvidia, CoreWeave, Oracle and Nebius fell Thursday. Google debuted a Gemini assistant that can act autonomously, generate code and complete work across web, mobile and desktop. Anthropic released Claude Haiku 5.5 and halved Sonnet 5.5 cache read prices.

  20. The Guardian · AINewsAI score38

    Britain's advertising directors fear AI will cut off the next generation's training

    AIWPP has opened a flagship AI-enabled production facility in east London as part of a £600m WPP Production business, with CEO Cindy Rose saying AI is ushering in a "golden age of marketing". Some AI-led ads can be made up to 60% cheaper, one industry source says, while critics including Luke Scott warn that a lost generation of directors may miss the hands-on training that built careers like Ridley Scott's.

  21. Adam RobinsonXAI score24

    MoltSets launches $27/mo unlimited B2B contact database API, claims to beat ZoomInfo and Apollo

    AISolo founder launches MoltSets, an API-only B2B contact database priced at $27/mo, claiming unlimited access, high quality, high coverage, high rate limits, and low pricing. The post says LinkedIn profiles across 47 million slugs are re-scraped every 3 days and emails re-validated every 3 days, with 15 cents per real-time carrier-verified mobile number through bolt-on plans.

  22. elvisXAI score60

    Meta researchers propose agent plasticity to measure self-improvement efficiency

    AIResearchers from UC Berkeley, Meta Superintelligence Labs, and other institutions introduce agent plasticity, the gain on held-out tasks per dollar of learning cost, with model weights frozen. The paper reports that in chess, Go, and Hex, Claude Fable 5 reaches the highest final score while GPT-5.6 Sol gains the most per dollar, and in NetHack only Claude Opus 5.5 improves significantly.

    Image from @omarsar0's post
  23. 404 MediaNewsAI score11

    Behind the Blog: 404 Media discusses AI and spirituality

    AI404 Media's Behind the Blog column discusses AI and spirituality, with Jason saying the outlet writes about AI's current capabilities and harms rather than dismissing it outright. He says reporters sometimes test AI tools while working on stories to write from an informed perspective. The excerpt does not say more about the spirituality discussion.

  24. TechCrunch · AINewsAI score44

    a16z's Olivia Moore says consumer AI revenue is mostly prosumer and many categories lack AI apps

    AIAndreessen Horowitz partner Olivia Moore released a report on the top 100 consumer AI apps, finding ChatGPT still leads by a wide margin while smaller players like Suno and ElevenLabs show staying power. Moore says almost all AI revenue comes from subscriptions and token usage, and that most consumer AI is prosumer AI. The report finds no top-100 entrants in social, dating, marketplace, retail, travel, finance, or health categories.

  25. Mike KnoopXAI score62

    Tufa Labs hits 88.06% on ARC-AGI-2, clearing the Kaggle bonus threshold

    AIMike Knoop says the 85% Grand Prize bonus threshold has been reached on Kaggle. The ARC Prize 2026 leaderboard lists Tufa Labs first at 88.06%, followed by Rabbithole at 80.56% and Yi-Chia Chen at 77.22%. Knoop says this will be the final year for ARC-AGI-2 on Kaggle and expects an open-source, low-cost, offline reproducible solution and model.

  26. Ethan MollickXAI score23

    Google's post-Gemini 4 challenge is product integration, Mollick argues

    AIEthan Mollick says Google's main challenge after Gemini 4 is what it does with a strong model. He argues that Anthropic and OpenAI are moving toward a single interface for many tasks using orchestrator agents. He says the fragmented products of the Gemini 3 era will not work for what comes next.