Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 30

Sep 30Wed
  1. Sakana AIOfficialAI score33

    Sakana AI's David Ha argues the future of AI lies in orchestrators

    AISakana AI co-founder and CEO David Ha published a Nikkei Asia op-ed titled "The future of AI belongs to the orchestrators." He argues that ever-larger models face limits, as open models close the gap within months and frontier inference costs can exceed the hourly wage of the people they assist. He also contends that sovereignty means supply-chain strength, not national isolation.

  2. Guillermo RauchXAI score22

    Vercel AI Gateway rejects dubious token promos, prioritizing trustworthy providers and data privacy

    AIVercel says its AI Gateway turns down "free token" promotions from companies making dubious Zero Data Retention claims, prioritizing the best providers over provider count. The company argues it has the largest trustworthy view of global AI token flows, citing real usage from 400k+ paying customers, thousands of enterprises, and zero markup. It says it invests as heavily in legal, compliance, privacy, and back-office operations as in engineering to serve trillions of tokens daily.

  3. Hamel HusainXAI score38

    Hamel Husain Reviews Claude's New Auto Eval Plugin for Evaluations

    AIHamel Husain has published a longer review of a new Claude Auto Eval plugin after many users asked about it. He invites readers to share their experiences using the plugin and how it went for them. The plugin is part of Claude's ability to help build evaluations and hillclimb on them, as described by @ClaudeDevs.

    Image from @HamelHusain's post
  4. Max ZeffXAI score34

    Greg Brockman drops second $25M donation to AI super PAC

    AIOpenAI co-founder Greg Brockman is no longer making the second $25 million donation he had promised to the super PAC Leading the Future, according to a New York Times scoop. The commentary notes that AI regulation has become a major national issue far faster than many expected since Brockman's August 2025 commitment.

  5. indigoXAI score81

    Google's Gemini 4 Argon debuts with limited access pending US government approval

    AIGoogle has announced Gemini 4 Argon, initially available only to trusted cyber defenders through its Fairwind Program while US government approval is pending. The author says the model is aimed at long-running software engineering, enterprise knowledge work, and cybersecurity tasks, with a 1 million token output limit. The post also gives promotional pricing of $2 per million input tokens and $10 per million output tokens, rising to $4 and $20 afterward, alongside a benchmark comparison.

    Why it matters: The post places Gemini 4 Argon's benchmark table beside GPT-6 Astra and Claude models, showing where each leads across coding, knowledge work, and cybersecurity tasks.

    Image from @indigox's post
  6. Apple Machine Learning ResearchOfficialAI score46

    Minimal Coding Agent Matches Elaborate ML Engineering Harnesses on Autonomous Tasks

    AIUnder equal time budgets and the same frontier LLM backbone, a single session of a minimal-harness coding agent with read, write, and bash primitives matched open-source state-of-the-art autonomous machine learning engineering harnesses. Apple researchers found the added orchestration and retrieval machinery redundant in large-scale ablation studies, pointing to the backbone model as the main driver of performance. They conclude that hand-crafted harnesses around strong models yield poor returns on current MLE benchmarks.

  7. Apple Machine Learning ResearchOfficialAI score36

    RLTL;DR: Self-Improvement Through Internalized Self-Generated Feedback

    AIApple researchers introduced RLTL;DR, a reinforcement learning method in which an agent writes its own one-line insight after each failed attempt and learns to map tasks to those insights. On challenging tool-calling and coding datasets filtered to Pass@128 = 0, standard GRPO training of a Qwen 3.5 9B Thinking policy stayed at 0% to 1% Pass@1, while RLTL;DR reached 14–31% with insights in context and 12–13% without them at evaluation. A compact variant, SFTL;DR, trained on just 4k task-insight tuples recovered nearly the full performance of RLTL;DR.

  8. Latent.SpaceXAI score67

    OpenAI details agent stack with Computer Use, Dots, and Decisions API

    AILatent.Space shares a podcast episode in which OpenAI's AriX and Nikunj Handa discuss OpenAI's new agent stack. The episode covers Computer Use, Dots cloud computers for agents, and the Decisions API, which the speakers say went from idea to product in weeks.

    Video from @latentspacepod's post
  9. Google ResearchOfficialAI score40

    Google's science AI tops CDC flu hospital admission forecasts this season

    AIThe CDC announced that Google's science AI model ranked highest among 39 eligible models for forecasting flu-related hospital admissions during the 2025-26 flu season. Google's forecasts were built with Empirical Research Assistance, an AI tool that generates computational solutions across scientific fields.

    Image from @GoogleResearch's post
  10. Matei ZahariaXAI score36

    Databricks AI Decide runs decision models in SQL and Spark queries

    AIMatei Zaharia announced that users can run decision models at scale within SQL and Spark queries. The post's quoted background describes Databricks' ai_decide function, which turns text into structured decisions over governed data for batch and REST API use.

  11. Yi TayXAI score46

    Gemini 4 Argon launches, reportedly outperforming astra and fable on many tasks

    AIGoogle DeepMind introduced Gemini 4 Argon, a new frontier model built for coding, enterprise knowledge work, and cybersecurity defense, rolling out to trusted testers through its Fairwind Program. Yi Tay says Gemini 4 outperforms astra and fable on many tasks, though the post gives no benchmark figures.

  12. Google · Innovation & AIOfficialAI score46

    Google AI Flu Model Ranks First in CDC FluSight Hospitalization Forecasts

    AIA flu forecasting model built with Google AI ranked first among 39 eligible models in the CDC's FluSight 2025-26 season evaluation for predicting U.S. flu-related hospital admissions. The model was developed using Empirical Research Assistance (ERA), an AI tool that generates optimization algorithms, and ERA's underlying technology is now available to trusted testers.

  13. whXAI score67

    Gemini 4 Argon previewed with frontier coding and cyber defense claims

    AIThe post quotes Google's Sundar Pichai introducing Gemini 4 Argon as an early look at the next model. It claims frontier performance in complex workflows, cyber defense, and software engineering, and says Google teams are using it for tasks from coding to quantum computing. The author adds that on FrontierSWE the model is very self-critical and often says "Eureka!", a personality they describe as a large improvement over previous Gemini models.

    Image from @nrehiew_'s post
  14. Google FlowOfficialAI score38

    Google's Gemini Omni Flash guide offers prompting tips for Flow videos.

    AIGoogle Flow publishes a guide to creative prompting with Gemini Omni Flash, covering video generation for films, marketing, and visual assets. The guide recommends high-level constraints, first and last frame visual anchors, tagged image, video, and storyboard ingredients, and granular mid-scene pacing edits. It also suggests transferring style and motion from reference images and videos.

  15. RunwayOfficialAI score22

    Runway AI Summit opens in San Francisco with world simulator keynote

    AIRunway's AI Summit is underway in San Francisco, with co-founder and co-CEO Anastasis Germanidis opening the day with a keynote. He argued that universal world simulators will be the most important technological development of our time. Further updates from the event are promised throughout the day.

    Video from @runwayml's post
  16. koray kavukcuogluXAI score62

    Google's Koray Kavukcuoglu Announces Gemini 4 Argon for Trusted Defenders First

    AIGoogle is sharing Gemini 4 Argon first with trusted defenders in its Fairwind Program, with frontier capabilities in coding, knowledge work, and cyber security defense. The model is also being rolled out to the US government, and the company plans broader availability as testing progresses and safeguards allow.

    Why it matters: The author is a Google DeepMind leader announcing the model directly, so the rollout limits to trusted defenders and government are the key detail to note.

  17. Guillermo RauchXAI score30

    Vercel Connect invites services to reach developers and AI agents

    AIGuillermo Rauch invites service providers to add themselves to Vercel Connect to reach over 20 million developers and the agents they build. He argues that connecting services is now the main challenge in building, and that Connect makes it easier and more secure for both agents and apps. Services submit by describing themselves, adding OAuth or API key auth, verifying with a real token, and sending it for review.

  18. Varun MohanXAI score40

    Google announces Gemini 4 Argon, a new frontier model for software tasks

    AIGoogle announced Gemini 4 Argon, a new frontier model that delivers frontier performance across complex software tasks, according to Varun Mohan. Thousands of Googlers have been using it internally in Antigravity, and it is rolling out first to trusted cyber defenders in the Fairwind Program, with broader availability to follow as soon as possible.

  19. Logan KilpatrickXAI score45

    Google's Logan Kilpatrick celebrates Gemini 4 Argon model launch

    AILogan Kilpatrick, Google's Gemini lead, praised the teams behind Gemini 4 Argon and said he looks forward to wider adoption. The post links to Google's blog announcement for the model, but it gives no benchmark scores, pricing, or capability details.

  20. Google AIOfficialAI score72

    Google announces Gemini 4 Argon, a frontier model with 1M output tokens

    AIGoogle AI announced Gemini 4 Argon, a new frontier model built for deep reasoning across long, complex workflows in software engineering, legal and finance knowledge work, and cybersecurity defense. Google says it is expanding the model's output token limit to 1M tokens. Argon is rolling out first to trusted cyber defenders in the Fairwind Program, with broader availability to follow as soon as possible.

    Why it matters: The benchmark table compares Gemini 4 Argon against GPT-6 Astra and Claude models across knowledge work, coding, and multimodal tasks, showing where it leads and trails.

    Image from @GoogleAI's post
  21. Google DeepMindOfficialAI score88

    Google DeepMind releases Gemini 4 Argon to trusted cyber defenders first

    AIGoogle DeepMind announced Gemini 4 Argon, rolling out first to trusted cyber defenders through its Fairwind Program. Argon will launch at an introductory price of $2 per million input tokens and $10 per million output tokens, with output limits raised to 1M tokens. The post cites a 77.9% score on DeepSWE v1.1 and 91.7% on LVBench, and says broad availability will follow safeguard testing.

    Why it matters: The post pairs Argon's benchmark claims with the phased release, pricing, and safeguard details, helping readers weigh its frontier-level capabilities against its access limits.

  22. Microsoft CopilotOfficialAI score23

    Microsoft's new Copilot combines Home, Code, and Autopilot in one app

    AIMicrosoft's new Copilot brings Home, Code, and Autopilot together in one place for creating custom apps, building decks, automating workflows, and resuming work. Users can start using the Copilot app now and try new features as they become available in Frontier.

    Image from @MSFTCopilot's post
  23. Google · Gemini appOfficialAI score91

    Google announces Gemini 4 Argon, rolling out first to trusted cyber defenders

    AIGoogle announced Gemini 4 Argon, a new frontier model rolling out first to trusted cyber defenders through its Fairwind Program. The model's output limit rises to 1M tokens from 64K, and its introductory API price is $2 per million input tokens and $10 per million output tokens. Google says broader availability to developers, enterprises, and consumers will follow after more testing of guardrails.

    Why it matters: The post pairs benchmark claims with a phased access plan, pricing, and safety measures, which helps readers judge how quickly Argon may reach developers.

  24. GammaOfficialAI score22

    Gamma launches Salesforce connector for generating presentations from live data

    AIGamma has released a Salesforce connector that lets users link their Salesforce account and describe what they need, with Gamma building presentations from live data. The post cites use cases including weekly pipeline reviews, client-ready QBR decks, rep coaching one-pagers, admin onboarding docs, and marketing performance reports.

    Video from @GammaApp's post
  25. eric zakariassonXAI score22

    Cursor promotes engineering bots in Grok bot marketplace

    AICursor's Eric Zakariasson announced engineering bots available in the Grok bot marketplace on x.ai. The bots can hand off coding tasks to Cursor, manage pull requests through GitHub and Origin plugins, and share video demos of what they build.

    Image from @ericzakariasson's post
  26. PerplexityOfficialAI score20

    Perplexity's embedding preview tops ConTEB benchmark average nDCG@10

    AIPerplexity's embedding preview achieves the highest average nDCG@10 among tested models on ConTEB, though not on every task. It also outperforms voyage-context-4 on chunk retrieval while using 8x less storage per vector, at 1 KB (1024 dims, int8) versus 8 KB (2048 dims, float32).

    Image from @perplexity_ai's post
  27. PerplexityOfficialAI score20

    Perplexity distills query-aware compression scores to train retrieval embedders

    AIPerplexity overcomes gold-chunk supervision limits by distilling relevance from its query-aware context compression model. The model scores every document token against the query, and those scores, aggregated into chunk-level targets, train the embedder to retrieve answer and supporting chunks.

    Image from @perplexity_ai's post