Skip to contentSkip to stories

Updated

All AI news

Oct 2

Oct 2Fri
  1. TransformerAI score55

    Human oversight may not prevent AI-driven military errors, analysis argues

    AIJoshua Keating argues that keeping a human in the loop on lethal AI decisions is not enough if the humans rely too heavily on AI outputs. He cites a CNN-reported case in which an analyst's AI-assisted report falsely identified a Chinese ship's cargo as nuclear components, nearly prompting a boarding during the Iran war. The piece links this to automation bias and to military AI cases in Gaza and Minab, and warns that AI integration early in a nuclear decision chain is harder to regulate than autonomous launch.

  2. GitHub Blog · AI & MLAI score23

    Three Skills Developers Need as AI Changes Their Work

    AIAI is changing developer work, and the article recommends three skills: directing AI agents, reviewing AI output instead of trusting the first answer, and using saved time for judgment-heavy problems such as customer needs and tradeoffs. It cites GitHub Copilot's built-in Rubber Duck agent, which uses a second model to critique plans, code, and tests. The author argues that developers remain responsible for outcomes while AI handles more implementation.

  3. a16z NewsAI score32

    The Case for Scaling America's Defense Manufacturing Base Beyond Prototypes

    AIVenture investors have funded defense-tech companies such as SpaceX, Anduril, and Castelion, but the article argues that production capacity in the supplier base is now the bottleneck. Most of America's machine shops and manufacturers are small, with 83% of machine shops employing fewer than 20 people, and 61% of tier-two-and-below defense manufacturers cite tooling, automation, or production-line limits as top expansion barriers.

  4. MIT Technology Review · AIAI score62

    AlphaGo's move 37 shows why LLMs do not truly reason, an AlphaGo team member argues

    AIThore Graepel, a core member of the AlphaGo team, argues that current large language models do not truly reason, despite chain-of-thought gains in math and coding. He says they lack an explicit, inspectable epistemic state, keep knowledge and reasoning intertwined in their weights, and often produce post-hoc explanations. He proposes systems that maintain an auditable epistemic state and evaluate each step by how much it resolves uncertainty.

  5. AI Futures ProjectAI score62

    Former OpenAI forecaster urges Senate to curb AI research automation race

    AIDaniel Kokotajlo, who leads the AI Futures Project, testified before a Senate subcommittee on September 30, 2026. He argued that Anthropic and OpenAI are racing toward superintelligence by automating AI research and development, and that his team thinks this could happen as early as 2028. He warned that declining monitorability and models that appear aligned during evaluations make misalignment harder to detect, and he recommended greater industry transparency and redirecting compute away from AI R&D.

  6. Lucas BeyerAI score45

    Lucas Beyer praises new coding benchmark for finding bugs in repos

    AILucas Beyer calls SWE-sweep a useful new benchmark, where agents must find and fix bugs in a repo checked out at an earlier commit, scored against unit tests from real later bugfixes. He notes two limitations: a model may find valid bugs that don't match the tested ones, and the construction makes training on the test set easy. He advises not overemphasizing small ranking differences once models score highly.

  7. Dongxi NLPAI score27

    LLMs replace condescending engineers by explaining code patiently in many formats

    AIThe author recalls a senior engineer who dismissed a newcomer's question with "oops, forgot," and says LLMs now answer patiently through text, diagrams, videos, and more. The post frames this shift as making dismissive gatekeeping obsolete, building on Andrej Karpathy's tips for turning LLM outputs into easier-to-read formats such as ASD-STE100 writing, diagrams, HTML pages, and generated explainer videos.

Oct 1

Oct 1Thu
  1. Latent.SpaceAI score60

    Recursive Language Models explained by MIT's Alex Zhang on coding agents

    AIA Latent.Space podcast episode features MIT researcher Alex Zhang explaining recursive language models (RLMs). He discusses why Claude Code, Codex, and Pi are basically the same, and how RLMs use code, context offloading, and recursive subagents to generalize across tasks. The episode also covers OpenAI's 10,000-agent, 130B-output-token experiment and academia's freedom to pursue ambitious research bets.

  2. Ben TossellAI score10

    Ben Tossell says AI tools fail non-technical users and must improve

    AIBen Tossell argues that AI products are not good enough for semi-technical users like himself, who are neither developers nor casual consumers. He says companies should stop blaming users for poor experiences and make their products better, since current results hurt the broader AI narrative. His reply to his own earlier post calls his experience with dot "💩".

  3. Kylie RobisonAI score6

    Kylie Robison Says Journalists Should Make News About Themselves

    AIJournalist Kylie Robison jokingly argues that it is her duty as a journalist to make this about her, linking to a post on X. The post is brief and offers no substantive news, and the quoted reply from @dwr only comments that "dot" is a useful name for a personal agent if it launches a dot-shaped hardware device.

  4. Harrison ChaseAI score33

    Harrison Chase Argues Every Agent Harness Needs a Durable Runtime

    AIHarrison Chase argues that every agent harness requires a durable runtime, citing pi-durable as an example alongside deepagents built on LangGraph. The post frames durable execution as a basic requirement for agent systems rather than an optional feature. Pi 1.0 shipped with Pi Durable, which the referenced @pidotdev post invites users to customize.

  5. Dongxi NLPAI score46

    Dongxi jokes about replacing remote consultants with Griffin AI agents

    AIThe author jokes about founding a consulting firm that would use agents for work, Griffin for meetings, and Griffin for interviews to fill remote roles. They then question whether remote engineers and consultancies would still be needed if that became reality. The quoted Tavus post says Griffin passed a video Turing test with 48% of live interlocutors believing it was human.