Skip to contentSkip to stories

Updated

#Expert opinion

Showing low-relevance items too. Hide low-relevance items

Apr 24

Apr 24Fri
  1. Ahmad Al-DahleAI score82

    Ahmad Al-Dahle says DeepSeek-V4's efficient 1M context is its key bet

    AIAhmad Al-Dahle argues that the most interesting part of DeepSeek-V4 is its bet on efficient ultra-long context rather than its benchmarks. He says this is the precondition for test-time scaling and long-horizon agents, and cites 27% of V3's FLOPs at 1M tokens. The quoted DeepSeek post announces DeepSeek-V4-Pro (1.6T total, 49B active) and DeepSeek-V4-Flash (284B total, 13B active), both open-sourced with 1M context and API access.

    Why it matters: The post argues that efficient 1M-token context, not benchmark scores, is the key bet behind DeepSeek-V4's design for test-time scaling and long-horizon agents.

Apr 22

Apr 22Wed
  1. Cognition Blog (Devin, Windsurf)AI score54

    Cognition says building cloud agents requires VM isolation, state snapshots, and org change

    AICognition argues that enterprises building cloud agents face three problems: shared container kernels, the inability to persist agent state across async gaps, and the scale of orchestration, governance, and integrations. The post says VM-level isolation with hypervisor-level snapshots was needed for Devin, and that organizations must also rebuild engineering processes around agent execution.

Apr 21

Apr 21Tue
  1. Cognition Blog (Devin, Windsurf)AI score72

    Cognition says multi-agent systems work when only one agent writes

    AICognition reports that multi-agent setups work best when writes stay single-threaded and extra agents contribute intelligence instead of actions. It describes a code-review loop where a clean-context review agent catches bugs in Devin-written PRs, averaging 2 bugs per PR with roughly 58% severe. The post also says the smart-friend pattern, pairing a smaller primary model with a stronger one, has not yet worked well with asymmetrically weaker primaries and is an open training problem.

    Why it matters: The post gives concrete findings on which multi-agent setups work, including clean-context code review and smart-friend escalation, and where they still fail.

  2. Awni HannunAI score14

    Awni Hannun argues top tech firms must be extreme co-design companies

    AIAwni Hannun argues the biggest companies cannot be just hardware, software, or AI, and must instead practice extreme co-design, which he says is where the global minima lie. He cites Nvidia as an example and says Apple should be, and hopefully will be, an extreme co-design company despite many calling it a hardware company after its recent transition.

Apr 20

Apr 20Mon
  1. Soumith ChintalaAI score36

    Soumith Chintala Critiques Dwarkesh's AGI Framing After Jensen Huang Podcast

    AISoumith Chintala says Jensen Huang understood AI ecosystems, trade, and policy far better than host Dwarkesh Patel in their podcast. He argues that no single model such as Mythos marks a critical phase change, since a state-of-the-art Chinese open-source model with three orders of magnitude more test-time compute and unpublished post-training advances would be a more realistic baseline. He also says American policy should use measured, continuous levers across a Western-controlled ecosystem rather than abrupt interventions.

Apr 18

Apr 18Sat

Apr 16

Apr 16Thu

Apr 14

Apr 14Tue
  1. Jan LeikeAI score14

    Jan Leike outlines top-down approach to automating alignment research

    AIJan Leike distinguishes two ways to automate alignment research: bottom-up, where researchers automate more of their existing work, and top-down, where specific subproblems are carved out for AI to solve. He says Anthropic's work mostly follows the bottom-up path, such as using Claude for coding, while this post focuses on the top-down approach.

Apr 9

Apr 9Thu
  1. Andrej KarpathyAI score45

    Karpathy says AI capability gap stems from uneven use and training

    AIAndrej Karpathy argues that people judging AI from free-tier ChatGPT or Advanced Voice Mode miss the strong capabilities of current agentic models like OpenAI Codex and Claude Code. He says gains are "peaky," concentrated in verifiable technical domains like programming and math that suit reinforcement learning and attract B2B investment, while writing and everyday advice improve less. Those who use frontier agentic tools professionally in these fields see far greater capability, which is why the two groups talk past each other.

Apr 7

Apr 7Tue

Apr 6

Apr 6Mon
  1. OpenAI Alignment Research BlogAI score31

    OpenAI opens applications for Safety Fellowship on AI safety and alignment research

    AIOpenAI announced applications for its Safety Fellowship, a pilot program supporting external researchers, engineers, and practitioners in safety and alignment research on advanced AI systems. The program runs from September 14, 2026 through February 5, 2027, with a monthly stipend, compute support, API credits, and mentorship, and fellows are expected to produce a substantial output such as a paper, benchmark, or dataset. Applications close May 3, and successful applicants will be notified by July 25.

Apr 5

Apr 5Sun

Apr 4

Apr 4Sat
  1. Andrej KarpathyAI score62

    Andrej Karpathy outlines an LLM-maintained markdown wiki workflow for personal research

    AIKarpathy describes using LLMs to compile raw source documents into a markdown wiki that he views in Obsidian, with the LLM writing and maintaining most of the wiki. He reports that at about 100 articles and 400K words, the LLM agent can answer complex questions directly from the wiki, and he also runs LLM health checks to find inconsistencies and gaps. He shares the underlying idea as an "idea file" that users can give to their own agents to build a customized version.

Apr 2

Apr 2Thu
  1. AI Futures ProjectAI score62

    AI Futures Project shortens Automated Coder timelines to mid 2028

    AIAI Futures Project moved Daniel Kokotajlo's Automated Coder median from late 2029 to mid 2028 and Eli's from early 2032 to mid 2030. The main reasons cited are a faster METR time horizon doubling time and the impressive results of Claude Opus 4.6. The authors also say progress in agentic coding has been faster than expected over the past 3 to 5 months.

Apr 1

Apr 1Wed
  1. Ahmad Al-DahleAI score12

    Ahmad Al-Dahle says incidents should drive systems design, not blame

    AIAhmad Al-Dahle argues that the best teams build systems that make right actions easy and wrong ones hard. He says strong cultures treat every incident as a systems design question rather than a matter of assigning blame. The quoted post by @bcherny attributes a recent mistake to a manual deploy step that should have been automated, and the team has since improved that automation.

Mar 30

Mar 30Mon
  1. Mckay WrigleyAI score22

    AI tools may soon use, clone, and extend any software autonomously

    AIMckay Wrigley predicts AI tools will within 6-12 months autonomously use any software, clone it in a weekend, monitor it for updates, and add custom features. He frames this as a future where users never need to operate their computer themselves. The prediction follows a referenced Claude Code update adding computer use in research preview for Pro and Max plans.

Mar 28

Mar 28Sat
  1. Andrej KarpathyAI score12

    Karpathy: LLMs can argue both sides, so beware sycophancy

    AIAndrej Karpathy reports that an LLM spent four hours strengthening his blog post's argument, then convinced him of the opposite when asked to argue the reverse. He concludes that LLMs are highly capable of arguing almost any direction, which makes them useful for forming opinions if users ask from multiple angles and watch for sycophancy.

Mar 26

Mar 26Thu
  1. Mckay WrigleyAI score22

    Mckay Wrigley urges developers to build MCP apps after Anthropic's rise

    AIMckay Wrigley argues that Anthropic has strong product taste, citing how it turns overlooked ideas into popular products once it commits to them. He says people were wrong to dismiss MCP and encourages developers to start building MCP apps. He also highlights bidirectional communication between users and models through MCP apps as a feature the masses have yet to discover.

    Image from @mckaywrigley's post
  2. Andrej KarpathyAI score47

    Karpathy wants agents to handle full app DevOps from one command

    AIAndrej Karpathy argues that the hardest part of building a deployed app is not the code but the DevOps work of assembling services, API keys, payments, auth, and deployment. He says the goal is for agents to handle this entire lifecycle as code, with agent-native CLI and API access instead of manual web clicking. He calls it a from-scratch redesign that is only now barely technically possible.

  3. Hamel HusainAI score38

    Data Scientists Face New Pressures as LLM APIs Let Teams Ship AI Without Them

    AIHamel Husain argues data scientists remain essential as foundation-model APIs let teams ship AI without them, because much of the work lies in evaluation, debugging, and metric design. He says teams often rely on generic off-the-shelf metrics and unverified LLM judges instead of examining their own data. He lists five eval pitfalls, starting with generic metrics, and recommends looking at traces and doing error analysis.

Mar 25

Mar 25Wed

Mar 24

Mar 24Tue
  1. Anthropic EngineeringAI score78

    How Anthropic built Claude Code auto mode to replace skipped permissions

    AIAnthropic describes Claude Code auto mode, which delegates approval of agent actions to model-based classifiers instead of manual prompts or skipped permissions. The classifier reviews tool calls before execution and a separate probe screens tool outputs for prompt injection. Anthropic reports a 0.4% false positive rate on real internal traffic and a 17% false negative rate on real overeager actions.

    Why it matters: The post explains the layered classifier design and its measured tradeoffs, showing how autonomous coding agents can cut approval fatigue without fully removing risk.