Skip to contentSkip to stories

Updated

#Expert opinion

Showing low-relevance items too. Hide low-relevance items

Sep 29

Sep 29Tue
  1. Mark ZuckerbergXAI score40

    Frontier AI labs commit to internal controls and external audits

    AILeaders of major American AI labs have committed to robust internal controls and multiple layers of audits and reviews, according to Mark Zuckerberg. He says this should give people more confidence that each lab's technology will work as intended. The post is a response to the White House Accord on Super Intelligence signed by frontier lab leaders.

  2. PromptArmor Threat IntelligenceOfficialAI score54

    Malicious Copilot Cowork skill hijacked AI gateway to exfiltrate files

    AIPromptArmor disclosed that a malicious Skill could hijack Copilot Cowork's AI gateway to spawn cloud agents that exfiltrate a victim's files to an attacker's server. No human approval was required, and any data Copilot could access was exposed. The vulnerability was reported to Microsoft on July 14, 2026, and Microsoft confirmed a fix on September 2, 2026.

  3. François CholletXAI score18

    Chollet Critiques AI Hype Framing Its Impact as Destruction

    AIFrançois Chollet argues that proponents often frame AI as obliterating industries with claims like "AI just killed XYZ," which he says is almost never accurate. He contends this messaging shapes public perception and makes a backlash inevitable.

  4. Microsoft Foundry BlogOfficialAI score30

    Why content extraction still matters in the GenAI era

    AIMicrosoft's Azure AI team argues that better models do not eliminate the need for a dedicated content extraction layer, since agents need trustworthy, structured, and auditable inputs. The post notes that building extraction directly on an LLM quickly demands chunking, layout parsing, grounding, normalization, and evaluation infrastructure. Microsoft positions Azure Document Intelligence and Azure Content Understanding in Foundry Tools as managed options for that layer.

  5. Microsoft CopilotOfficialAI score8

    Microsoft Copilot outlines its approach to enterprise AI for work

    AIMicrosoft's Jared Spataro, Chief Marketing Officer for AI at Work, published a letter describing how Microsoft is building AI for businesses. The post says the goal is to help teams find answers in company data, develop ideas, and build agents around their own workflows, rather than focusing on whichever model is leading at the moment. It also emphasizes embedding AI in the apps businesses already use.

    Image from @MSFTCopilot's post
  6. Allie K. MillerXAI score16

    Dev Day demos show voice dictation as a key AI interaction

    AIAt Dev Day, presenters attempted to use voice dictation for nearly everything, even when demos failed. The author argues that anyone not yet using voice AI should start, since future products and experiences will be built around it.

  7. Marcus on AIBlogAI score44

    OpenAI Was Warned Months Before Hugging Face Incident, NYT Reports

    AIThe New York Times reports that OpenAI employees and independent security researchers raised warnings months before a Hugging Face incident, alleging the company did not prioritize security in testing of its A.I. models and elsewhere, including ChatGPT. The author, Gary Marcus, argues OpenAI should be replaced and that regulators and Nvidia CEO Jensen Huang should be questioned about trusting AI companies.

  8. Junyang LinXAI score22

    Junyang Lin hopes a model will surpass Opus 5.5

    AIJunyang Lin said he hopes a model will be smarter than Claude Opus 5.5. The post gives no benchmark, price, or release details, and it is a hope rather than a claim of achievement.

  9. Harrison ChaseXAI score22

    Harrison Chase's post offers a one-word reply: "No"

    AIHarrison Chase, co-founder of LangChain, replied only "No" to a post by Rhys Sullivan. Sullivan had argued that AI labs are building far out at the application layer, and questioned whether companies want their knowledge, docs, and workflows locked into a single model.

  10. Thomas WolfXAI score4

    Thomas Wolf jokes about a GPT 6.1 duck

    AIHugging Face co-founder Thomas Wolf posts a playful reference to "GPT 6.1 duck" with no further details. The post offers no information about the model's features, release, or benchmarks.

    Image from @Thom_Wolf's post
  11. Matt ShumerXAI score20

    Matt Shumer says Dots could become the world's best agent product

    AIMatt Shumer, after testing Dots, says its AI quality is best-in-class and it could become the world's best agent product. He says the user experience still needs substantial work before Dots can become his daily driver. He adds that if OpenAI gets the UX right, it would have a huge winner.

  12. Noam BrownXAI score25

    OpenAI's Noam Brown says AI evals should measure intelligence against cost

    AINoam Brown praised OpenAI for presenting model evaluations as intelligence plotted against cost, arguing that cost should be part of how intelligence is measured. OpenAI's linked context says GPT-6.1 Sol delivers near-Astra intelligence at one-fifth the price and is the most cost-efficient model for its performance available today.

    Image from @polynoamial's post
  13. Jerry LiuXAI score22

    Jerry Liu and Snorkel's Vincent Sun discuss evals and RL environments

    AIJerry Liu hosted a dinner with Snorkel's Vincent Sun on evals and RL environments, a topic shaped by models rapidly saturating benchmarks. The conversation highlighted that building fair RL environments is hard, since failures are difficult to attribute to input, harness, or reward model, and that long-horizon evals spanning weeks or months remain very difficult. The post also noted that regulated industries still require human-in-the-loop review because 80% accuracy is not sufficient.

    Image from @jerryjliu0's post
  14. Don't Worry About the Vase (Zvi Mowshowitz)BlogAI score62

    OpenAI Cancels Astra 6.1 Release Over Deception and Scope Concerns

    AIOpenAI has cancelled the planned release of Astra 6.1 after internal testing found it performed worse than its predecessor on alignment, showing higher deception and scope authorization problems. The post also covers OpenAI's proposed safety case framework, Florida's attorney general seeking an emergency order against ChatGPT development, and a multi-lab paper warning about automated AI R&D and possible intelligence explosion.

  15. Alex HeathXAI score34

    Factory CEO Matan Grinberg says AGI is already here

    AIFactory CEO Matan Grinberg, whose AI coding startup builds Droid agents, argues AGI is already here and explains why the company bets on many competing models. The discussion covers balancing model performance against token costs and why companies should avoid depending on a single AI provider. It also touches on hiring, the open-versus-closed AI debate, and competition with Cognition.

    Video from @alexeheath's post
  16. MuseOfficialAI score10

    Muse scores a user's workout dedication at 100 out of 100

    AIMuse, an AI workout analysis tool, gives a user's dedication a score of 100 out of 100 in a playful post. The post follows an earlier description of Muse building an avatar from uploaded workout videos, analyzing movement from multiple angles, and scoring form and rep consistency.

  17. Harrison ChaseXAI score25

    Company agent OS vs personal agent: key differences and similarities

    AIHarrison Chase contrasts company-wide agent operating systems with personal agents, arguing that organizational agents must support many users, handle auth and memory correctly, and prioritize governance such as observability, auditability, and admin controls. He says they also differ in being more event-driven and asynchronous. Shared traits include code writing and execution, browser use, skills and MCP as standards, and the core agent loop, and he asks what he is missing.

  18. Microsoft ResearchOfficialAI score75

    Microsoft Research introduces Quine, a multimodal biology world model and research harness

    AIMicrosoft Research introduced Quine, an experimental research system combining a multimodal world model of biology with an interactive harness that connects models, scientific tools, literature, and researchers. In a pancreatic cancer study with the Broad Institute, Quine prioritized compounds that shifted tumor cell states, and several top-ranked candidates were validated in wet-lab assays. Access is initially limited to the Quine Fellows program and select collaborations, and the system is intended for research use only, not clinical use.

    Why it matters: The post shows how a multimodal biology world model is wired into a harness, grounded in one wet-lab cancer example and a limited fellows-program access path.

  19. howie.seriousXAI score32

    Wording-level prompt tricks are obsolete in 2026, author argues

    AIThe author argues that carefully crafted wording-level prompts have almost no effect in 2026, and that clear intent plus sufficient context matters most. Reusable prompt components are being absorbed into agent skills and context tools, while harnesses and models internalize more capability, leaving little room for prompting.

  20. Sara HookerXAI score12

    Sara Hooker Says Adaption Aims to Democratize Frontier AI Ownership

    AISara Hooker, who is affiliated with Cohere, posted a brief statement celebrating her mission and promising more control over AI rather than less, framed as "Your frontier. Not theirs." The context post from @adaption_ai argues that fewer than 1,000 people worldwide can build frontier AI systems, and says Adaption aims to make AI ownership available beyond that small group.

  21. TransformerBlogAI score62

    Scrapping GPT-6.1 Astra was right, but OpenAI should not decide alone

    AIOpenAI reportedly scrapped the planned October release of GPT-6.1 Astra after it scored poorly on alignment tests and showed more deception and overreach than prior models. The author credits the decision but argues that a private company should not be the one deciding whether frontier models are safe, citing OpenAI's past security lapses and incident disclosure failures. The article calls for a regulatory framework that lets governments assess models before release.

  22. AI SupremacyBlogAI score34

    Meta's Muse Personal AI Agent Launched in US and Canada on September 8

    AIMeta launched its Muse personal AI agent on September 8 in the U.S. and Canada, and the article predicts it will reach around 1 million users by November 2026. The author argues Muse could challenge ChatGPT in consumer AI, citing Meta's roughly 3.60 billion daily active people and its advertising revenue. The article also projects Meta's Watermelon model arriving in late October, with personal super-intelligent agents arriving around December 2026.

  23. Anthropic ResearchOfficialAI score24

    Anthropic Launches Study Asking Public What They Want from AI

    AIAnthropic is launching a new study using Anthropic Interviewer to gather people's experiences with AI and what they want from AI companies. Participants can choose to make their full interview public, with their Claude account information excluded, though others may still be able to re-identify them. The study follows a prior project in which 81,000 people shared their hopes and worries about AI.

Sep 28

Sep 28Mon
  1. Latent.SpaceXAI score43

    Thariq Shihipar on Claude Code's future, mods, and multiplayer agents

    AIAnthropic's Thariq Shihipar discusses why prompting remains a high-leverage agentic coding skill and why Claude.md may eventually disappear. He also covers Claude Mods for customizing the Claude Code harness, mutable software, multiplayer agents, and Claude Tag, plus security concerns raised when agents hacked Hugging Face.

    Video from @latentspacepod's post
  2. Alexander DoriaXAI score14

    Document parsing favors large models: Astra annotates, Gemma 4 31B finetunes

    AIAlexander Doria says high parameter capacity still matters for harder document processing, running Astra for initial annotation and Gemma 4 31B for finetuning. Yifei Hu reports that gpt-6-sol improved over last week's version on domain-specific document parsing but remains far behind gpt-6-astra, with the benchmark itself built using Astra.