Skip to contentSkip to stories

Updated

All AI news

Oct 6

Oct 6Tue
  1. Google ResearchAI score13

    Google sponsors COLM 2026 with 28 papers and 16 workshops

    AIGoogle, as a Diamond Sponsor of COLM 2026, will have researchers from Google Research and Google DeepMind present 28 papers and participate in 16 workshops at the Hilton Union Square in San Francisco from October 6–9. Attendees can visit Google booth #107 to explore work on agentic systems, reasoning, and multimodal AI.

  2. a16z NewsAI score30

    General Medicine Raises $120M Series B to Build Healthcare Shopping Platform

    AIGeneral Medicine has raised a $120M Series B to build a healthcare "general store" that gives patients clinical guidance, transparent pricing, and help accessing care in one place. Its offering combines its own medical group with a marketplace of external primary care physicians and specialists, plus a catalog of over 2,900 medications, labs, imaging, consultations, and select procedures. The company also plans to offer the same infrastructure to health systems, payors, and pharmacies.

  3. ARC PrizeAI score14

    ARC Prize hosts Frontier AI benchmarking dinner with Snorkel AI during SF Tech Week

    AIARC Prize is hosting its first Frontier AI Benchmarking Ecosystem Dinner with Snorkel AI during SF Tech Week. The event will bring together 40 experts from nonprofit benchmarking organizations, industry, academia, and government to discuss AI benchmark design and building a more scalable, transparent evaluation ecosystem. Attendees named in the post include representatives from METR, Epoch, Vals, Artificial Analysis, Harvard, and Stanford.

  4. Paige BaileyAI score20

    Google's Gemini apps rank high in new a16z consumer AI rankings

    AIGoogle has five products in a16z's top 50 consumer AI web apps, with the Gemini app at #2, NotebookLM at #9, Google AI Studio at #11, Google Labs at #18, and Antigravity at #44. Paige Bailey highlights Google AI Studio, a developer workbench, reaching the overall consumer top ten. Steven B. Johnson notes Google is the only company with two products in the top ten and four in the top twenty.

  5. ARC PrizeAI score25

    ARC Prize finds DeepSeek V4.1 Flash high reasoning gains no clear edge

    AIOn ARC-AGI-1, DeepSeek V4.1 Flash scored 88.5% at high reasoning versus 90.5% at low, with high using 35% more output tokens without consistently better answers. On ARC-AGI-2, the reported per-task cost of max reasoning ($0.129) appears slightly lower than high ($0.133), but after excluding incomplete tasks caused by API issues, max is about 4.5% more expensive per task.

  6. Google ResearchAI score51

    Google's PDFM location embeddings improve five global public health tasks

    AIGoogle Research reports that Population Dynamics Foundation Model (PDFM) embeddings, built from search trends, mobility, built environment, and weather signals, were tested by partners across five public health tasks. The embeddings improved results in cross-border MMR vaccination coverage, dengue forecasting, postpartum depression screening, and cholera outbreak prediction, and matched census inputs for cardiovascular mortality nowcasting.

  7. SemiAnalysisAI score18

    ClusterMAX rates FarmGPU underperform on Slurm and Kubernetes testing

    AISemiAnalysis rated FarmGPU as ClusterMAX Underperform after its Slurm layer failed to advertise GPU resources and Kubernetes exposed no RDMA devices for scale-out networking. The post credits FarmGPU's Grafana monitoring, provisioning notes, and trustworthy technical team, while noting the team may be stretched thin across small clusters.

  8. MIT News · AIAI score22

    MIT launches MIT for America initiative to strengthen STEM education nationwide

    AIMIT has launched MIT for America, a strategic initiative to strengthen STEM education across the United States, from kindergarten through community college. Its programs target mathematical thinking and problem solving, living and working with AI, and hands-on design and fabrication. The initiative builds on the MIT for America Calculus Project, which pairs MIT students and alumni with school districts for calculus tutoring.

  9. The Next PlatformAI score20

    When One Datacenter Is No Longer Enough: Cisco on Scale-Across AI Networking

    AICisco SVP Rakesh Chopra discusses the "Scale-Across" approach to networking AI training workloads spread across multiple data centers. He describes how Silicon One architecture and Intelligent Collective Networking aim to manage synchronized GPU traffic over long-distance fiber links. The interview covers power efficiency and hardware-accelerated MACsec and IPsec security.

  10. Merve NoyanAI score72

    Mistral Large 4 will open its weights at the end of October

    AIMistral announced Mistral Large 4, which it describes as a natively multimodal model with 1T parameters and 49B active. Mistral says it is available via API now, with open weights to follow at the end of October, and a Hugging Face page is listed for the release.

    Why it matters: The quoted Mistral announcement gives specific size, activation, and API details, and the open-weights timing matters for teams weighing open model options.

  11. Simon WillisonAI score36

    Mistral's Pelican SVG Test Passes, Tied to Mistral Large 4 Context

    AISimon Willison reports that Mistral can now generate his pelican SVG test, shared via a Markdown SVG renderer. The post links to a rendered result but gives no benchmark or scoring details. Background from Mistral's own announcement describes Mistral Large 4 as a 1T-parameter, natively multimodal model with 49B active parameters, available via API today and with open weights planned for end of October.

  12. Aravind SrinivasAI score42

    Perplexity Computer plays real-time StarCraft against itself, Blue wins 2-5

    AIPerplexity's Computer ran two agents playing StarCraft against each other in real time, with the game never paused while each agent thought. Blue, playing with 41 Dragoons, lost the final match 2-5 to Red, which used High Templar and Psionic Storm after Blue failed to scout Red's build. Each agent received only its own fog-of-war-limited game state, and video input was not provided.

  13. Interconnects (Nathan Lambert)AI score52

    Nathan Lambert argues the open-weight cyber risk debate is missing trade-offs

    AINathan Lambert argues that policy debates on open-weight model cyber risks lack nuance, because banning open models may not reduce risk and could weaken American competitiveness. He says closed frontier APIs have been tied to most documented cyber attacks, and that restricting open models while closed models keep advancing could widen the offense-defense gap. He also argues that Chinese labs' safety practices are shaped by their own government and society, and that the claimed risk of models like Claude Mythos has been overstated.