Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 8

Oct 8Thu
  1. Latent SpaceBlogAI score59

    Periodic Labs argues AI scientists need physical experiments, not just more data

    AIPeriodic Labs' Liam Fedus and Ekin Dogus Cubuk explain why scientific discovery differs from math and coding, and why experiments remain the ground truth. They describe reinforcement learning grounded in physical experiments, AI-driven materials characterization, and the view that failed experiments can be valuable training data. The transcript was truncated before the discussion of giving lab instruments "140 IQ" was completed.

  2. SunoOfficialAI score22

    Suno launches Albums for bundling songs into full releases

    AISuno announced that Albums are now live, letting users combine songs into a full release, set artwork, arrange the tracklist, and publish when ready. Existing playlists can be converted into Albums without rebuilding them from scratch.

    Video from @suno's post
  3. Google GemmaOfficialAI score27

    EmbeddingGemma 2 developer guide released by Google

    AIGoogle Gemma has published a developer guide for EmbeddingGemma 2, with code snippets to help developers start searching beyond text. The post directs readers to the full guide on the Google Developers Blog.

  4. Google GemmaOfficialAI score44

    Google publishes a developer guide for EmbeddingGemma 2 multimodal embeddings

    AIGoogle Gemma announces a developer guide showing how to embed text, code, images, video, audio, and interleaved inputs with EmbeddingGemma 2 using the sentence-transformers library. The guide outlines a four-step workflow: loading the model, embedding text and code with task prompts, embedding multimodal inputs, and optionally truncating dimensions with Matryoshka.

    Image from @googlegemma's post
  5. AWS Machine Learning BlogOfficialAI score27

    Share SageMaker HyperPod GPU clusters across teams with isolation and fair scheduling

    AIAWS published a reference architecture for running multiple teams on one Amazon SageMaker HyperPod EKS cluster, with each team isolated in its own Kubernetes namespace. The design combines AWS IAM Identity Center for authentication, per-team SageMaker AI domains, HyperPod Task Governance for fair resource allocation, and namespace-level cost allocation for per-team spend visibility.

  6. MarkTechPostNewsAI score58

    JetBrains releases Mellum2.1, a 12B MoE open model for coding agents

    AIJetBrains has released Mellum2.1, a 12B mixture-of-experts thinking model with 2.5B active parameters, under Apache 2.0 on Hugging Face. Post-training reinforcement learning in real software repositories raised SWE-bench Verified from 2.0 to 47.0, according to JetBrains' self-reported results. Qwen3.5-9B still leads on SWE-bench Pro, GPQA Diamond and AIME, and GGUF builds start at 7.0 GB for local use.

  7. Goodfire ResearchOfficialAI score57

    Goodfire deploys probe-based cyber monitors on Kimi K3 with a judge cascade

    AIGoodfire Research describes probe-based cyber monitors for Kimi K3 and GLM 5.3 deployed on a production inference stack. The probe filters suspicious exchanges before an LLM judge reviews them, reaching about 93% recall at a 5.5% benign-session interruption rate at roughly 50x lower judge cost. In FAR.AI's red-teaming, the monitor reduced universal jailbreaks to zero across 140 tested strategies.

  8. Vercel DevelopersOfficialAI score36

    StepFun's Step 5 Preview model now available on Vercel AI Gateway

    AIVercel says StepFun's flagship Step 5 Preview, built for agentic coding, research, and finance, is now live on AI Gateway. The model offers a 1M-token context window, accepts text and image input, and uses a 600B-parameter mixture-of-experts design with 27B parameters active.

  9. Grok BotOfficialAI score38

    Grok Bot can now help run Shopify stores

    AIGrok Bot can now connect to Shopify to check orders, track inventory, and keep product listings up to date. The feature lets merchants ask it questions about their business after linking their store. Shopify's related announcement says the new connectors also allow merchants to build teams of agents to help run their business.

    Image from @bot's post
  10. ClaudeDevsOfficialAI score33

    Anthropic credits work with Messages API, Managed Agents, and Agent SDK

    AIAnthropic's credits can be used with the Messages API, Claude Managed Agents, and the Agent SDK. They cannot be used for interactive Claude Code sessions, but they also apply in third-party harnesses that accept a Claude API key.

  11. ClaudeDevsOfficialAI score46

    Anthropic adds monthly API credits for Max and Team plans

    AIAnthropic now provides monthly Claude Platform API credits to Max and Team subscribers: $100 for Max 5x, $200 for Max 20x, and up to $500 pooled for Team. The credits can be used on any Claude model, including in code or third-party harnesses.

  12. Arena.aiOfficialAI score30

    Arena raises Series B led by Lightspeed and Khosla Ventures

    AIArena has announced a Series B round co-led by Lightspeed and Khosla Ventures, with participation from Salesforce Ventures, 01 Advisors, Dell Technologies Capital, and Endeavor Catalyst. Existing investors Andreessen Horowitz, Felicis, a16z's AMP, QuantumLight, and The House Fund also supported the round.

  13. Arena.aiOfficialAI score55

    Arena raises $200M Series B and launches Alignment Index for AI agents

    AIArena announced a $200M Series B at a $3.1B valuation and released its Alignment Index, a benchmark measuring agent safety and alignment. The index is built from 90K+ real-world agent sessions across 27 models and tracks Unauthorized Action, False Attribution, and Deceptive Completion. OpenAI's GPT-6.1-Sol leads with a score of 87.9, ahead of Claude-Opus-5.5 at 83.2 and Grok-4.7 at 82.7.

    Video from @arena's post
  14. Arena.aiOfficialAI score22

    Arena reports GPT-6 model variants' false attribution rates

    AIArena found that some models misquote users while others credit users with others' work in false attribution cases. GPT-6 Luna and Astra rarely misquoted users, at 15.6% and 28.6%, but often misattributed statements, at 53.1% and 48.2%. Sibling model GPT-6 Sol had the highest rate of misstating the user's history, at 23.5%.

    Image from @arena's post
  15. OpenAI NewsOfficialAI score26

    How Oracle turns days of work into minutes with ChatGPT and Codex

    AIOracle is using ChatGPT Work and Codex to turn specialist knowledge into fast, repeatable workflows across recruiting, engineering, and operations. The source does not provide figures, timelines, or specific results beyond the headline's claim that days of work can take minutes.

  16. GoogleOfficialAI score40

    Google AI estimates gestational age within four days in clinical study

    AIIn a prospective clinical study, Google's models pinpointed gestational age within 4 days of accuracy. The company says that precision could meaningfully affect clinical care, and that extending such tools to low-resource settings could help reduce maternal deaths and close care gaps worldwide.

    Image from @Google's post
  17. Boris PowerXAI score34

    Boris Power says progress on a result has been remarkable

    AIBoris Power, who owns OpenAI's account, praised the pace of progress on a result he called remarkable. The context from @0xdoug reports a PR merged and a tightened bound from κ = 2⁻¹⁸² to κ = 2⁻¹⁵, a community effort across contributors.

  18. Thomas WolfXAI score67

    Carbon-A open model and database find 566 million candidate genes across 22,617 species

    AIThomas Wolf says Carbon-A, an open model that finds genes directly in DNA, has been released with a database of 566.34 million candidate genes across 22,617 species. The team reports wet-lab validation of several new genes in cats, chickens and arabidopsis, and RNA evidence for 239 genes missing from reference annotations of common species.

    This story has a top pick“Carbon-A open model and database predict 566 million gene candidates across 22,617 species”

  19. ZyphraOfficialAI score34

    Zyphra's lossless method cuts communication for MoE expert routing

    AIZyphra says its approach is lossless: the same tokens still reach the same experts, with unchanged architecture, routing decisions, and training objective. By reorganizing where experts and tokens live, it reduces the communication needed to perform the same computation.

  20. ZyphraOfficialAI score38

    Zyphra reports up to 2.63x faster MoE token exchange in Megatron-LM

    AIZyphra reports that its MoE training optimizations speed up token exchange by 1.16x to 2.63x and full training steps by up to 1.41x in Megatron-LM on 8 to 64 GPUs. The gains are largest when each token uses more experts and those experts span several nodes.

  21. ZyphraOfficialAI score32

    MoE training spends 45-60% of step time on cross-node token exchange

    AIIn Zyphra's runs, MoE token exchange between experts consumed 13-24% of step time on one node and 45-60% across four nodes. Because experts are spread across GPUs and nodes, tokens must be sent to their experts and returned, making this communication a major training cost as models scale.

    Image from @ZyphraAI's post
  22. The Verge · AINewsAI score52

    Google's experimental AI Edge Foresight transcribes meetings fully offline on Mac

    AIGoogle has released AI Edge Foresight, a free experimental note-taking app that transcribes meetings and audio files entirely offline on macOS. It runs on the on-device EmbeddingGemma 2 model and turns shorthand notes into polished notes based on the transcript. Google says files, meeting audio, and notes never leave the computer, and the app is currently optimized only for Macs with Apple Silicon.

  23. CohereOfficialAI score20

    Cohere hosts live webinar on future of search and retrieval

    AICohere is hosting a live webinar on the future of search and retrieval, covering Embed 5, Parse 5, and its new retrieval methodology, RCP-nDCG. The post is a livestream announcement and does not include details of the methodology or model performance.