Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Oct 8

Oct 8Thu
  1. LlamaIndex 🦙OfficialAI score8

    Why LlamaIndex defaults to Markdown output for document parsing

    AILlamaIndex says Markdown is its default output for document parsing because it preserves headings, lists, and tables, which helps models read content correctly. The post notes that parsers can extract every word yet lose which column a number belongs to, forcing models to guess. For tables with merged headers, LlamaIndex switches to HTML.

    Image from @llama_index's post
  2. Aravind SrinivasXAI score22

    Perplexity Decider ranks first on DecisionBench at lowest cost

    AIPerplexity's Decider V1.1 ranked first on DecisionBench while also having the lowest cost, according to a post highlighting the result. The benchmark results cited include 949 shared text cases, 93.9% accuracy, a 534 ms median latency, and $0.016 per 1k decisions.

  3. Stanford HAIOfficialAI score22

    Stanford HAI leaders urge keeping people central as AI transforms research

    AIStanford HAI associate directors Risa Wechsler and Russ Altman told incoming Stanford students, faculty, and staff that AI agents can help researchers write code and tackle more ambitious questions. They stressed that AI-generated results need rigorous, reproducible methods, measured uncertainty, and careful attention to missing data, systematic errors, and biased models. Altman also argued that labs should preserve mentorship and interdisciplinary collaboration while adopting AI tools.

  4. Hacker News · Show HN, AI (20+ points)BlogAI score23

    Show HN: Jevman lets AI models play Pac-Man against the arcade ghosts

    AIJevman is an open-source Pac-Man benchmark where AI models play 100 games each against the classic scripted ghosts. Each model gets a maze state at every junction and returns a direction probability, with answers over 2 seconds replaced by a backup rule. Community models can join the leaderboard by submitting games that CI replays to verify their scores.

  5. LangChainOfficialAI score34

    LangChain's Restock agent buys office supplies through Slack with approval

    AILangChain has built Restock, an office supply agent that works inside Slack and can find real products, prepare purchases, and pay for them. A person approves each order, which is reviewed in Slack and approved through Stripe's Link agent wallet, built on MPP and Managed Deep Agents.

    Video from @LangChain's post
  6. GoodfireOfficialAI score21

    Goodfire's probes run during inference with no added latency

    AIGoodfire reports that running its probes during model inference maintains the same throughput with no added latency. The company attributes this to infrastructure engineering, including kernel-level optimizations and a custom inference server.

  7. Daniel HanXAI score38

    Unsloth adds OS-level sandboxing for Linux, Mac, and Windows

    AIUnsloth now supports OS-level sandboxing on Linux via bwrap, on Mac via seatbelt, and on Windows via Microsoft's MXC. Per-tool-call latency is under 100ms across all three, and its software-style sandboxing with regex AST checks adds about 3ms. The Windows integration was built in collaboration with Microsoft.

  8. Latent SpaceBlogAI score59

    Periodic Labs argues AI scientists need physical experiments, not just more data

    AIPeriodic Labs' Liam Fedus and Ekin Dogus Cubuk explain why scientific discovery differs from math and coding, and why experiments remain the ground truth. They describe reinforcement learning grounded in physical experiments, AI-driven materials characterization, and the view that failed experiments can be valuable training data. The transcript was truncated before the discussion of giving lab instruments "140 IQ" was completed.

  9. SunoOfficialAI score22

    Suno launches Albums for bundling songs into full releases

    AISuno announced that Albums are now live, letting users combine songs into a full release, set artwork, arrange the tracklist, and publish when ready. Existing playlists can be converted into Albums without rebuilding them from scratch.

    Video from @suno's post
  10. Google ResearchOfficialAI score14

    Google Research demos EnvHarness for co-evolving LLM agents and environments at COLM 2026

    AIGoogle Research is presenting EnvHarness, a flexible framework that enables co-evolution between LLM agents and their training environments, at the #COLM2026 Google booth #107 today at 11:00 AM PT. The post notes that static environments limit agent growth, and EnvHarness is described as a plug-in architecture that dynamically reshapes environment behaviors to improve reinforcement learning and adaptability.

  11. Google GemmaOfficialAI score27

    EmbeddingGemma 2 developer guide released by Google

    AIGoogle Gemma has published a developer guide for EmbeddingGemma 2, with code snippets to help developers start searching beyond text. The post directs readers to the full guide on the Google Developers Blog.

  12. Google GemmaOfficialAI score44

    Google publishes a developer guide for EmbeddingGemma 2 multimodal embeddings

    AIGoogle Gemma announces a developer guide showing how to embed text, code, images, video, audio, and interleaved inputs with EmbeddingGemma 2 using the sentence-transformers library. The guide outlines a four-step workflow: loading the model, embedding text and code with task prompts, embedding multimodal inputs, and optionally truncating dimensions with Matryoshka.

    Image from @googlegemma's post
  13. elvisXAI score34

    Monsoon ASR dataset cuts Bengali Whisper word error rate to 7.65%

    AIVoice Arena's Monsoon ASR dataset fine-tuned Whisper Medium on Bengali FLEURS, reducing LLM word error rate from 85.27% to 7.65%. The corpus spans 100,000 hours across 50 languages, and Voice Arena says more than 80 organisations have asked to license it since its launch a week ago.

  14. AWS Machine Learning BlogOfficialAI score27

    Share SageMaker HyperPod GPU clusters across teams with isolation and fair scheduling

    AIAWS published a reference architecture for running multiple teams on one Amazon SageMaker HyperPod EKS cluster, with each team isolated in its own Kubernetes namespace. The design combines AWS IAM Identity Center for authentication, per-team SageMaker AI domains, HyperPod Task Governance for fair resource allocation, and namespace-level cost allocation for per-team spend visibility.

  15. jasonXAI score8

    Jason Liu on learning CAD and 3D printing with AI as a magic tool

    AIJason Liu says he lacked the vocabulary, not the ambition, to design objects in CAD for 3D printing, so he learned printing, part efficiency, and shape terminology. He argues AI is like magic that still requires learning the right spells and incantations, and that it has made him want to understand the subject more deeply.

  16. MarkTechPostNewsAI score58

    JetBrains releases Mellum2.1, a 12B MoE open model for coding agents

    AIJetBrains has released Mellum2.1, a 12B mixture-of-experts thinking model with 2.5B active parameters, under Apache 2.0 on Hugging Face. Post-training reinforcement learning in real software repositories raised SWE-bench Verified from 2.0 to 47.0, according to JetBrains' self-reported results. Qwen3.5-9B still leads on SWE-bench Pro, GPQA Diamond and AIME, and GGUF builds start at 7.0 GB for local use.

  17. elvisXAI score22

    Interface ring lets users control AI agents by voice from hand

    AINatura AI's Interface is a ring that lets users press and hold to speak requests to AI agents such as Claude Code, Codex, or Hermes, then release to send them. The post argues that screenless interfaces may define the next phase of agent use, since handing work to agents is currently slowed by pulling out a phone. Early-adopter pricing is $99, with shipping slated for January.

  18. The Robot ReportNewsAI score34

    Jabil Says Humanoid Robots Are Moving Toward Tens-of-Thousands Production Volumes

    AIJabil senior director Thomas Brown says humanoid robots are entering a phase of tens of thousands of units, where manufacturability, cost structure, and quality become central. He says Jabil works with developers to cut costs for scale, while compute and memory prices remain a pain point, and that humanoids make sense in factories and warehouses while mobile arms still suit high-speed tasks.

  19. Unsloth AIOfficialAI score44

    Unsloth adds Windows OS-level sandboxing via Microsoft's mxc

    AIUnsloth now supports OS-level sandboxing on Windows by integrating Microsoft's open-source mxc repository for sandboxed code execution. The integration adds under 100 ms of overhead, according to the post. A setup guide is available in Unsloth's documentation.

    Image from @UnslothAI's post
  20. Goodfire ResearchOfficialAI score57

    Goodfire deploys probe-based cyber monitors on Kimi K3 with a judge cascade

    AIGoodfire Research describes probe-based cyber monitors for Kimi K3 and GLM 5.3 deployed on a production inference stack. The probe filters suspicious exchanges before an LLM judge reviews them, reaching about 93% recall at a 5.5% benign-session interruption rate at roughly 50x lower judge cost. In FAR.AI's red-teaming, the monitor reduced universal jailbreaks to zero across 140 tested strategies.

  21. Vercel DevelopersOfficialAI score36

    StepFun's Step 5 Preview model now available on Vercel AI Gateway

    AIVercel says StepFun's flagship Step 5 Preview, built for agentic coding, research, and finance, is now live on AI Gateway. The model offers a 1M-token context window, accepts text and image input, and uses a 600B-parameter mixture-of-experts design with 27B parameters active.

  22. Grok BotOfficialAI score38

    Grok Bot can now help run Shopify stores

    AIGrok Bot can now connect to Shopify to check orders, track inventory, and keep product listings up to date. The feature lets merchants ask it questions about their business after linking their store. Shopify's related announcement says the new connectors also allow merchants to build teams of agents to help run their business.

    Image from @bot's post
  23. ClaudeDevsOfficialAI score33

    Anthropic credits work with Messages API, Managed Agents, and Agent SDK

    AIAnthropic's credits can be used with the Messages API, Claude Managed Agents, and the Agent SDK. They cannot be used for interactive Claude Code sessions, but they also apply in third-party harnesses that accept a Claude API key.

  24. ClaudeDevsOfficialAI score46

    Anthropic adds monthly API credits for Max and Team plans

    AIAnthropic now provides monthly Claude Platform API credits to Max and Team subscribers: $100 for Max 5x, $200 for Max 20x, and up to $500 pooled for Team. The credits can be used on any Claude model, including in code or third-party harnesses.