Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 2

Oct 2Fri
  1. Baseten BlogOfficialAI score70

    Baseten's agent-built VibeQwen engine beats vLLM on Qwen-3.6 decode speed

    AIBaseten tested the MetaInfer skills-only approach by having Claude Code build an inference engine, VibeQwen, for Qwen-3.6-35B-A3B in NVFP4 on a single B200. On single-stream text, VibeQwen decoded 90% faster than a tuned vLLM 0.25.1 deployment (1,792 vs. 943 TPS) and cut time to first token from 28 ms to 12 ms, with a 71% throughput gain at concurrency 32. The author notes this was an outcome-focused run that allowed some numerically different outputs as long as accuracy stayed at or above the BF16 baseline.

    Why it matters: The post tests a skills-only inference engine method on a real model and states the speed and accuracy constraints used, helping readers judge how far such automated optimization can be trusted.

  2. Aravind SrinivasXAI score44

    Perplexity Computer builds a 3D map of NYC restaurants

    AIPerplexity's Computer built a 3D map of nearly 26,000 restaurants and cafes across New York City's five boroughs. Users can search by dish or neighborhood and step inside places such as Peter Luger and Grand Central Oyster Bar. The post frames such projects as ones an agent can run for hours to produce something substantial.

  3. Aravind SrinivasXAI score62

    Perplexity open-sources models, an inference engine, and security tools

    AIPerplexity has released several open source projects, including the pplx-decider-v1-27b multimodal decision model, the pplx-embed-v2-context-9b-preview contextual embeddings model, and the Lily local inference engine for Apple silicon. The post also lists the 0.6B on-device PII-Tracer classifier with its PII-TRACE benchmark, the WANDR research agent benchmark, and the Numbat and Bumblebee security tools, and says more open source releases are coming soon.

    Why it matters: The post lists several named open source releases with specific benchmark figures, helping readers scan which tools and models Perplexity has recently published.

  4. ClaudeDevsOfficialAI score29

    Anthropic adds "You should Know" plugin to Claude Code

    AIAnthropic is adding a new Claude Code plugin called "You should Know" that scans Claude's output for important information users might otherwise miss. It can be enabled with the command /plugin enable cc-plugin-you-should-know@builtin.

    Video from @ClaudeDevs's post
  5. Claude Code · GitHub ReleasesOfficialAI score38

    Claude Code v2.1.288 is released with fixes and new controls

    AIAnthropic released Claude Code v2.1.288, adding $.ui.selection() for mods, a built-in gh api for cloud sessions without the GitHub CLI, and --max-findings for /code-review. The release also fixes many issues, including mid-response API timeouts, resume and compaction bugs, and auto mode denials and model switching on Bedrock and Mantle.

  6. PyTorch BlogOfficialAI score47

    Helion Linear Backend Boosts vLLM Hopper GPU Inference Throughput Over CUTLASS and DeepGEMM

    AIThe vLLM team integrated Helion, a PyTorch-native kernel DSL, into vLLM's linear backend, using per-shape autotuning to select among Standard GEMM, Split-K, and Swap-AB variants. On NVIDIA Hopper GPUs, the Helion backend outperformed the default CUTLASS and DeepGEMM backends across the evaluated models, with more than 10% throughput gains for some workloads. The work focuses on FP8 and INT8 quantized GEMM.

  7. Amazon ScienceOfficialAI score22

    Amazon and Georgia Tech launch joint science hub for robotics and engineering research

    AIAmazon and Georgia Tech launched a joint science hub to support research in robotics, industrial and systems engineering, aerospace technology, and foundational and emerging technologies. Georgia Tech researchers will also use Sprout, the robotics development platform from Fauna Robotics, to advance work in human-robot interaction and perception.

  8. GitHub Copilot ChangelogOfficialAI score34

    Copilot code review gains API access and Balanced default effort level

    AIGitHub Copilot code review can now be requested through the REST and GraphQL APIs, with an optional review effort level set per request. Balanced became the default review effort level for new and existing repositories and organizations as of September 28, 2026, while users who explicitly selected Lite keep that setting. The changes are generally available to Copilot Pro, Pro+, Max, Business, and Enterprise plans.

  9. AI at MetaOfficialAI score21

    Meta's Muse Spark helps find exception to biology-inspired algebra rule

    AIMeta researchers used Muse Spark to uncover a counterexample to a proposed rule about mathematical structures inspired by biology, then developed and refined an alternative characterization. The work is detailed in a linked paper on solvable evolution algebras and a conjecture by Garcia-Martinez and Perez-Rodriguez.

  10. AI at MetaOfficialAI score22

    Muse Spark helps prove finite-time blow-up in a laser-inspired wave model

    AIWith help from Muse Spark, researchers proved that a wave in a laser-inspired model must blow up in finite time under the conditions studied. The result comes from a tug-of-war between one effect squeezing the wave inward and another spreading it out. The paper is titled finite-time blow-up of radial negative-energy solutions for the mass-critical biharmonic nonlinear Schrödinger equation.

    Image from @AIatMeta's post
  11. AI at MetaOfficialAI score61

    Meta shares six math papers from mathematician-AI collaborations on open problems

    AIAI at Meta says mathematicians used Muse Spark 1.1 and Muse Spark 1.2 in Thinking Mode through the standard meta.ai chat interface to find solutions to open problems. The company is sharing six resulting papers, each marking which passages were drafted primarily by humans or AI, with mathematicians guiding the work and a second group reviewing it.

    Why it matters: The post shows AI models helping mathematicians on open problems, with human guidance, peer review, and disclosure of AI-drafted passages, which clarifies how such collaborations are documented.

  12. eric zakariassonXAI score47

    xAI releases experimental TypeScript SDK with Grok models and tools

    AIxAI has released an experimental TypeScript SDK, installable via npm install @xai-official/sdk, that covers text, voice, image, and video in one package. It provides access to the latest Grok models along with server-side tools including real-time X search, web search, code execution, and remote MCP.

    Video from @ericzakariasson's post
  13. DatabricksOfficialAI score44

    Omnigent: open-source meta-harness coordinating Claude Code and Codex agents

    AIDatabricks' new open-source meta-harness, Omnigent, lets multiple coding agents such as Claude Code and Codex share sessions, rules, and security policies in one system. A walkthrough by @leonvz demonstrates forking work across agents, multi-agent review and debate with Debby, and splitting implementation across subagents with Polly.

    Video from @databricks's post
  14. ChatGPTOfficialAI score60

    ChatGPT adds Finances for subscriptions, budgets, credit, and investments

    AIChatGPT now offers Finances, which can find forgotten subscriptions, flag unfamiliar or duplicate charges, and track recurring bills that have increased. It also provides weekly updates, monthly spending breakdowns, budget building, credit score tracking, debt payoff planning, emergency fund estimates, and investment mix and concentration views. Users can access it at

    Why it matters: The post lists concrete finance features across budgeting, debt, and investments, showing how a general assistant is expanding into personal money management.

  15. ChatGPTOfficialAI score60

    Finances in ChatGPT rolls out to Free and Go users in the U.S.

    AIChatGPT's Finances feature is rolling out to Free and Go users in the U.S. Users can securely connect their accounts through Plaid and Experian to get answers based on their own financial information.

    Why it matters: The post names the rollout scope and the account connection method, which helps readers judge how the feature handles personal financial data.

    Video from @ChatGPT's post
  16. Harrison ChaseXAI score38

    LangSmith Custom Apps lets teams build trace review UIs in-workspace

    AILangChain's LangSmith Custom Apps lets agent teams build their own review UI over their traces and publish it directly into the workspace. Developers build the interface on their LangSmith data, while hosting, authentication, and permissions are handled by the platform.

  17. CursorOfficialAI score42

    Cursor's Rollouts detects deployment regressions and launches cloud agent fixes

    AICursor introduced Rollouts, a tool that writes a monitoring plan and watches changes as they deploy to catch regressions before users see them. When Rollouts detects a regression, it identifies the offending PR and opens an issue, and one click starts a cloud agent to fix it. Rollouts usage credits are included through Oct 3.

  18. GitHubOfficialAI score44

    GitHub Copilot adds Project HydraFusion and new models to model picker

    AIGitHub has made the Project HydraFusion research preview available in the GitHub Copilot app and @code, where it orchestrates multiple models rather than acting as a single model. New models from Anthropic (Fable 5.1 and Opus 5.5) and OpenAI (GPT-6.1 Sol) are also now selectable in the Copilot model picker.

  19. Together AIOfficialAI score34

    Together AI shares how its team uses AI to boost collective productivity

    AITogether AI's CPO and product team outlined how they use AI to make the whole team more productive, not just individuals. The approach includes a shared context repo readable by any AI harness, cutting half a day of research to about 5 minutes, and evals that test their product the way agents actually use it.

  20. Ant LingOfficialAI score27

    Ling-3.1-flash now available free on OpenRouter

    AIAnt Ling has made Ling-3.1-flash available on OpenRouter at no cost, inviting users to try it and share feedback. The post provides no details on model size, benchmarks, context length, or pricing beyond the free access.

  21. Harrison ChaseXAI score26

    LangChain improves memory for managed Deep Agents in enterprise settings

    AIHarrison Chase says LangChain is improving memory in managed Deepagents, noting that memory is difficult to get working well in company settings. The linked background post describes user memory in Managed Deep Agents 0.8, which lets an agent remember the people it works with.

  22. ElevenLabsOfficialAI score32

    ElevenLabs earns FedRAMP 20x Class A certification for federal voice AI

    AIElevenLabs has achieved FedRAMP 20x Class A certification, covering ElevenAgents and its Text to Speech and Speech to Text APIs when run in Zero Retention Mode with US data residency. The company says the published evidence package and FedRAMP Marketplace listing should shorten security reviews for public sector and regulated buyers. The certification extends ElevenLabs for Government, its dedicated offering for federal agencies.

    Image from @ElevenLabs's post
  23. Microsoft AIOfficialAI score41

    Microsoft MAI voice models now available on LiveKit for agents

    AIMicrosoft's MAI-Transcribe-2-Streaming, MAI-Voice-2.1, and MAI-Voice-2.1-Flash are now live on LiveKit for building voice agents. LiveKit says MAI-Transcribe-2-Streaming debuts at #1 on the Artificial Analysis accuracy leaderboard, and suggests pairing it with MAI-Voice-2.1-Flash for efficient, expressive voice agents.

  24. Microsoft CopilotOfficialAI score20

    Copilot Code lets more people build apps and workflows

    AIMicrosoft's Copilot Code is designed to help more people turn ideas into apps, workflows, and solutions for their work. Microsoft Copilot EVP Jacob Andreou discusses how the product expands who gets to build.

    Video from @MSFTCopilot's post
  25. Google AIOfficialAI score62

    Google launches Project Suncatcher prototype satellite to test TPUs in orbit

    AIGoogle AI announced that its Project Suncatcher prototype satellite, built with Planet, has launched into orbit on SpaceX's Transporter-18 rideshare mission. The initial mission will gather data on how Google TPUs handle the physical stress and extremes of spaceflight. The post says low Earth orbit systems could generate up to 8x more solar power than on Earth, and that future work may link multiple satellite constellations for scaled machine learning.

    Why it matters: The post explains a space-based machine learning prototype and why orbit's near-constant sunlight matters, which helps readers weigh the idea's practical potential.

    Video from @GoogleAI's post
  26. Kilo (acq. by Anaconda)OfficialAI score36

    Ling 3.1 Flash is free in Kilo Code until October 13

    AIKilo Code is offering Ling 3.1 Flash for free until October 13, with the model served by Novita Labs. Ant Ling's background post describes the model as roughly 560B total parameters with about 25B active per token and a context window of up to 1M tokens. Ant Ling says it scores 1,673 Elo on GDPVal-AA v2.1, 75.16 on FrontierSWE, and 65.35 on HealthBench Professional, and plans to open-source it soon.