Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 2

Oct 2Fri
  1. Aravind SrinivasAI score62

    Perplexity open-sources models, an inference engine, and security tools

    AIPerplexity has released several open source projects, including the pplx-decider-v1-27b multimodal decision model, the pplx-embed-v2-context-9b-preview contextual embeddings model, and the Lily local inference engine for Apple silicon. The post also lists the 0.6B on-device PII-Tracer classifier with its PII-TRACE benchmark, the WANDR research agent benchmark, and the Numbat and Bumblebee security tools, and says more open source releases are coming soon.

  2. Claude Code · GitHub ReleasesAI score38

    Claude Code v2.1.288 is released with fixes and new controls

    AIAnthropic released Claude Code v2.1.288, adding $.ui.selection() for mods, a built-in gh api for cloud sessions without the GitHub CLI, and --max-findings for /code-review. The release also fixes many issues, including mid-response API timeouts, resume and compaction bugs, and auto mode denials and model switching on Bedrock and Mantle.

  3. SGLangAI score38

    SGLang v0.5.20 adds Intel XPU support and faster RL rollouts

    AISGLang has released v0.5.20, bringing Intel XPU into standard releases alongside RL sampling masks that make rollouts more reliable with up to 52% faster decode. The update also adds Unified Radix Tree SWA branching-point caching, which the project says lifts cache hit rate about 20 points and cuts TTFT by roughly one-third, plus up to 12.5× faster ROCm model loading. New models named in the release include GLM-5.3-Flash, Qwen3.8-Flash-Next, K2 Horizon, Hy4-Preview, FastH3, and VDN-H3.

  4. SGLangAI score39

    SGLang adds a scoring API and multi-item scoring for decision models

    AISGLang's update adds a /v1/score endpoint that returns scores for requested labels such as Yes/No or A/B/C, avoiding the label loss of generate with top-k logprobs. Its multi-item scoring computes shared context once and keeps each candidate isolated, with 16-candidate p95 on Qwen3-8B dropping from 54.1 ms (Generate) to 20.6 ms.

  5. SGLangAI score28

    SGLang's /v1/decisions API turns Qwen3.8-27B into a decision model

    AISGLang demonstrated Qwen3.8-27B as a multimodal decision model that beat Pokémon FireRed's Elite Four and champion with sub-100 ms decisions from live game state. The company says its native /v1/decisions API lets LLMs and VLMs be used for classification and scoring. It also announced /v1/systemone for running Jev-like open models with the TypeSafe SDK.

  6. SGLangAI score58

    SGLang v0.5.21 adds native decisions API and new model support

    AISGLang has released v0.5.21 with a native Decisions API that turns an LLM or VLM into a low-latency classifier and scorer. The release also lets /v1/score rerank search or RAG results in one call, lets PD instances switch between prefill and decode without restarting, and adds support for models including DeepSeek-V4.1 Flash, Kimi K3, and GLM-5.3-Flash on AMD MI355X. The announcement reports a 22% faster first token on long prompts for DeepSeek-V4.1 Flash and 20.6% higher prefill throughput for Kimi K3 in PD serving.

    Image from @sgl_project's post
  7. LiveKitAI score23

    AssemblyAI Universal 3.6 Pro now live in LiveKit Inference

    AIAssemblyAI's Universal 3.6 Pro speech-to-text model is now available in LiveKit Inference, with 45% fewer wrong yes/no confirmations and about 30% less background speech transcribed. It supports 32 languages plus code-switching and endpointing that waits out phone numbers and emails, at the same $0.45/hr price, accessible by switching to universal-3-6-pro.

    Image from @livekit's post
  8. PyTorch BlogAI score24

    PyTorch Certified Associate Gets New Four-Module Certification Pathway

    AIThe Linux Foundation Education has launched a PyTorch Certified Associate (PTCA) Certification Pathway that combines four self-paced learning modules with the PTCA exam. The pathway includes 15–17 hours of self-paced learning and hands-on labs covering tensors, data handling, model development, and performance optimization. The source recommends additional hands-on practice before taking the exam.

  9. GitHub Copilot ChangelogAI score34

    Copilot code review gains API access and Balanced default effort level

    AIGitHub Copilot code review can now be requested through the REST and GraphQL APIs, with an optional review effort level set per request. Balanced became the default review effort level for new and existing repositories and organizations as of September 28, 2026, while users who explicitly selected Lite keep that setting. The changes are generally available to Copilot Pro, Pro+, Max, Business, and Enterprise plans.

  10. DatabricksAI score44

    Omnigent: open-source meta-harness coordinating Claude Code and Codex agents

    AIDatabricks' new open-source meta-harness, Omnigent, lets multiple coding agents such as Claude Code and Codex share sessions, rules, and security policies in one system. A walkthrough by @leonvz demonstrates forking work across agents, multi-agent review and debate with Debby, and splitting implementation across subagents with Polly.

    Video from @databricks's post
  11. ChatGPTAI score60

    ChatGPT adds Finances for subscriptions, budgets, credit, and investments

    AIChatGPT now offers Finances, which can find forgotten subscriptions, flag unfamiliar or duplicate charges, and track recurring bills that have increased. It also provides weekly updates, monthly spending breakdowns, budget building, credit score tracking, debt payoff planning, emergency fund estimates, and investment mix and concentration views. Users can access it at

  12. Kilo (acq. by Anaconda)AI score36

    Ling 3.1 Flash is free in Kilo Code until October 13

    AIKilo Code is offering Ling 3.1 Flash for free until October 13, with the model served by Novita Labs. Ant Ling's background post describes the model as roughly 560B total parameters with about 25B active per token and a context window of up to 1M tokens. Ant Ling says it scores 1,673 Elo on GDPVal-AA v2.1, 75.16 on FrontierSWE, and 65.35 on HealthBench Professional, and plans to open-source it soon.

  13. Jerry LiuAI score44

    LlamaIndex Extract v2.5 hits 93–96% on dense table extraction benchmarks

    AILlamaIndex released Extract v2.5, a set of document extraction agents that it says reach 93%–96%+ accuracy on long-list extraction, including records spanning pages. The post claims the agents outperform frontier VLMs, which it says stop early, miss repeated records, and struggle to attribute values to sources, while LlamaIndex attributes every extracted value to its source. The agents are available through LlamaParse.

    Video from @jerryjliu0's post
  14. Google · AI blogAI score58

    Google recaps September 2026 AI launches, led by Gemini 4 Argon

    AIGoogle's September 2026 roundup highlights Gemini 4 Argon, a frontier model with a 1-million-token output limit aimed at complex tasks such as cybersecurity defense. Argon is rolling out first to trusted cyber defenders through the Fairwind Program, with developer, enterprise, and consumer access to follow after guardrail feedback. The post also covers Gemini 3.8 Flash, Connected Apps in Gemini, and WeatherNext 3.

  15. merveAI score36

    llama.cpp adds support for decision models on modest hardware

    AIllama.cpp now supports decision models, which route tickets, moderate content, or choose an agent's next step by returning a probability for every option. Five open models from 144M to 27B parameters are supported at launch, and the team says more will follow in the coming days. Because most decision models do not need large GPUs, they are a good fit for llama.cpp, and a Hugging Face blog post explains how to set them up.