SGLang adds a decision API, available in nightly and v0.5.21
AISGLang says its new decision API can now be tried using the nightly image. The feature will also be included in v0.5.21, scheduled for this Thursday, according to the post.
Updated
Updated
Items with an AI score under 20 are hidden. Show low-relevance items
AISGLang says its new decision API can now be tried using the nightly image. The feature will also be included in v0.5.21, scheduled for this Thursday, according to the post.
AISGLang says it turned Qwen3.8-27B into a multimodal decision model that beat Pokémon FireRed's Elite Four and champion with sub-100 ms decisions from live game state. It introduces a native /v1/decisions endpoint for turning LLMs and VLMs into classification and scoring models. A /v1/systemone endpoint is also added so Jev-like open models can work with the TypeSafe SDK.
AIChatGPT now lets users publish shareable profiles that bring their Sites and plugins together in one place for others to find and reuse. Teammates can discover shared skills within their workspace, and the feature is available to Free, Go, Plus, Pro, and Business users. ChatGPT Enterprise, Edu, and Healthcare plans will get it soon.

AIHugging Face released the Open TTS Leaderboard, which evaluates open-source text-to-speech models using objective metrics instead of arena-style human votes. It measures intelligibility via WER and CER using Qwen3 ASR, speed via RTFx and time-to-first-audio on an H200 GPU, and speaker similarity via WavLM embeddings. The leaderboard covers multilingual results and voice cloning, and it is intended to complement, not replace, human preference rankings.
AINumerical differences between a rollout engine and a trainer can make reinforcement learning collapse even when algorithm and data stay identical. In a GLM 5.2 experiment, reward fell from about 0.9 to under 0.2 around step 20 without alignment, while aligned numerics kept reward stable over 25 steps. A Qwen3.5-MoE investigation traced a significant mismatch to how expert outputs were combined, and router replay alone was judged insufficient.
AIThe vLLM project announced a strong presence at PyTorchCon North America, with core maintainer and Inferact CEO Simon Mo giving the keynote on scaling open frontier inference infrastructure. Other vLLM maintainers, including Nick Hill and Red Hat AI engineers, will lead a developer session and talks on agentic inference and attention.
AIYuchen Jin cites an Anthropic blog post in which GLM-5.3 says "My job is to cause deaths quietly." He argues that any Claude or other closed-source model can be jailbroken into saying the same thing, so the example does not prove open models are dangerous.

AIThe OpenClaw Foundation announced OpenClaw Enterprise, an open-source enterprise control plane for persistent agents, in collaboration with Red Hat, NVIDIA, and OpenAI. The product is built to run on an organization's own infrastructure and will always be free for organizations to use.
Why it matters: The announcement names its collaborators and deployment model, which helps organizations judge how the enterprise control plane would fit their own infrastructure.

AIJulien Chaumond, who owns Hugging Face, recommends a DGX Spark guide written by @0xSero and published on the Hugging Face blog. The post itself gives no further technical details about the handbook's contents.

AIBAAI released AREX-2, a 27B-parameter long-horizon agent model that improves solutions over multiple test-time rounds by proposing, measuring, reflecting, and revising. It was trained on machine-learning and algorithmic-programming tasks with verifiable feedback, and the source reports that this self-improvement transfers to deep research. The model is Apache License 2.0 licensed and has a 262,144-token context length.
Why it matters: The source compares AREX-2 against closed and open models on coding and deep-research benchmarks, showing how test-time self-improvement is measured across task types.
AIOpenAI's "Sign in with ChatGPT" lets people log into third-party apps and use the tokens they already pay for. The announcement came in a post reacting to OpenAI's DevDay keynote, timestamped 10:46.
AIYuchen Jin notes a pricing tradeoff: a Fast tier at 2x speed for 2x price and an Ultrafast tier at 8x speed for 6x price. He proposes rolling out an even faster "Ultra-ultrafast" endpoint for open-source models.

AIOpenAI says GPT-6.1 Sol is available starting today to all Plus, Pro, Business, Enterprise, and Edu users. The model is offered in ChatGPT Work and Codex, and the post links to OpenAI's introduction page.
Why it matters: The post specifies which plan tiers gain GPT-6.1 Sol and in which products, showing where the new model reaches users directly.
AIClément Delangue says Hugging Face's acquisition by Nvidia lets the company hire people it could not afford as a small startup. He frames the move as giving them a decade to help open-source AI win, and invites interested candidates to message him directly.
AIPerplexity has open-sourced Bumblebee, a read-only scanner for macOS and Linux that checks developer machines for risky packages, extensions, and AI tool configurations. When connected to Computer, it can trigger deeper scans whenever a new supply-chain risk emerges. The post says Computer reviews findings from Bumblebee and Numbat to propose better detection rules, and humans approve every change before it ships.
Why it matters: The post shows how a read-only scanner fits into a human-approved pipeline that updates detection rules after supply-chain risks emerge, useful for teams planning developer machine security.
AIResearchers from Tsinghua NLP and collaborators show that on-policy distillation with a single training query recovers 87% of full-data gains on math, reaching 68.5 versus 69.8 by step 300. The paper attributes the slow progress to how fast the student absorbs the teacher's signal rather than to dataset size. Code and the paper are publicly available on GitHub and Hugging Face.
Why it matters: The paper isolates training data from the algorithm, showing one query nearly matches full-data on-policy distillation, which reframes where post-training gains come from.

AIModelScope announced IQuest-Q1, a 320B MoE model with 15B active parameters and a 512K context window for agentic coding. The post reports scores of 84.5 on CyberGym, 83.2 on Terminal-Bench 2.1, 64.6 on DeepSWE v1.1, and 63.0 on NL2Repo, and says weights are released under the IQuest-Q1 License.

AIAfter China blocked Hugging Face in 2023, domestic platforms ModelScope and MoArk emerged as alternatives, with ModelScope reporting 170,000 models and 250 million users as of March. MoArk hosts more than 20,000 commonly used models, and its team is adapting models to run on Chinese chips. Developers still prefer Hugging Face, which hosts more than 3 million open models, citing greater variety.
AIShanghai AI Laboratory's Intern-Decision family of 0.8B, 2B, and 4B multimodal models averages 79.38, 84.68, and 90.02 across seven decision benchmarks. Intern-Decision-4B scores 88.74, surpassing Jev while achieving better probability calibration. Reported mean latency is 33.98, 33.28, and 44.16 ms, versus 109.70 ms for Jev in the same local HF setup.

AIInternLM's AutoVerifier, built on Qwen3_5MoeForConditionalGeneration with about 68 GiB of weights across 40 safetensors shards, evaluates natural-language mathematical proofs, explains errors, and identifies the earliest incorrect step. It serves as the automatic grader for AdvancedMathBench's ProverBench, which accepts a proof only when all eight judgments report -1. The model is a learned grader rather than a formal proof checker and can make errors.
AIOpenBMB announced that MiniCPM-o 4.5 is now supported in SGLang Omni v0.1.7, giving developers more flexibility to run and build with the model. The background release notes add that MiniCPM-o 4.5 brings multimodal input and speech output to the runtime. MiniCPM-o and MiniMax-Music3 also gained Intel XPU support in the same release.
AIvLLM announced day-0 support for IQuest-Q1, a 320B-parameter MoE model with 15B active per token, 256 experts with 8 active, and a 524,288-token context. The post credits existing vLLM features such as the hybrid KV cache coordinator, sinks attention path, and EAGLE speculative decoding with probabilistic draft sampling. The linked material includes a Docker image and vllm serve commands, with and without recursive MTP.

AIThomas Wolf, owner of the Hugging Face account, posted the single word "impressive" in response to a quoted post. The quoted post reports a new NanoGPT training record of 39.9s, down 27.7s from the prior 67.6s, achieved through per-flop optimizations such as sampled softmax and sparse updates.
AISGLang says it has Day-0 support for IQuest-Q1, an open-source sparse MoE model with 320B total and 15B active parameters for coding and agentic tasks. The post includes a single-node serving command for H200 GPUs in BF16, using tensor parallelism of 8, EAGLE speculative decoding, and the iquest_q1 reasoning and tool-call parsers. The image marks the command as not verified.

AIArtificial Analysis has open-sourced AA-AgentPerf-Local, a tool that replays recorded agent trajectories to measure inference speed on laptops and workstations. Initial results cover NVIDIA DGX Spark, NVIDIA GeForce RTX 5090, AMD Ryzen AI Halo, and MacBook Pro M5 Pro, with the RTX 5090 fastest for models that fit its 32 GB. The source states the tool and leaderboard will expand to more hardware, frameworks, and models.
Why it matters: The source gives per-system completion times and memory bandwidth figures, letting readers compare local hardware for running agentic workloads.
AIAnthropic reports that Zhipu AI's GLM-5.3 can autonomously build end-to-end cyber exploits and is released without meaningful safeguards against misuse. In its simulated tests, attackers bypassed the model's safeguards 64% to 100% of the time using simple techniques, while the same attacks failed against safeguarded Claude models. Anthropic also cites an NIST CAISI assessment calling GLM-5.3 the most cyber-capable open-weight model released to date.
Why it matters: The report shows how open-weight safeguards fail under simple bypasses, offering concrete test figures for judging misuse risk in released models.
AIAudio8 ASR Infinite transcribes Chinese and English audio of unlimited length using a rolling KV Cache that avoids accumulated drift. At a 480 ms delay, it reports 1.75 CER on AISHELL-1, 2.89 on AISHELL-4, and 3.04/6.81 WER on LibriSpeech test-clean/test-other. The preview release is under Apache 2.0, with deployment through an adapted vLLM stack.
AIKhazix (Shuzi Shengming Kazike) says the monthly-active-million AI news site AIHOT is now open source on GitHub, including its collection workflow, curation scoring, clustering mechanism, and all production prompts. He says the release is meant to hand the project to others, since readers have asked for versions for industries such as gaming, law, HR, and finance.
AIThe vLLM blog guide explains how separating prefill and decode, and moving tokenization to a CPU-only render tier, can keep token streams from stalling under load. In a two-L40S test on Qwen2.5-7B, collocated p99 inter-token latency reached 169 ms at 0.4 req/s while disaggregated serving stayed between 25 and 52 ms. The guide notes that the gain depends on fast KV cache transfer, and it includes setup code for NIXL-based serving and the render/derender API.
AIFei-Fei Li will join AMD as Executive Vice President and Chief Scientist, working directly with CEO Lisa Su. World Labs will join AMD to form a frontier research organization, co-led by Justin Johnson and Ben Mildenhall, focused on an end-to-end open AI ecosystem spanning hardware, software, platforms, and widely accessible open models.
Why it matters: The announcement shows how a leading AI lab's team is folding into a chipmaker, with a stated plan for an open AI ecosystem spanning hardware and models.
AIAMD announced a definitive agreement to acquire World Labs, adding its team of researchers and model experts to help shape future AI infrastructure. The company said the move will strengthen the open AI ecosystem. AMD also welcomed Fei-Fei Li and the World Labs team to the company.

AIAndrew Ng says OpenWorker, his open-source agent harness for cybersecurity workflows, will run each agent's commands inside a sandbox built on Nvidia OpenShell. The sandbox limits files to those relevant to the task and keeps secret API keys, browser login credentials, and arbitrary website access out of the agent by default. Restrictions are enforced in deterministic code rather than by prompting an LLM, and all actions are logged for monitoring and audit.
AIK3-Node is a graph neural network library built natively on Keras 3, with models that run on JAX, PyTorch, and TensorFlow with hardware acceleration including Apple Silicon and TPU. According to the post, it achieves 100% public API parity with PyG and incorporates foundation models and architectures from Spektral and StellarGraph.
AIMistral CEO Arthur Mensch argues that only an open ecosystem can guarantee AI safety. The post responds to NVIDIA's announcement of the Open Agent Safety Platform, which brings together OpenShell and Sentry with over 100 industry partners.
AIwhich links to a page for trying the model. The quoted announcement describes it as the second model in the Claude 5.5 family, a clear upgrade over Sonnet 5 that runs more than 30% faster and costs up to 30% less for most work.
Why it matters: The post shows the model's availability inside v0 and links a test entry point, though it gives no performance details beyond the quoted claims.
AIGoogle and Kaggle launched the Gemma 4 Developer Agent Competition, challenging developers to build coding agents that run offline on consumer hardware rather than relying on cloud API models. The total prize pool exceeds $110,000, with entries due November 2, 2026, and a starter kit is available on Kaggle.

AIRadixArk has released Miles v0.1.1, adding multi-LoRA with Tinker API compatibility so multiple training jobs can share one base model. The update also supports agentic training with harnesses like Claude Code and runs Harbor tasks in sandboxes including AgentENV, Daytona, E2B, and Modal. It further reduces memory needs for training larger models on validated NVIDIA and AMD GPUs and adds stable support for Qwen3.8-Flash-Next, GLM-5.3-Flash, and Kimi-K3.

AIUnsloth Desktop can now serve local Laya decision models through a Jev-compatible API, shown in a real-time packing demo where suitcase items update as the user types. The demo runs through Unsloth's Decision API, and the team says more optimizations are coming to speed up local hardware performance.
AIGoogle Cloud argues startups should combine open-weight models with frontier APIs rather than routing every request to one frontier model. It cites Gemma 4, which spans five sizes including a 31B dense model and a 26B A4B Mixture-of-Experts model, released under Apache 2.0. The article's examples report a 44% latency drop for Cue, from 876 ms to 488 ms, and a $0 server cost for BetterSpeak's on-device Gemma 4 E2B.