Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Oct 9

TodayOct 9Fri
  1. Artificial AnalysisOfficialAI score22

    HeyGen Voice ranks first in Artificial Analysis assistant and customer service arenas

    AIHeyGen Voice ranks first in the Artificial Analysis Controlled Voice Arena for Assistants at 1,214 and Customer Service at 1,212. It ranks second in Knowledge Sharing at 1,158 and fourth in Entertainment at 1,178. By accent, it ranks first in English (UK) at 1,190 and second in English (US) at 1,206, behind Qwen-Audio-3.1-TTS-Plus at 1,220.

    Image from @ArtificialAnlys's post
  2. Artificial AnalysisOfficialAI score32

    HeyGen Voice scores 83.1% on pronunciation robustness benchmark

    AIHeyGen Voice scores 83.1% overall on Artificial Analysis's Pronunciation Robustness benchmark, ranking #10 of 29 models. The benchmark has human reviewers judge whether text-to-speech models pronounce challenging text correctly against pre-agreed accepted pronunciations. HeyGen Voice places #2 for preserving exact sequences at 83.3%, behind SpaceXAI TTS at 85.7%, and scores 77.0% on expanding shorthand, while Eleven v4 leads that category at 94.1%.

    Image from @ArtificialAnlys's post
  3. RadixArkOfficialAI score28

    Miles v0.1.2 adds Kubernetes support and torchtitan training backend

    AIRadixArk releases Miles v0.1.2, adding score centering for stable async RL and an experimental native Kubernetes backend. RL jobs can run as ordinary cluster workloads, and orchestration can restart while training continues. The release also adds torchtitan as a third training backend alongside Megatron and FSDP, plus support for DeepSeek-V4.1-Flash and MiMo-V2.6-Flash-RL.

    Image from @radixark's post
  4. 🚨 AI News | TestingCatalogXAI score41

    Pine AI launches Pine Computer, a cloud runtime for agentic tasks

    AIPine AI launched Pine Computer, a cloud computer, harness, and runtime layer built for agentic tasks. On the publisher's SaaS-Bench v1.1, it posts a 78.3% checkpoint score against 74.3% for Opus 5 with Claude Code, but completes fewer whole tasks, 27.4% against 31.1%. Instead of simulating clicks and screenshots, it reads web pages as structured data, and access is through a private beta waitlist.

    Image from @testingcatalog's post
  5. SantiagoXAI score44

    Pine launches agentic cloud computers with built-in AI agents

    AIPine has released a cloud computer service with a built-in AI agent that applications can control through its SDK. Developers give the agent a plain-English task, and it can use a browser, files, and a shell while the app receives notifications and final outputs. Pine's Stanley Wei says the computer is built for AI rather than humans.

    Video from @svpino's post
  6. LlamaIndex 🦙OfficialAI score12

    LlamaParse keeps nested tables and values intact in earnings decks

    AILlamaIndex says its LlamaParse keeps all 18 values from Micron's latest earnings deck under the correct headers, despite nested tables with two business units and repeated row names. The post presents this as a sample document rather than a benchmark result.

    Image from @llama_index's post
  7. elvisXAI score34

    Elvis Saravia urges builders to focus on agent harnesses and environments

    AIElvis Saravia says AI models are already smart, but they need better harnesses and environments, with major cost implications. He recommends reading a report on how Pine Computer can help teams, and says he will test it himself and share more later. The quoted post from Stanley Wei argues that real-world AI tasks remain slow, expensive and unreliable because AI runs on computers built for humans, and announces Pine Computer.

    Image from @omarsar0's post
  8. Ai2OfficialAI score13

    Ai2 describes a fair-share GPU scheduler for research budgets

    AIAi2 says managers now assign GPU-time budgets to research programs and projects. Its fair-share scheduler compares recent usage with those budgets and moves work from underused allocations ahead in the queue.

    Image from @allen_ai's post
  9. elvisXAI score40

    Syren Video learns your style to build AI videos from prompts

    AISyren Video, a new agentic video tool, learns preferred graphics, motion, and editing rhythm from a user's library and generates new videos from a prompt. Users refine the results through chat, and the tool is free to try in a browser or through Claude MCP, per the company's announcement. The post's author says the education sector is exploring it.

  10. GuizangXAI score22

    Guizang releases a one-click Grok bot for daily AI news videos

    AIGuizang says he turned his workflow into a Grok bot that users can install with one click. The bot runs on Grok's cloud virtual machine to collect content, write code, and render a daily morning AI news video without using a local computer.

  11. CoW SwapXAI score34

    CoW Protocol launches pay-per-quote API for bots and AI agents

    AICoW Protocol has launched x402.cow.fi, a service offering pay-per-request trading quotes for bots and AI agents with no API key or sign-up required. Each quote costs $0.001, payable in USDC on Base, Ethereum, or BNB Chain, or in $COW on Base. The service is built on x402.

    Image from @CoWSwap's post
  12. QbitAINewsAI score67

    Aether AI shows CRIS-0 robot recovering from disturbances via causal reasoning

    AIAether AI, founded by UCSD assistant professor Biwei Huang, has released official demos of its CRIS-0 causal intelligence system for robots. In tests, the robot recovered from external disturbances in 9 of 10 random trials, typically within about 2 seconds, and stopped within 0.2 seconds when a human hand entered the workspace during a microwave-door task.

  13. ModelScopeOfficialAI score28

    Corvus-Gov-3B: a 3B Chinese government-domain dialogue model

    AIModelScope released Corvus-Gov-3B, a compact model tuned for Chinese policy Q&A, public-service consultation, and internal government or enterprise assistants. It was fine-tuned on one million Chinese government-domain dialogue samples and built on Llama 3.2 3B Instruct using LoRA SFT via LLaMA Factory. The model is released under Apache 2.0.

    Image from @ModelScope2022's post
  14. OpenBMBOfficialAI score28

    MiniCPM5-2B runs at 37 tok/s on iPhone Air

    AIOpenBMB reports that its MiniCPM5-2B model runs at 37 tokens per second on an iPhone Air. The post presents this as evidence that small open multimodal models can run on mobile devices without a cloud GPU, with NobodyWho noting the model is available in its Chat app.

  15. MarkTechPostNewsAI score44

    Underdog Releases Saluki 27B, a 2-Bit Qwen3.8-27B That Beats the Original at Tool Calling

    AIUnderdog has released Saluki 27B under Apache 2.0, a 2-bit GGUF of Qwen3.8-27B that fits in 7.89 GB, versus 54 GB for the full BF16 model. On Underdog Bench, Saluki scores 88 against 84 for the full model, and it raises parallel tool-call accuracy to 42 from 35. It runs on stock llama.cpp, but math and reasoning drop sharply, with AIME 2025 at 79.2 versus 96.7.

  16. IThome · AINewsAI score55

    Odyssey-3 world model scores 66.1 on Physics-IQ Verified benchmark

    AIOdyssey announced the Odyssey-3 series of foundation world models, with Odyssey-3 Pro scoring 66.1 on the Physics-IQ Verified video-to-video benchmark, the highest recorded on that leaderboard. The series includes a standard version balancing physical accuracy and generation cost, and a Pro version with stronger physics prediction. The preview supports first-person and third-person navigation and lets users move the camera, take actions, or trigger events while the model predicts environmental changes in real time.

  17. PandailyNewsAI score38

    KingKong Technology Open-Sources Jumper Crab Robot Software Stack

    AIKingKong Technology has open-sourced the software stack for Jumper, a six-legged crab-style robot it designed, including its MuJoCo model, simulation scenes, reinforcement learning training and deployment tooling. Jumper has 22 degrees of freedom, measures about 400 by 400 by 200 mm, weighs about 1.8 kg and lists a maximum jump height of 400 mm or more. The mechanical CAD files, bill of materials, PCB designs and electrical schematics are not public, and RKNN inference on the real board has not yet been validated.

  18. vLLMOfficialAI score42

    vLLM Semantic Router team releases Decision 2.0 multi-question classification models

    AIThe vLLM Semantic Router team has released Decision 2.0, which answers multiple questions about one input in a single forward pass and outputs per-option probabilities. The post presents this as useful for routing and classification. A quoted post from Xunzhuo Liu says Decision 2.0 includes six open decision models ranging from 0.6B to 27B parameters, each topping same-size open models on the Jev Decision Index 0.3.

  19. Simon WillisonBlogAI score14

    ttok 1.0 switches default tokenizer to GPT-5 family and GPT-6

    AISimon Willison released ttok 1.0, changing its default tokenizer from GPT-4 to a newer GPT-5/GPT-6 tokenizer after finding the old default was outdated. OpenAI has not confirmed that GPT-6 shares the GPT-5 tokenizer, but William Liu's experiment found all seven GPT models tested reported 44,794 tokens on a 31-fixture corpus.

  20. QbitAINewsAI score44

    Sharpa unveils D01 humanoid robot, W02 dexterous hand, and AE01 haptic glove at IROS

    AISharpa launched D01, a fully self-developed humanoid robot with electronic skin covering the whole body and tactile coverage of the upper body, sensing forces from 0.1 to 20N at 100Hz. It also unveiled the W02 dexterous hand, which has 21 active degrees of freedom, about 30% smaller than the W01, and the AE01 exoskeleton data glove with 22 encoders for teleoperation and data collection.

Oct 8

Oct 8Thu
  1. meng shaoXAI score43

    Unsloth integrates Microsoft's mxc sandbox for Windows AI agent isolation

    AIUnsloth has integrated Microsoft's open-source mxc sandboxing system into Windows as an OS-level sandbox for isolating AI agent code execution. Its High mode provides real operating-system isolation that confines tool calls to specified directories, while its Low mode adds language-level checks that block dangerous commands and shell escapes. Both modes also strip secret environment variables and enforce resource limits such as 8GB memory and 600-second CPU time.

    Image from @shao__meng's post
  2. Higgsfield AI 🧩OfficialAI score36

    Higgsfield Katana adds community presets for Claude video editing

    AIHiggsfield has released community presets for Higgsfield Katana, its AI video editing tool available inside Claude. Users can pick a preset for motion graphics, 3D animations, product launches, fashion, car, travel, or aura-farming edits, then add their own characters, products, or clothes to recreate it in Claude. More presets are coming soon.

    Video from @higgsfield's post
  3. Jerry LiuXAI score10

    Jerry Liu announces dots, linking to dot.com

    AIJerry Liu (@jerryjliu0) announced a product called dots and linked to dot.com. The post provides no details about its features, capabilities, or pricing.

  4. OpenClaw🦞OfficialAI score34

    OpenClaw shares recent feature updates and upcoming roadmap plans

    AIOpenClaw says it has added many new features and quality-of-life improvements over the past few months. A video covers new models, multiplayer, interactive dashboards, memory and skills, meetings and voice, and easier Mac setup. It also previews plans for the coming months.

  5. Simon WillisonBlogAI score22

    ttok 0.4 Adds --list-models Command for Counting Tokens with tiktoken

    AISimon Willison has released ttok 0.4, a command-line tool for counting tokens built on OpenAI's open source tiktoken library. The update adds a --list-models command, fixes a Click warning, and updates CI, and the tool can run through uvx, for example with cat file.txt | uvx ttok.

  6. meng shaoXAI score77

    Theo open-sources tsc-rs, a Rust port of the TypeScript 7 compiler

    AITheo, creator of the T3 Stack, open-sourced tsc-rs, a line-by-line Rust port of Microsoft's Go-native TypeScript 7 compiler, type checker, and language server under MIT, pinned to typescript-go commit 673a5f17. The author reports tsc-rs is about 1.61× faster than tsc 7 and about 2.95× faster than bun check on six real-app benchmarks on an Apple M4 Pro. The port passes all 181,711 ported Go tests, and CLI output matches the Go version on 120 open-source repos except for known edge cases such as monorepo rootDir and tsc -b incremental output.

    Why it matters: The post reports a benchmarked, test-verified Rust port of the TypeScript 7 compiler, with pinned upstream and stated edge cases useful for judging its compatibility.

    Image from @shao__meng's post