Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Oct 3

Oct 3Sat
  1. Aravind SrinivasXAI score28

    Perplexity plans to run its agent sandboxes on NVIDIA Vera CPUs

    AIPerplexity says it aims to vertically integrate its agentic infrastructure by owning its sandboxes and optimizing them for the best silicon. Aravind Srinivas claims NVIDIA's Vera is far better than x86, with more details promised as Perplexity Computer begins rolling out on Vera.

Oct 2

Oct 2Fri
  1. Hamel HusainXAI score35

    Hamel Husain criticizes a Claude Code mod demo as hard to follow

    AIHamel Husain says he cannot understand a demo video for a new Claude Code modding feature, calling it visual slop. He suggests the feature may be cool but argues demos should be understandable to humans. The background post says Claude Code can now be modded to change behavior, customize the UI, or add features via TypeScript or Claude-built mods installed through /plugin.

  2. ollamaOfficialAI score29

    Cloudflare's Clef decision models now available on Ollama

    AIOllama now offers Cloudflare's decision models, Clef (27B) and Clef Flash (9B), which classify images, label bug reports, and route support tickets. Users can run them locally with the commands ollama pull clef and ollama pull clef-flash.

    Image from @ollama's post
  3. NVIDIA AIOfficialAI score33

    Nemotron 3 Diarization tracks overlapping speakers on Hugging Face

    AINVIDIA's Nemotron 3 Diarization model identifies who spoke when, including during overlapping speech, and is now available on Hugging Face. It supports up to eight speakers and has 100M parameters. The post thanks users for downloads and trending activity and shares a follow-up answering community questions.

    Video from @NVIDIAAI's post
  4. IThome · AINewsAI score36

    Analyst Dumps Airbnb, Buys Meta After Testing Meta's Muse AI Agent

    AIIndependent analyst Mostly Borrowed Ideas said he sold his Airbnb stake and added to Meta after testing Meta's Muse AI agent for about 10 days. He said Muse browsed Airbnb like a human, then found a farmhouse stay about 60% cheaper by booking directly with the host, suggesting AI agents could bypass booking platforms. He acknowledged Muse is slow, with a five-hotel price comparison taking 14 minutes.

  5. Replit ⠕OfficialAI score40

    Replit adds interactive charts, new models, and Jev integration

    AIReplit chat now generates interactive charts when users ask Replit Agent to visualize data. Users can also choose GPT-6.1 Sol from OpenAI or Claude Sonnet 5.5 from Anthropic when building with Agent, or stay in auto mode. Jev is available through Replit AI Integrations for classifying content, routing requests, and scoring leads without managing API keys.

    Video from @Replit's post
  6. Prime IntellectOfficialAI score10

    Prime Intellect shows how to run GLM-5.3 via CLI or OpenAI SDK

    AIPrime Intellect's post gives a quick start for querying z-ai/glm-5.3 through its inference service, using the command prime inference chat with a sample haiku prompt. Developers can alternatively point any OpenAI-compatible SDK at

  7. Prime IntellectOfficialAI score34

    vLLM's block-major KV layout halves NVLink transfer time

    AIvLLM changed its KV cache layout to block-major BLHNC, cutting transfer descriptors about 10x and halving mean KV transfer time on NVLink. The original slowdown came from fragmented KV layout that split one 200K-token request into 32K tiny copies, making NVLink slower than InfiniBand.

    Image from @PrimeIntellect's post
  8. Prime IntellectOfficialAI score20

    Prime Intellect: DEP8 cuts prefix-cache pressure versus TEP8 on same GPUs

    AIPrime Intellect reports that DEP8 provides about 5x the prefix-cache capacity of TEP8 on the same GPUs. The post argues that fast KV retrieval alone does not ensure fast first tokens, since cached KV often sat ready while requests waited to join a batch. Halving the prefill budget reduced median queue wait time and time to first token (TTFT).

    Image from @PrimeIntellect's post
  9. Prime IntellectOfficialAI score38

    Prime Intellect stores MLA KV cache in NVFP4 for more cached tokens

    AIPrime Intellect compresses the MLA latent KV cache to NVFP4, reducing each row from 576 to 352 bytes. This fits about 50% more cached tokens per decoder compared with FP8. Its native sparse-MLA kernel unpacks the format on-chip, and the company is contributing that kernel to FlashInfer as an experimental operation.

    Image from @PrimeIntellect's post
  10. Prime IntellectOfficialAI score38

    GLM-5.3 served on GB200 NVL72 at 100+ tokens/s per user

    AIPrime Intellect served GLM-5.3 on GB200 NVL72 while targeting 100+ end-to-end tokens per second per user for concurrent agent tasks. At that interactivity bar, a 1:4 prefill-to-decode ratio delivered the most throughput, supporting 66 sessions per prefill group at 101 tokens/s per user and 100 output tokens/s per GPU.

    Image from @PrimeIntellect's post
  11. Prime IntellectOfficialAI score23

    Prime Intellect optimizes long-context agent serving across three paths

    AIPrime Intellect says long-context agent serving depends on retaining history, scheduling new work, and moving cached state efficiently. It optimized three paths separately: prefill topology and scheduling, compressed KV with a fused attention kernel, and a transfer-friendly cache layout.

  12. Prime IntellectOfficialAI score20

    Prime Intellect launches Prime Inference for serving AI model tokens

    AIPrime Intellect has introduced Prime Inference, an inference service it says has served trillions of tokens for reinforcement learning and dedicated customer deployments. The company argues that owning your intelligence requires owning your inference, and the post promises to unpack its inference stack.

    Video from @PrimeIntellect's post
  13. Guillermo RauchXAI score26

    Vercel's Jev arrives in the AI SDK for Python

    AIVercel has added Jev to the AI SDK for Python, and the team tested it in two experiments: detecting whether typed text is Python or English, and writing Python one decision at a time. The main post is a short endorsement praising a writeup about Jev and Python.

  14. ModalOfficialAI score6

    Modal thanks attendees of its Runtime event for joining

    AIModal thanked everyone who attended its Runtime event, highlighting nine tracks, more than 30 speakers, and a room full of engineers running AI in production. The post offers no further details about the sessions or announcements.

    Image from @modal's post
  15. MiniMax Design (H3)OfficialAI score10

    MiniMax urges users to update to the latest version

    AIMiniMax's Hailuo AI account urges users to update to the latest version through a linked design page. The post gives no version number, feature details, or release notes.

  16. Baseten BlogOfficialAI score70

    Baseten's agent-built VibeQwen engine beats vLLM on Qwen-3.6 decode speed

    AIBaseten tested the MetaInfer skills-only approach by having Claude Code build an inference engine, VibeQwen, for Qwen-3.6-35B-A3B in NVFP4 on a single B200. On single-stream text, VibeQwen decoded 90% faster than a tuned vLLM 0.25.1 deployment (1,792 vs. 943 TPS) and cut time to first token from 28 ms to 12 ms, with a 71% throughput gain at concurrency 32. The author notes this was an outcome-focused run that allowed some numerically different outputs as long as accuracy stayed at or above the BF16 baseline.

    Why it matters: The post tests a skills-only inference engine method on a real model and states the speed and accuracy constraints used, helping readers judge how far such automated optimization can be trusted.

  17. Aravind SrinivasXAI score44

    Perplexity Computer builds a 3D map of NYC restaurants

    AIPerplexity's Computer built a 3D map of nearly 26,000 restaurants and cafes across New York City's five boroughs. Users can search by dish or neighborhood and step inside places such as Peter Luger and Grand Central Oyster Bar. The post frames such projects as ones an agent can run for hours to produce something substantial.

  18. Higgsfield AI 🧩OfficialAI score10

    Higgsfield AI Influencer Studio offers five free generations

    AIHiggsfield AI is promoting its AI Influencer Studio, letting users try the tool with up to five free generations. The post links to the studio page but provides no details on features, models, or pricing.

  19. Higgsfield AI 🧩OfficialAI score9

    Higgsfield AI launches AI Influencer tool for creating virtual personas

    AIHiggsfield AI has introduced AI Influencer, a tool for creating your own AI influencer and inserting them into any trend. Users can try it with up to 5 free generations on Higgsfield and in the ChatGPT extension, and the post says it is powered by Genjutsu.

    Video from @higgsfield's post
  20. Aravind SrinivasXAI score62

    Perplexity open-sources models, an inference engine, and security tools

    AIPerplexity has released several open source projects, including the pplx-decider-v1-27b multimodal decision model, the pplx-embed-v2-context-9b-preview contextual embeddings model, and the Lily local inference engine for Apple silicon. The post also lists the 0.6B on-device PII-Tracer classifier with its PII-TRACE benchmark, the WANDR research agent benchmark, and the Numbat and Bumblebee security tools, and says more open source releases are coming soon.

    Why it matters: The post lists several named open source releases with specific benchmark figures, helping readers scan which tools and models Perplexity has recently published.

  21. ClaudeDevsOfficialAI score29

    Anthropic adds "You should Know" plugin to Claude Code

    AIAnthropic is adding a new Claude Code plugin called "You should Know" that scans Claude's output for important information users might otherwise miss. It can be enabled with the command /plugin enable cc-plugin-you-should-know@builtin.

    Video from @ClaudeDevs's post
  22. Claude Code · GitHub ReleasesOfficialAI score38

    Claude Code v2.1.288 is released with fixes and new controls

    AIAnthropic released Claude Code v2.1.288, adding $.ui.selection() for mods, a built-in gh api for cloud sessions without the GitHub CLI, and --max-findings for /code-review. The release also fixes many issues, including mid-response API timeouts, resume and compaction bugs, and auto mode denials and model switching on Bedrock and Mantle.