Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 3

Oct 3Sat
  1. ClineOfficialAI score35

    Ling 3.1 Flash is available free in Cline until October 13

    AICline says Ling 3.1 Flash is now available in its platform and free through October 13. The 560B total parameter mixture-of-experts model activates 25B parameters and is described as on par with open-weights models Kimi K3 and DeepSeek V4 Pro.

    Image from @cline's post
  2. IndexTeam (Bilibili) · new models on Hugging FaceOfficialAI score22

    Index-Echo-S2ST-9B-FP4 released as NVFP4 quantized speech translation model

    AIIndexTeam released Index-Echo-S2ST-9B-FP4, an NVFP4 (W4A4) quantization of the Index-Echo-S2ST-9B speech-to-speech translation model, with only its text LLM backbone quantized. Perplexity rose from 3.8218 to 3.9650 (+3.75%) on a fixed corpus, while zh→en and en→zh outputs were semantically equivalent, and full FP4 speedup requires an NVIDIA Blackwell GPU.

  3. IndexTeam (Bilibili) · new models on Hugging FaceOfficialAI score27

    Index-Echo-S2ST-2B FP4 Quantized Speech-to-Speech Translation Model Released on Hugging Face

    AIIndexTeam released Index-Echo-S2ST-2B-FP4, an NVFP4 (W4A4) quantized version of the Index-Echo-S2ST-2B speech-to-speech translation model, with only the text LLM backbone quantized and the audio components kept in BF16. On a fixed corpus, perplexity rose from 5.9332 to 6.4980 (+9.52%), while zh->en and en->zh generations matched the original. Full FP4 acceleration requires an NVIDIA Blackwell GPU, and the model loads via compressed-tensors in vLLM or transformers.

  4. IndexTeam (Bilibili) · new models on Hugging FaceOfficialAI score20

    IndexTeam releases NVFP4 quantized Index-Echo-S2TT-9B speech translation model

    AIIndexTeam published an NVFP4 (W4A4) quantized version of its Index-Echo-S2TT-9B speech-to-text translation model, quantizing only the text LLM backbone while keeping the audio tower and other components in BF16. On an NVIDIA A100, perplexity rose from 3.4155 to 3.5113 (+2.81%), with zh->en and en->zh outputs semantically equivalent under greedy decoding. Full FP4 speedup requires an NVIDIA Blackwell GPU, while older GPUs get only memory reduction.

  5. IndexTeam (Bilibili) · new models on Hugging FaceOfficialAI score20

    IndexTeam releases NVFP4 quantized Index-Echo-S2TT-2B speech translation model

    AIIndexTeam has published an official NVFP4 (W4A4) quantized version of its Index-Echo-S2TT-2B speech-to-text translation model on Hugging Face. Only the text LLM backbone is quantized, while the audio tower, connector, and speech-synthesis components remain in BF16. Perplexity rises 5.80%, from 4.8772 to 5.1599, on a fixed corpus, and full FP4 speedup requires an NVIDIA Blackwell GPU.

  6. IndexTeam (Bilibili) · new models on Hugging FaceOfficialAI score29

    Index-Nailong-2B-FP4 Released as NVFP4 Quantized Translation Model

    AIIndexTeam has released Index-Nailong-2B-FP4, an official NVFP4 (W4A4) quantization of its Index-Nailong-2B multilingual translation model, which supports 150 languages. The checkpoint keeps lm_head, embeddings, and MoE router gates in BF16, and a perplexity test on a fixed corpus rose from 3.2806 to 3.4998 (+6.68%), while zh->en and en->zh outputs matched BF16 semantically. Full FP4 acceleration requires an NVIDIA Blackwell GPU; on Hopper or Ampere, vLLM provides only memory savings, so the FP8 build is recommended.

  7. IndexTeam (Bilibili) · new models on Hugging FaceOfficialAI score23

    Index-Homura-9B-FP4 released with NVFP4 quantization for translation model

    AIIndexTeam released Index-Homura-9B-FP4, an official NVFP4 (W4A4) quantization of the Index-Homura-9B translation model from the Index-Translate family. On a fixed corpus, perplexity rose from 2.5386 in BF16 to 2.6245, a 3.38% increase, and zh->en generations matched the original. Full FP4 compute acceleration requires an NVIDIA Blackwell GPU, while older GPUs get only weight-only memory savings and the FP8 build is recommended for them.

  8. Aravind SrinivasXAI score34

    Perplexity Computer adds inline interactive visualizations on request

    AIPerplexity's Computer can now generate inline visualizations when users ask it to "Visualize" a topic, producing interactive widgets and animations within the thread. The feature is best used on Standard or High effort, and an example given is an inline 3D cutaway of a jet engine.

  9. Amjad MasadXAI score42

    Amjad Masad and Alex Atallah discuss AI independence and specialized agents

    AIAmjad Masad of Replit and Alex Atallah of OpenRouter discuss why AI independence and model diversification matter for enterprises. They argue that depending on a single lab risks lock-in and that specialized agents may outperform one general superagent. The post presents the conversation as a podcast episode, the first Atallah has done since Stripe acquired OpenRouter.

  10. X.PINXAI score67

    Huawei says Ascend has overtaken Nvidia in China without giving figures

    AIHuawei chairman Eric Xu said at Huawei Connect that Ascend now leads Nvidia in China, based on Huawei's own data, but did not give a market share. Bernstein forecasts about 50% for Huawei and 8% for Nvidia this year, and Xu says mainland process nodes, not chip design, are the bottleneck. DeepSeek reportedly plans to deploy at least 160,000 Ascend 950DT chips in Inner Mongolia.

    Why it matters: The source sets Huawei's self-reported lead against Bernstein's share estimates and the process-node and capacity limits Huawei cites, which shows how the claim is constrained.

  11. DatabricksOfficialAI score27

    Databricks Genie One adds ontology, uploads, and scheduled tasks

    AIDatabricks has rolled out a set of updates to Genie One spanning context, data access, collaboration, and automation. Genie Ontology is enabled by default to provide business-aware context, and workspace instructions can apply organizational data conventions to every prompt. Users can also upload Word documents, images, CSVs, spreadsheets, and PDFs, query Unity Catalog tables with schema preview and one-click access requests, and automate recurring work with scheduled tasks that reference past runs.

    Video from @databricks's post
  12. CohereOfficialAI score22

    Cohere explains how hard training samples improved North Small Translate

    AICohere reports that after the first training step, its model could already translate over 90% of the training documents, which created a data problem. Kocmi describes how the most difficult samples were used to strengthen North Small Translate's capabilities. This post is part 3 of a six-part thread.

    Video from @cohere's post
  13. Max ZeffXAI score45

    Former OpenAI safety staffer says culture, not rules, needs fixing

    AIMax Zeff quotes former OpenAI safety team member David Robinson, who resigned this week, saying he regrets not staying to push for staffing and culture changes. The quoted passage says colleagues were too busy sprinting to consider or make major changes. The Atlantic piece argues that the fix lies in culture rather than specific rules or new laws.

  14. SemiAnalysisXAI score34

    AMD reaches above 90% parity on upstream vLLM gating tests

    AIAMD has reached above 90% parity on upstream vLLM gating test groups this week, according to SemiAnalysis. The milestone followed months of work by AMD maintainers, including Andreas, and vLLM CI lead Kevin, plus SemiAnalysis supplying additional AMD GPUs to vLLM CI.

    Image from @SemiAnalysis_'s post
  15. Guillermo RauchXAI score52

    Vercel confirms a KVM zero-day found through its sandbox bounty program

    AIVercel says it confirmed a zero-day vulnerability in KVM, the Linux virtualization standard, through its Vercel Sandbox bounty program. The author credits researcher Paulos and other researchers for helping build a more secure sandbox for agents, and says a full writeup is coming. A screenshot shows Vercel awarding a $50,000 bounty for the report, which the screenshot describes as a guest-to-host root escape.

  16. PixVerseOfficialAI score22

    PixVerse launches Short Drama plugin for AI agent video creation

    AIPixVerse has introduced Short Drama, a plugin that lets an AI agent turn a user's scene description into a finished video. The plugin takes text covering characters, setting, action, and mood, and produces video through the agent workflow, giving teams concrete material to review and develop.

    Video from @PixVerse's post
  17. Aravind SrinivasXAI score28

    Perplexity plans to run its agent sandboxes on NVIDIA Vera CPUs

    AIPerplexity says it aims to vertically integrate its agentic infrastructure by owning its sandboxes and optimizing them for the best silicon. Aravind Srinivas claims NVIDIA's Vera is far better than x86, with more details promised as Perplexity Computer begins rolling out on Vera.

Oct 2

Oct 2Fri
  1. Hamel HusainXAI score35

    Hamel Husain criticizes a Claude Code mod demo as hard to follow

    AIHamel Husain says he cannot understand a demo video for a new Claude Code modding feature, calling it visual slop. He suggests the feature may be cool but argues demos should be understandable to humans. The background post says Claude Code can now be modded to change behavior, customize the UI, or add features via TypeScript or Claude-built mods installed through /plugin.

  2. ollamaOfficialAI score29

    Cloudflare's Clef decision models now available on Ollama

    AIOllama now offers Cloudflare's decision models, Clef (27B) and Clef Flash (9B), which classify images, label bug reports, and route support tickets. Users can run them locally with the commands ollama pull clef and ollama pull clef-flash.

    Image from @ollama's post
  3. NVIDIA AIOfficialAI score33

    Nemotron 3 Diarization tracks overlapping speakers on Hugging Face

    AINVIDIA's Nemotron 3 Diarization model identifies who spoke when, including during overlapping speech, and is now available on Hugging Face. It supports up to eight speakers and has 100M parameters. The post thanks users for downloads and trending activity and shares a follow-up answering community questions.

    Video from @NVIDIAAI's post
  4. IThome · AINewsAI score36

    Analyst Dumps Airbnb, Buys Meta After Testing Meta's Muse AI Agent

    AIIndependent analyst Mostly Borrowed Ideas said he sold his Airbnb stake and added to Meta after testing Meta's Muse AI agent for about 10 days. He said Muse browsed Airbnb like a human, then found a farmhouse stay about 60% cheaper by booking directly with the host, suggesting AI agents could bypass booking platforms. He acknowledged Muse is slow, with a five-hotel price comparison taking 14 minutes.

  5. Replit ⠕OfficialAI score40

    Replit adds interactive charts, new models, and Jev integration

    AIReplit chat now generates interactive charts when users ask Replit Agent to visualize data. Users can also choose GPT-6.1 Sol from OpenAI or Claude Sonnet 5.5 from Anthropic when building with Agent, or stay in auto mode. Jev is available through Replit AI Integrations for classifying content, routing requests, and scoring leads without managing API keys.

    Video from @Replit's post
  6. Prime IntellectOfficialAI score34

    vLLM's block-major KV layout halves NVLink transfer time

    AIvLLM changed its KV cache layout to block-major BLHNC, cutting transfer descriptors about 10x and halving mean KV transfer time on NVLink. The original slowdown came from fragmented KV layout that split one 200K-token request into 32K tiny copies, making NVLink slower than InfiniBand.

    Image from @PrimeIntellect's post
  7. Prime IntellectOfficialAI score20

    Prime Intellect: DEP8 cuts prefix-cache pressure versus TEP8 on same GPUs

    AIPrime Intellect reports that DEP8 provides about 5x the prefix-cache capacity of TEP8 on the same GPUs. The post argues that fast KV retrieval alone does not ensure fast first tokens, since cached KV often sat ready while requests waited to join a batch. Halving the prefill budget reduced median queue wait time and time to first token (TTFT).

    Image from @PrimeIntellect's post
  8. Prime IntellectOfficialAI score38

    Prime Intellect stores MLA KV cache in NVFP4 for more cached tokens

    AIPrime Intellect compresses the MLA latent KV cache to NVFP4, reducing each row from 576 to 352 bytes. This fits about 50% more cached tokens per decoder compared with FP8. Its native sparse-MLA kernel unpacks the format on-chip, and the company is contributing that kernel to FlashInfer as an experimental operation.

    Image from @PrimeIntellect's post
  9. Prime IntellectOfficialAI score38

    GLM-5.3 served on GB200 NVL72 at 100+ tokens/s per user

    AIPrime Intellect served GLM-5.3 on GB200 NVL72 while targeting 100+ end-to-end tokens per second per user for concurrent agent tasks. At that interactivity bar, a 1:4 prefill-to-decode ratio delivered the most throughput, supporting 66 sessions per prefill group at 101 tokens/s per user and 100 output tokens/s per GPU.

    Image from @PrimeIntellect's post
  10. Prime IntellectOfficialAI score23

    Prime Intellect optimizes long-context agent serving across three paths

    AIPrime Intellect says long-context agent serving depends on retaining history, scheduling new work, and moving cached state efficiently. It optimized three paths separately: prefill topology and scheduling, compressed KV with a fused attention kernel, and a transfer-friendly cache layout.

  11. Prime IntellectOfficialAI score20

    Prime Intellect launches Prime Inference for serving AI model tokens

    AIPrime Intellect has introduced Prime Inference, an inference service it says has served trillions of tokens for reinforcement learning and dedicated customer deployments. The company argues that owning your intelligence requires owning your inference, and the post promises to unpack its inference stack.

    Video from @PrimeIntellect's post
  12. Guillermo RauchXAI score26

    Vercel's Jev arrives in the AI SDK for Python

    AIVercel has added Jev to the AI SDK for Python, and the team tested it in two experiments: detecting whether typed text is Python or English, and writing Python one decision at a time. The main post is a short endorsement praising a writeup about Jev and Python.