Cloudflare's Clef model trends No. 1 on Hugging Face
AICloudflare's Clef is trending at No. 1 on Hugging Face, according to a post from Hugging Face CEO Clément Delangue. The post does not provide details about what Clef is or what it does.
Updated
Updated
AICloudflare's Clef is trending at No. 1 on Hugging Face, according to a post from Hugging Face CEO Clément Delangue. The post does not provide details about what Clef is or what it does.
AIFractalyze optimized Qwen3-Omni on vLLM-Omni for a single RTX 5090, using AWQ-4bit at batch size 1 with text prompts. In its tests, time to first audio dropped from 213ms to 23ms compared with stock vLLM-Omni. vLLM hopes the optimizations will be contributed upstream to benefit more users.
AIThe PyTorch Accelerator Integration Working Group released updates on its H1 2026 progress toward standardizing how new hardware connects to the framework. Key workstreams include the Cross-Repository CI Relay (CRCR), which automatically reports downstream backend test results to a shared dashboard, and refactored test suites that decouple PyTorch's 600,000-plus tests from specific accelerators.
AILiquid AI has released d1-3B, a 3B parameter multimodal model post-trained to return calibrated, typed answers to yes/no, choice, and score questions in one forward pass. The source reports a Decision Index 0.2.1 score of 48.57, the highest among models under 10B in its table, and 8 ms per decision on an NVIDIA RTX 4090.
Why it matters: The source gives benchmark scores against named peer models and edge latency figures across several hardware targets, helping readers judge fit for on-device decision pipelines.
AIAfter NVIDIA acquired SchedMD, the SLURM scheduler's support for non-NVIDIA chips has allegedly worsened, and AMD built a competing scheduler called spur. The author says NVIDIA has not kept SLURM hardware neutral despite its earlier pledge, and questions whether Hugging Face will face the same fate after NVIDIA's acquisition of it.
AIGuillermo Rauch says his new project has a hand-written README for human readers, while its internal documentation is written in AI-style English for agents. He argues that blogs, tweets, and READMEs are human communication and should be written by people to connect with other readers.
AITeknium announced that an ESP32 is now running at home, with no further technical details given in the post. The post is a brief update that quotes an @adolandev post about Hermes Gadget, an open SDK for a small device that speaks to a user's own Hermes model and can be tested with a desktop simulator.
AIMicrosoft and Hugging Face released ThinkingBox, a benchmark that grades AI agents on the terminal backend state and side effects they leave behind rather than their final responses. Each of 507 stateful business tasks runs 20 times from a clean backend, and the post reports pass@1, pass@20, and observed 20/20 counts, plus cost per successful and per dependable task across 18 models. The harness and dataset are available on Hugging Face, with the OpenEnv interface for running evaluations.
Why it matters: The post shows why checking the database state, not tool calls or final replies, exposes agent failures, and gives a repeat-run method for judging reliability.
AIStepFun's Step 5 Preview has debuted on the Vals leaderboard, ranking #7 among open-weight models just ahead of Qwen 3.8 Max. The model costs $2.54 per task.
AINous Research highlighted that Linus Sebastian discussed his Hermes setup experience on the latest Linus Tech Tips podcast. The post celebrated that he pronounced the name Hermes correctly.
AILangChain's ModelRouterMiddleware uses Jev to read the first message and select a model that handles the entire run. Because Jev is cheap, developers can also re-select a model after each tool result using a custom hook. The router was demonstrated in a quick project by @dbreunig built with DSPy and Jev.
AIIndexTeam has released Index-Nailong-2B-FP4, an official NVFP4 (W4A4) quantization of its Index-Nailong-2B multilingual translation model, which supports 150 languages. The checkpoint keeps lm_head, embeddings, and MoE router gates in BF16, and a perplexity test on a fixed corpus rose from 3.2806 to 3.4998 (+6.68%), while zh->en and en->zh outputs matched BF16 semantically. Full FP4 acceleration requires an NVIDIA Blackwell GPU; on Hopper or Ampere, vLLM provides only memory savings, so the FP8 build is recommended.
AIIndexTeam released Index-Homura-9B-FP4, an official NVFP4 (W4A4) quantization of the Index-Homura-9B translation model from the Index-Translate family. On a fixed corpus, perplexity rose from 2.5386 in BF16 to 2.6245, a 3.38% increase, and zh->en generations matched the original. Full FP4 compute acceleration requires an NVIDIA Blackwell GPU, while older GPUs get only weight-only memory savings and the FP8 build is recommended for them.
AIAmjad Masad of Replit and Alex Atallah of OpenRouter discuss why AI independence and model diversification matter for enterprises. They argue that depending on a single lab risks lock-in and that specialized agents may outperform one general superagent. The post presents the conversation as a podcast episode, the first Atallah has done since Stripe acquired OpenRouter.
AINathan Lambert says it is a huge failure of the AI ecosystem that people respond to some unspecified situation with accusations of evil and shame. He says he will work to shift those norms, arguing that positive paths forward exist even though the path will be hard.
AIAMD has reached above 90% parity on upstream vLLM gating test groups this week, according to SemiAnalysis. The milestone followed months of work by AMD maintainers, including Andreas, and vLLM CI lead Kevin, plus SemiAnalysis supplying additional AMD GPUs to vLLM CI.
AIAMD made 502 commits to core vLLM over 90 days, accounting for 13% of all organizational contributions and almost four times Nvidia's count. The company credited Red Hat, IBM, Embedded LLM, and Inferact for building the project alongside it.
AIThe creator Khazix open-sourced AIHOT, a monthly-active-million AI hot-news site, including its collection workflow, curation scoring, clustering mechanism, and production prompts. The main post reports the project gained 5.1K GitHub stars within a few days and thanks supporters on X.
AIThe author reports that a dual RTX 5070 Ti setup running Qwen Flash rose from 200 prefill and 10 decode to 2200 prefill and 67 decode, now on a single card, using Strata and a custom PR. The post argues that such consumer-hardware speeds, once limited to top-end machines, could pressure the economics of selling model compute via API.
AIvLLM Semantic Router has released Decision 2.0, featuring open models ranging from 0.6B to 27B parameters. The models can answer 64 questions about a request in a single forward pass, are licensed under Apache-2.0, and load with Transformers.
AIOllama now offers Cloudflare's decision models, Clef (27B) and Clef Flash (9B), which classify images, label bug reports, and route support tickets. Users can run them locally with the commands ollama pull clef and ollama pull clef-flash.
AIOllama has published Clef Flash and Clef as new models in its model library, with links to their library pages. The post provides no further details on parameters, benchmarks, or capabilities.
AICMU researchers released SMDD-Bench, a benchmark of 502 small-molecule drug design tasks that use RDKit, ADMET-AI, and Boltz-2 as feedback loops. The authors argue that long-horizon planning, exploration, and learning from imperfect feedback remain open problems beyond math and coding, and the benchmark is available in Prime Intellect's Environments Hub for training with prime-rl.
AIOpenClaw has integrated Tencent Hunyuan's AI-Infra-Guard (AIG) into ClawScan, its open-source command-line security tool. Every skill and plugin uploaded to ClawHub now runs through AIG as part of its security review.
AIPrime Intellect reports that DEP8 provides about 5x the prefix-cache capacity of TEP8 on the same GPUs. The post argues that fast KV retrieval alone does not ensure fast first tokens, since cached KV often sat ready while requests waited to join a batch. Halving the prefill budget reduced median queue wait time and time to first token (TTFT).
AIPrime Intellect compresses the MLA latent KV cache to NVFP4, reducing each row from 576 to 352 bytes. This fits about 50% more cached tokens per decoder compared with FP8. Its native sparse-MLA kernel unpacks the format on-chip, and the company is contributing that kernel to FlashInfer as an experimental operation.
AIPrime Intellect served GLM-5.3 on GB200 NVL72 while targeting 100+ end-to-end tokens per second per user for concurrent agent tasks. At that interactivity bar, a 1:4 prefill-to-decode ratio delivered the most throughput, supporting 66 sessions per prefill group at 101 tokens/s per user and 100 output tokens/s per GPU.
AIVercel has added Jev to the AI SDK for Python. The team tested it in two experiments: detecting whether typed text is Python or English as it is entered, and writing Python one decision at a time.
AIAt an event, @zainhas and @parthsareen of Ollama discussed evaluating open models against production traffic rather than relying only on benchmarks. The post offers no further details on methods, models, or results.
AITogether AI's Director of Field Engineering, Rochelle Mattern, delivered a keynote on shortlisting and evaluating open models. The post is a brief follow-up noting the speaker lineup, with no further details on specific models, benchmarks, or results.
AIThe vLLM team integrated Helion, a PyTorch-native kernel DSL, into vLLM's linear backend, using per-shape autotuning to select among Standard GEMM, Split-K, and Swap-AB variants. On NVIDIA Hopper GPUs, the Helion backend outperformed the default CUTLASS and DeepGEMM backends across the evaluated models, with more than 10% throughput gains for some workloads. The work focuses on FP8 and INT8 quantized GEMM.
AIChatGPT's Finances feature is rolling out to Free and Go users in the U.S. Users can securely connect their accounts through Plaid and Experian to get answers based on their own financial information.
AIKeras is moving to a pluggable backend design, with MLX and PaddlePaddle backends upcoming as add-on libraries. The team is reducing the operations needed to ship new backends and streamlining unit testing so a single harness can test all ops, such as casting consistency. KerasHub also gains many new models and is shifting its preprocessing from tf-text to PyGrain.
AINVIDIA reports that fine-tuning Nemotron 3.5 ASR reduced word error rate on Najdi and Hijazi Saudi Arabic from 55% to 30%. The post says a new tutorial shows how to adapt the model for other dialects and languages.
AINathan Lambert and Tom Zick have unveiled Trillium Labs, a new non-profit focused on the open science of frontier AI. The lab plans to build open post-training recipes and expand into open infrastructure to study topics such as RSI, reward hacking, and multi-agent systems. It is hiring, fundraising, and seeking compute, with support from Halcyon Futures and Schmidt Sciences.
AIRunway co-founder and co-CEO Cristóbal Valenzuela closed the Runway AI Summit by reflecting on the evolution of generative video. He also introduced Continuum, the latest project from Runway Labs. Full talks will be available on the Runway summit site after sign-up.
AIllama.cpp can now run decision models on-device, according to Clément Delangue of Hugging Face. He says the setup is free, fast, and private, and gives the command llama serve -hf ggml-org/Kev-4B-GGUF to start it.
AIllama.cpp now supports decision models, which route tickets, moderate content, or choose an agent's next step by returning a probability for every option. Five open models from 144M to 27B parameters are supported at launch, and the team says more will follow in the coming days. Because most decision models do not need large GPUs, they are a good fit for llama.cpp, and a Hugging Face blog post explains how to set them up.
AIHugging Face and collaborators published a guide to multi-harness RL that trains models through a capture proxy without changing the agent harness. The proxy records the token ids and logprobs vLLM samples, and the source reports LFM2.5-2.6B rising from 42% to 54% after training across four harnesses. Fine-tuning on 3,189 successful rollouts from Qwen3.8-27B plateaued at 47.5%, below both RL runs, and the capture proxy, trainer, tasks, SFT data, training code, and seven trained models are released openly.
Why it matters: The source gives a concrete method for training models across several agent harnesses, with measured gains and a note that imitation learning underperformed RL.
AINVIDIA's DGX Spark 64GB configuration will be available from Acer, ASUS, Dell, Gigabyte, HP and MSI on Oct. 23, starting at $4,999. It supports models up to 100 billion parameters on device, and two units can be clustered via NVIDIA Sync Cluster Assistant to pool 128GB of memory and support up to 200 billion parameters. NVIDIA says the clustered setup delivers up to 1.7x the performance of a single system in its Qwen 3.8 27B test.