Ollama adds Clef Flash and Clef models to its library
AIOllama has published Clef Flash and Clef as new models in its model library, with links to their library pages. The post provides no further details on parameters, benchmarks, or capabilities.
Updated
Updated
Showing low-relevance items too. Hide low-relevance items
AIOllama has published Clef Flash and Clef as new models in its model library, with links to their library pages. The post provides no further details on parameters, benchmarks, or capabilities.
AICMU researchers released SMDD-Bench, a benchmark of 502 small-molecule drug design tasks that use RDKit, ADMET-AI, and Boltz-2 as feedback loops. The authors argue that long-horizon planning, exploration, and learning from imperfect feedback remain open problems beyond math and coding, and the benchmark is available in Prime Intellect's Environments Hub for training with prime-rl.
AIOpenClaw has integrated Tencent Hunyuan's AI-Infra-Guard (AIG) into ClawScan, its open-source command-line security tool. Every skill and plugin uploaded to ClawHub now runs through AIG as part of its security review.
AIInferact invites developers to a free vLLM Meetup in Toronto, co-hosted with Cohere, where co-founder and lead maintainer Roger Wang will discuss where the vLLM project is heading next. Engineers from Cohere and NVIDIA are also scheduled to speak, and spots are limited, with registration through a link in the thread.
AISGLang's community recapped its sessions at AI Infra Summit in Santa Clara, where core contributor Alex Nails and Samsung's Vasanthi Jagatha presented SGLang HiCache with Samsung Cognos for the KV cache bottleneck. LinkedIn's Sundara Raman Ramachandran also described running latency-critical ranking on SGLang at global scale. The post promotes upcoming SGLang events.

AIPrime Intellect reports that DEP8 provides about 5x the prefix-cache capacity of TEP8 on the same GPUs. The post argues that fast KV retrieval alone does not ensure fast first tokens, since cached KV often sat ready while requests waited to join a batch. Halving the prefill budget reduced median queue wait time and time to first token (TTFT).

AIPrime Intellect compresses the MLA latent KV cache to NVFP4, reducing each row from 576 to 352 bytes. This fits about 50% more cached tokens per decoder compared with FP8. Its native sparse-MLA kernel unpacks the format on-chip, and the company is contributing that kernel to FlashInfer as an experimental operation.

AIPrime Intellect served GLM-5.3 on GB200 NVL72 while targeting 100+ end-to-end tokens per second per user for concurrent agent tasks. At that interactivity bar, a 1:4 prefill-to-decode ratio delivered the most throughput, supporting 66 sessions per prefill group at 101 tokens/s per user and 100 output tokens/s per GPU.

AIVercel has added Jev to the AI SDK for Python. The team tested it in two experiments: detecting whether typed text is Python or English as it is entered, and writing Python one decision at a time.
AIAt an event, @zainhas and @parthsareen of Ollama discussed evaluating open models against production traffic rather than relying only on benchmarks. The post offers no further details on methods, models, or results.

AITogether AI's Director of Field Engineering, Rochelle Mattern, delivered a keynote on shortlisting and evaluating open models. The post is a brief follow-up noting the speaker lineup, with no further details on specific models, benchmarks, or results.

AIThe vLLM team integrated Helion, a PyTorch-native kernel DSL, into vLLM's linear backend, using per-shape autotuning to select among Standard GEMM, Split-K, and Swap-AB variants. On NVIDIA Hopper GPUs, the Helion backend outperformed the default CUTLASS and DeepGEMM backends across the evaluated models, with more than 10% throughput gains for some workloads. The work focuses on FP8 and INT8 quantized GEMM.
AISGLang has released v0.5.20, bringing Intel XPU into standard releases alongside RL sampling masks that make rollouts more reliable with up to 52% faster decode. The update also adds Unified Radix Tree SWA branching-point caching, which the project says lifts cache hit rate about 20 points and cuts TTFT by roughly one-third, plus up to 12.5× faster ROCm model loading. New models named in the release include GLM-5.3-Flash, Qwen3.8-Flash-Next, K2 Horizon, Hy4-Preview, FastH3, and VDN-H3.
AISGLang's update adds a /v1/score endpoint that returns scores for requested labels such as Yes/No or A/B/C, avoiding the label loss of generate with top-k logprobs. Its multi-item scoring computes shared context once and keeps each candidate isolated, with 16-candidate p95 on Qwen3-8B dropping from 54.1 ms (Generate) to 20.6 ms.
AISGLang demonstrated Qwen3.8-27B as a multimodal decision model that beat Pokémon FireRed's Elite Four and champion with sub-100 ms decisions from live game state. The company says its native /v1/decisions API lets LLMs and VLMs be used for classification and scoring. It also announced /v1/systemone for running Jev-like open models with the TypeSafe SDK.
AISGLang has published full release notes for version v0.5.21 on GitHub. The post itself gives no further details beyond linking to the release page.
AISGLang has released v0.5.21 with a native Decisions API that turns an LLM or VLM into a low-latency classifier and scorer. The release also lets /v1/score rerank search or RAG results in one call, lets PD instances switch between prefill and decode without restarting, and adds support for models including DeepSeek-V4.1 Flash, Kimi K3, and GLM-5.3-Flash on AMD MI355X. The announcement reports a 22% faster first token on long prompts for DeepSeek-V4.1 Flash and 20.6% higher prefill throughput for Kimi K3 in PD serving.

AICline says its desktop app works with ClinePass and free models including DeepSeek-V4.1-Flash and Space Bunny Alpha. Users start by opening Customize, choosing Connectors, and signing in with Cline.
AIChatGPT's Finances feature is rolling out to Free and Go users in the U.S. Users can securely connect their accounts through Plaid and Experian to get answers based on their own financial information.
AIKeras is moving to a pluggable backend design, with MLX and PaddlePaddle backends upcoming as add-on libraries. The team is reducing the operations needed to ship new backends and streamlining unit testing so a single harness can test all ops, such as casting consistency. KerasHub also gains many new models and is shifting its preprocessing from tf-text to PyGrain.
AINVIDIA reports that fine-tuning Nemotron 3.5 ASR reduced word error rate on Najdi and Hijazi Saudi Arabic from 55% to 30%. The post says a new tutorial shows how to adapt the model for other dialects and languages.
AINathan Lambert and Tom Zick have unveiled Trillium Labs, a new non-profit focused on the open science of frontier AI. The lab plans to build open post-training recipes and expand into open infrastructure to study topics such as RSI, reward hacking, and multi-agent systems. It is hiring, fundraising, and seeking compute, with support from Halcyon Futures and Schmidt Sciences.

AIRunway co-founder and co-CEO Cristóbal Valenzuela closed the Runway AI Summit by reflecting on the evolution of generative video. He also introduced Continuum, the latest project from Runway Labs. Full talks will be available on the Runway summit site after sign-up.
AIllama.cpp can now run decision models on-device, according to Clément Delangue of Hugging Face. He says the setup is free, fast, and private, and gives the command llama serve -hf ggml-org/Kev-4B-GGUF to start it.

AIllama.cpp now supports decision models, which route tickets, moderate content, or choose an agent's next step by returning a probability for every option. Five open models from 144M to 27B parameters are supported at launch, and the team says more will follow in the coming days. Because most decision models do not need large GPUs, they are a good fit for llama.cpp, and a Hugging Face blog post explains how to set them up.
AIHugging Face and collaborators published a guide to multi-harness RL that trains models through a capture proxy without changing the agent harness. The proxy records the token ids and logprobs vLLM samples, and the source reports LFM2.5-2.6B rising from 42% to 54% after training across four harnesses. Fine-tuning on 3,189 successful rollouts from Qwen3.8-27B plateaued at 47.5%, below both RL runs, and the capture proxy, trainer, tasks, SFT data, training code, and seven trained models are released openly.
Why it matters: The source gives a concrete method for training models across several agent harnesses, with measured gains and a note that imitation learning underperformed RL.

AINVIDIA's DGX Spark 64GB configuration will be available from Acer, ASUS, Dell, Gigabyte, HP and MSI on Oct. 23, starting at $4,999. It supports models up to 100 billion parameters on device, and two units can be clustered via NVIDIA Sync Cluster Assistant to pool 128GB of memory and support up to 200 billion parameters. NVIDIA says the clustered setup delivers up to 1.7x the performance of a single system in its Qwen 3.8 27B test.
AIPrime Inference is a serving platform for frontier open-source models, offering serverless endpoints and reserved capacity on Prime's GPU infrastructure across multiple datacenters. Its first public deployment, GLM-5.3, went live on OpenRouter on September 22, and the post reports a near-zero tool-call error rate and 100% uptime since launch. The post also describes GLM-5.3 serving on GB200 NVL72 with prefill/decode disaggregation and NVFP4 KV compression.
Why it matters: The post separates scheduler, KV-cache, and tool-call fixes, showing concretely which bottlenecks shape production serving of open frontier models.
AIIdeogram praised photographers for their craft and said it hopes to help them save hours of work. The post contains no concrete product features, model names, or figures, and it does not specify what kind of editing help is offered.
AIResearchers Maureen de Seyssel, Jie Chi, and Zakaria Aldeneh found that strengthening language discrimination during pretraining reduces the performance gap between multilingual and monolingual HuBERT speech models. In a controlled English/French setting, phone-ABX error fell from 11.6% to 10.4%, close to the monolingual 10.8%, while lexical sWUGGY scores rose from 52.1% to 56.7%. The gains were largest when language discrimination was introduced in the first training iteration.
AIGLM 5.3 and GLM 5.3 Flash are now available in Cursor. GLM 5.3 Max is the best-scoring open-weight model on CursorBench 4.0.

AIMeta's Jagged Flash Attention kernel, built with TLX on NVIDIA Blackwell B200, outperforms FlashAttention-4 (May 2026 version) on GEM's jagged shapes by about 13% on the forward pass and about 50% on the backward pass. The TLX attention kernel is roughly 3.2K lines of Triton-level code, about 3× shorter than FA4's ~10K-line CuteDSL kernels. The benchmarks use bfloat16 on B200.
AITogether AI highlighted a talk by Yogish Baliga arguing that open-weight models have closed the gap with closed models for most use cases. The post cites Vercel's AI Gateway data showing agentic workloads on open models rose from 30% to 78% in 13 weeks.

AINVIDIA has released PixelUMM, an encoder-free unified multimodal model with 15,199,672,064 parameters that handles text, image, and video understanding and generation directly in pixel space. It represents images as 16-by-16 RGB pixel patches on a Qwen3-8B language backbone, with iterative denoising for generation. The checkpoint is licensed for non-commercial research or evaluation only, while the source code is under Apache License 2.0.
AIDSPy's official account highlighted a recorded office hours session showcasing Jev use cases and explainers. The quoted context says the session covered DSPy's Jev/System One implementation, the ReAnchor optimizer, and design patterns for selective compaction, tool approvals, and subagent delegation.
AIBlack Forest Labs says customers can contact the company for commercial weights for its FLUX 3 image model. The post also points users to a playground to try the model and to a product page with more information.
AIBlack Forest Labs announces FLUX 3 Image, which supports multi-turn edits that leave other pixels unchanged, layout control via bounding boxes, generation up to 4K, and up to 10 reference images. Commercial weights are available for companies running image generation at scale, and an open weights version is launching in the coming weeks.
AIPerplexity has published the weights for pplx-decider-v1-27b, a model fine-tuned from Qwen3.8-27B with a 250k-token context window. The weights are available on Hugging Face, and the post points developers to the Perplexity Decisions API quickstart for getting started.
AIComfyUI announces the Comfy Developer Platform, which includes the Comfy API and Comfy Router for building applications. The post is a broadcast link without further details on features, pricing, or availability.
AIHugging Face shows that training LFM2.5-2.6B with RL inside the agent harnesses themselves lifted held-out task success from 42% to 54% across four harnesses. Before training, the model solved 62% of tasks in Mini-SWE-Agent but only 33% in Claude Code, so the same model behaved very differently per harness. The approach uses an OpenEnv capture proxy to record tokens and logprobs, Harbor for tasks and sandboxes, and TRL's async GRPO trainer, with 31% fewer tool calls on already-solved tasks; training in OpenCode alone mostly improved OpenCode.