Ling-3.1-flash now free on OpenCode with 262K context
AIOpenCode is offering inclusionAI's latest model, Ling-3.1-flash, for free. The model has 560B total parameters, 25B active parameters, and a 262K context window.
Updated
Updated
Items with an AI score under 20 are hidden. Show low-relevance items
AIOpenCode is offering inclusionAI's latest model, Ling-3.1-flash, for free. The model has 560B total parameters, 25B active parameters, and a 262K context window.
AIRunway co-founder and co-CEO Cristóbal Valenzuela closed the Runway AI Summit by reflecting on the evolution of generative video. He also introduced Continuum, the latest project from Runway Labs. Full talks will be available on the Runway summit site after sign-up.
AIAnt Ling has made Ling-3.1-flash available on OpenRouter at no cost, inviting users to try it and share feedback. The post provides no details on model size, benchmarks, context length, or pricing beyond the free access.
AIAi2 released AstaBrief 8B, an open-weights model that turns a research question and retrieved literature excerpts into a cited report, along with its training data. The model runs as Fast mode in Asta, averaging 51.1 seconds per report versus 178.5 seconds for Thinking mode, about 3.5x faster. The post also describes filtering synthetic training data by citation density and building DPO pairs judged by two models that agreed.
Why it matters: The post explains how supervised fine-tuning, preference data, and citation-density filtering were used to build a cited-report model, which is useful for teams training their own models.
AILiquid AI shared a Hugging Face guide to multi-harness reinforcement learning for coding agents, in which a proxy records the token ids and logprobs vLLM samples so training works without changing the harness. Per the quoted post, LFM2.5-2.6B rose from 42% to 54% after training across four harnesses at once, and imitation fine-tuning on 3,189 rollouts from Qwen3.8-27B plateaued at 47.5%, below both RL runs. The proxy, trainer, tasks, SFT data, training code and seven trained models are described as open.
Why it matters: The guide explains how to train one model with RL across several coding agent harnesses without modifying the harnesses, using a proxy that records token ids and logprobs.
AIllama.cpp can now run decision models on-device, according to Clément Delangue of Hugging Face. He says the setup is free, fast, and private, and gives the command llama serve -hf ggml-org/Kev-4B-GGUF to start it.

AIGoogle announces a next-generation federated learning system that uses Trusted Execution Environments to provide verifiable, auditable data anonymization. The system publishes access policies to a public transparency log and is deployed in Gboard, which has launched English and Japanese next-word prediction models with stronger privacy guarantees and improved accuracy. Training time has also sped up significantly because computation moved to the server and is parallelized across many machines.
Why it matters: The post shows how Trusted Execution Environments make federated learning's privacy claims externally verifiable, rather than relying on trust in the server operator.
AIllama.cpp now supports decision models, which route tickets, moderate content, or choose an agent's next step by returning a probability for every option. Five open models from 144M to 27B parameters are supported at launch, and the team says more will follow in the coming days. Because most decision models do not need large GPUs, they are a good fit for llama.cpp, and a Hugging Face blog post explains how to set them up.
AIHugging Face and collaborators published a guide to multi-harness RL that trains models through a capture proxy without changing the agent harness. The proxy records the token ids and logprobs vLLM samples, and the source reports LFM2.5-2.6B rising from 42% to 54% after training across four harnesses. Fine-tuning on 3,189 successful rollouts from Qwen3.8-27B plateaued at 47.5%, below both RL runs, and the capture proxy, trainer, tasks, SFT data, training code, and seven trained models are released openly.
Why it matters: The source gives a concrete method for training models across several agent harnesses, with measured gains and a note that imitation learning underperformed RL.

AIGeorgi Gerganov announced that decision models are now available in the latest llama.cpp builds through a new /v1/systemone endpoint. The endpoint supports local, private inference with multiple open models, and more are planned.
AIinclusionAI has released Ling 3.1 Flash, a hybrid reasoning mixture-of-experts model with 560B total parameters, of which 25B are active. The source does not provide benchmark scores, pricing, context length, or availability details.
AINVIDIA's DGX Spark 64GB configuration will be available from Acer, ASUS, Dell, Gigabyte, HP and MSI on Oct. 23, starting at $4,999. It supports models up to 100 billion parameters on device, and two units can be clustered via NVIDIA Sync Cluster Assistant to pool 128GB of memory and support up to 200 billion parameters. NVIDIA says the clustered setup delivers up to 1.7x the performance of a single system in its Qwen 3.8 27B test.
AIAi2 released AstaBrief 8B, a model that turns a research question and retrieved literature excerpts into a cited report, along with its training data. In Asta's Generate a report feature, Fast mode averages 51.1 seconds per report versus 178.5 seconds for Thinking mode, about 3.5x faster. The model is built on Qwen3-8B with supervised fine-tuning and DPO, and institutions can run its open weights on their own infrastructure.
Why it matters: The post explains the data filtering and one-pass generation choices behind a fast open-weights report model, showing what worked and what did not.
AIDeepSeek released packaged desktop versions of DeepSeek Harness for macOS and Windows. Linux users can install it from the @deepseek-ai/dsh package on npm.
AIPrime Inference is a serving platform for frontier open-source models, offering serverless endpoints and reserved capacity on Prime's GPU infrastructure across multiple datacenters. Its first public deployment, GLM-5.3, went live on OpenRouter on September 22, and the post reports a near-zero tool-call error rate and 100% uptime since launch. The post also describes GLM-5.3 serving on GB200 NVL72 with prefill/decode disaggregation and NVFP4 KV compression.
Why it matters: The post separates scheduler, KV-cache, and tool-call fixes, showing concretely which bottlenecks shape production serving of open frontier models.
AIResearchers Maureen de Seyssel, Jie Chi, and Zakaria Aldeneh found that strengthening language discrimination during pretraining reduces the performance gap between multilingual and monolingual HuBERT speech models. In a controlled English/French setting, phone-ABX error fell from 11.6% to 10.4%, close to the monolingual 10.8%, while lexical sWUGGY scores rose from 52.1% to 56.7%. The gains were largest when language discrimination was introduced in the first training iteration.
AITeknium, back from vacation, says the Hermes update should be around 4x faster for everyone. The post gives no further details on what changed or how the speedup was measured.

AIcheap typed answers for the small calls inside a harness (routing, approvals, judging), big model for the rest
AIGLM 5.3 and GLM 5.3 Flash are now available in Cursor. GLM 5.3 Max is the best-scoring open-weight model on CursorBench 4.0.

AITogether AI highlighted a talk by Yogish Baliga arguing that open-weight models have closed the gap with closed models for most use cases. The post cites Vercel's AI Gateway data showing agentic workloads on open models rose from 30% to 78% in 13 weeks.

AIComfyUI released YUI, described as the first animated short film made with Comfy Agent, a 4:20 film its creator made in three days. Comfy Agent handled shot regeneration, side-by-side model comparisons, and character consistency checks that previously required weeks of manual testing and generation.
AINVIDIA has released PixelUMM, an encoder-free unified multimodal model with 15,199,672,064 parameters that handles text, image, and video understanding and generation directly in pixel space. It represents images as 16-by-16 RGB pixel patches on a Qwen3-8B language backbone, with iterative denoising for generation. The checkpoint is licensed for non-commercial research or evaluation only, while the source code is under Apache License 2.0.
AIPerplexity open-sourced pplx-decider-27b, a state-of-the-art multimodal Decision Model, and offers it through a new Decisions API. Input tokens cost $0.04 per million and output tokens are free, with further price cuts promised in the coming days.
AIARC Prize has published full ARC-AGI results for Alibaba's Qwen3.8-27B, with the testing run hosted by Baseten. The post links the public leaderboard, the open-source benchmarking repository for reproducing the results, and the testing policy.
AIDSPy's official account highlighted a recorded office hours session showcasing Jev use cases and explainers. The quoted context says the session covered DSPy's Jev/System One implementation, the ReAnchor optimizer, and design patterns for selective compaction, tool approvals, and subagent delegation.
AIBlack Forest Labs announces FLUX 3 Image, which supports multi-turn edits that leave other pixels unchanged, layout control via bounding boxes, generation up to 4K, and up to 10 reference images. Commercial weights are available for companies running image generation at scale, and an open weights version is launching in the coming weeks.
AIPerplexity has published the weights for pplx-decider-v1-27b, a model fine-tuned from Qwen3.8-27B with a 250k-token context window. The weights are available on Hugging Face, and the post points developers to the Perplexity Decisions API quickstart for getting started.
AIGuillermo Rauch argues that the future is verification engineering, spanning proofs, end-to-end tests, benchmarks, and linters. He expects some of these tests to be deterministic and others agentic, and he says the approach looks great. The quoted post introduces e2e, an open-source agentic testing framework that mixes deterministic and agentic APIs and runs locally or in CI.
AIAnthropic is adding pre-built sample mods to the Claude Code Playground repo for developers to adapt. The post points readers to a getting started guide on creating Claude Code mods.
AICloudflare has open-sourced Clef and Clef-Flash, its first homegrown decision models, now available on Workers AI. Harrison Chase praised the release as an open-weights entrant and argued that agent harnesses should be able to swap their decision model as easily as their main model.
AIGitHub Security Lab has released gh-secure, a tool that gives repositories basic security hygiene with a single command to help stop secrets from leaking into public code. It can be installed and used in the Copilot app, Copilot CLI, or GitHub CLI.

AIGoodfire says AI cybersecurity risks are already here and that biosecurity risks are next, as models improve at biology. The post presents this as both an opportunity for science and medicine and a reason for stronger security. It quotes Demis Hassabis announcing SynthID for biology, a watermarking approach for AI-generated proteins, published in Nature with SynthID Bio tools open sourced.
AICloudflare's Workers AI team has released Clef, its first models, two fast decision models that it says top benchmarks for quality and latency. They are available hosted on Workers AI or as open weights on Hugging Face.
AIComfyUI announces the Comfy Developer Platform, which includes the Comfy API and Comfy Router for building applications. The post is a broadcast link without further details on features, pricing, or availability.
AIGoogle has introduced the @google/stitch CLI, letting users generate screens and design systems without leaving the terminal. It connects to local coding agents and can send a local dev server snapshot to Stitch. The tool complements the existing Stitch MCP and SDK, and can also be driven through agents such as Antigravity.
AIHugging Face shows that training LFM2.5-2.6B with RL inside the agent harnesses themselves lifted held-out task success from 42% to 54% across four harnesses. Before training, the model solved 62% of tasks in Mini-SWE-Agent but only 33% in Claude Code, so the same model behaved very differently per harness. The approach uses an OpenEnv capture proxy to record tokens and logprobs, Harbor for tasks and sandboxes, and TRL's async GRPO trainer, with 31% fewer tool calls on already-solved tasks; training in OpenCode alone mostly improved OpenCode.
AIHugging Face says ml-intern is an open-source ML engineering and research harness usable free on local setups, and it is also hosted on Hugging Chat with no-code access. A second hosted option runs on Hugging Face infrastructure, where ml-intern selects the cheapest GPU for a task so models can be trained for a few dollars. MaziyarPanahi reportedly trained a model by prompting alone for $6.60 on an NVIDIA A100 in 16 minutes.
AICloudflare released Clef and Clef-flash, two decision models hosted on Workers AI and open-sourced on Hugging Face under Apache 2.0, and launched a reinforcement learning fine-tuning service. In Cloudflare's tests, Clef classified a domain in 2.2s versus 4.7s for gpt-oss-120b, and the models are Jev-API compatible. The company is offering fine-tuning first through a forward-deployed engineering team, with a self-serve platform planned later.
AIHugging Face's Merve Noyan promoted a four-chapter YouTube series teaching viewers how to train AI agents using only open-source tools. The post links to the series' first video and is a brief call to learn agentic training.
AISara Hooker shared a full technical report on a method for creating datasets from zero seed data, linked via alphaXiv. She also pointed readers to an interactive tool at adaptionlabs.ai where users can begin inventing datasets.