Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 1

Oct 1Thu
  1. ZyphraOfficialAI score38

    Zyphra Research explains how local memory aids positional sense in LLMs

    AIZyphra Research explains how language models track word order without explicitly encoding position in attention. The post says local memory layers that read nearby words help global attention layers preserve sequence information. The source is a short teaser thread, so no further technical details are given.

    Image from @ZyphraAI's post
  2. Josh WoodwardXAI score34

    Google launches Stitch CLI to generate design ideas from terminal

    AIGoogle has introduced the @google/stitch CLI, letting users generate screens and design systems without leaving the terminal. It connects to local coding agents and can send a local dev server snapshot to Stitch. The tool complements the existing Stitch MCP and SDK, and can also be driven through agents such as Antigravity.

  3. Lewis Tunstall @ COLM 🌉XAI score44

    Training LFM2.5-2.6B inside four agent harnesses boosts held-out tasks

    AIHugging Face shows that training LFM2.5-2.6B with RL inside the agent harnesses themselves lifted held-out task success from 42% to 54% across four harnesses. Before training, the model solved 62% of tasks in Mini-SWE-Agent but only 33% in Claude Code, so the same model behaved very differently per harness. The approach uses an OpenEnv capture proxy to record tokens and logprobs, Harbor for tasks and sandboxes, and TRL's async GRPO trainer, with 31% fewer tool calls on already-solved tasks; training in OpenCode alone mostly improved OpenCode.

    Video from @_lewtun's post
  4. merveXAI score46

    Hugging Face clarifies ml-intern options, one trained model for $6

    AIHugging Face says ml-intern is an open-source ML engineering and research harness usable free on local setups, and it is also hosted on Hugging Chat with no-code access. A second hosted option runs on Hugging Face infrastructure, where ml-intern selects the cheapest GPU for a task so models can be trained for a few dollars. MaziyarPanahi reportedly trained a model by prompting alone for $6.60 on an NVIDIA A100 in 16 minutes.

  5. Google · Gemini appOfficialAI score60

    Google launches Guided Vision in Gemini Live for blind and low-vision users

    AIGoogle is launching Guided Vision in Gemini Live on compatible Android devices, letting users share their camera for spoken descriptions and follow-up questions. The model was trained with Aira on tens of thousands of hours of visual interpretation and tested by more than 1,000 members of Aira's Trusted Tester network. The feature is not a medical device, mobility aid, or navigation tool, and it requires Android 9 or later.

    Why it matters: The launch shows how a real-time visual model was trained and tested with blind and low-vision users, a practical reference for accessibility-focused AI design.

  6. RunwayOfficialAI score34

    Runway's real-time video model generates interactive portals for any concept

    AIRunway says its real-time video model generates animated, interactive overlays called portals that illustrate any concept, place, or simulation while users browse. The post describes the feature as a demonstration without listing pricing, availability, or technical specifications.

    Video from @runwayml's post
  7. RunwayOfficialAI score42

    Runway's Project Continuum previews real-time video computer interfaces

    AIRunway Labs introduced Project Continuum, an operating system research application built around real-time video interfaces. The early look shows four interaction concepts: Portals, Visual Thinking, Responsive Video Interfaces, and Interactive Worlds. The post says Interface World Models such as Solaris are reimagining what an interface can be.

    Video from @runwayml's post
  8. Hugging FaceOfficialAI score23

    Hugging Face Chat adds MCP support to bring your data in

    AIHugging Face announced that users can now use MCPs (Model Context Protocol servers) to bring their own data into Hugging Face Chat's ml-intern mode. The post links to the feature at and gives no further details on setup or supported integrations.

    Video from @huggingface's post
  9. Cloudflare Blog · AIOfficialAI score58

    Cloudflare releases open-source Clef decision models and an RL fine-tuning service

    AICloudflare released Clef and Clef-flash, two decision models hosted on Workers AI and open-sourced on Hugging Face under Apache 2.0, and launched a reinforcement learning fine-tuning service. In Cloudflare's tests, Clef classified a domain in 2.2s versus 4.7s for gpt-oss-120b, and the models are Jev-API compatible. The company is offering fine-tuning first through a forward-deployed engineering team, with a self-serve platform planned later.

  10. Meta NewsroomOfficialAI score22

    Ranveer Singh Becomes Ray-Ban and Ray-Ban Meta Brand Ambassador in India

    AIMeta names Ranveer Singh the first Brand Ambassador for Ray-Ban and Ray-Ban Meta in India and launches Ray-Ban Meta (Gen 3) there, starting at INR 44,300. Gen 3 offers up to nine hours of battery life, a 12 MP camera, and a 6-mic array that cuts more than 90% of background noise. Ray-Ban Meta Audio, weighing 43 grams, is coming soon.

  11. OpenRouter · New modelsBlogAI score36

    Pareto 26.10 Preview: A Multimodal Model for Research, Coding and Agents

    AIPareto 26.10 Preview is a multimodal composite model built for research, coding, and agentic workflows. It is described as delivering frontier-level performance across a broad range of general-purpose tasks, though the source excerpt is a preview and provides no benchmark scores, parameter counts, pricing, or availability details.

  12. Cloudflare Blog · AIOfficialAI score52

    Cloudflare AI Search reaches general availability with image embeddings and OCR

    AICloudflare's AI Search is now generally available, adding native image embeddings for visual retrieval, OCR for scanned PDFs, and a 10 MiB file limit up from 4 MiB. Billing begins November 1, 2026, charged per ingested token, stored GB-month, and query, with a free monthly allotment on all Workers plans.

  13. Cloudflare Blog · AIOfficialAI score62

    Cloudflare OS opens managed agent workspace waitlist with GitHub and Google Workspace support

    AICloudflare is opening a waitlist for fully managed Cloudflare OS deployments, where organizations configure a custom domain, Cloudflare Access policies, and an AI Gateway. The update lets agents mount existing GitHub repositories to explore code, fix bugs, and open pull requests, and read, draft, and send Gmail while accessing Google Drive. Built-in document, presentation, and spreadsheet tools can now export to Excel, CSV, PDF, Markdown, and HTML, with Word and PowerPoint export coming soon.

    Why it matters: The post shows how a managed agent workspace connects to GitHub and Google Workspace, which matters for teams weighing self-hosting against a managed deployment.

  14. Amazon ScienceOfficialAI score38

    Amazon Science explains graph-centric agentic AI for network root cause analysis

    AIAmazon Science says it built a root cause analysis approach that combines a network digital twin graph, cascaded graph algorithms, and agentic AI orchestration. The approach was demonstrated with NTT DOCOMO at Mobile World Congress, where it identified root causes in minutes on commercial networks. The article traces graph-based network modeling from topology and alarm correlation graphs to graph neural networks.

  15. JetBrains AI BlogOfficialAI score75

    JetBrains Air enters early access as an agent system inside its IDEs

    AIJetBrains has opened the Early Access Program for Air, an agentic development experience available as a plugin on JetBrains Marketplace or in the 2026.3 EAP builds of its IDEs. Air works with existing agents such as Codex, GitHub Copilot, Junie, and Cursor, and it ships with no agents installed. Free Junie Lite runs are offered, while cloud runs require a JetBrains AI subscription.

    Why it matters: The post explains how Air brings existing agents into the IDE, showing a concrete workflow for managing parallel agent sessions alongside code review tools.

  16. WanOfficialAI score62

    Alibaba's Wan 3.0 ranks first overall on Artificial Analysis video leaderboard

    AIAlibaba's Wan 3.0 ranks #1 overall on the new Artificial Analysis AA-Video-T2V v2.0 text-to-video leaderboard, priced at $12 per minute of video. The benchmark judges models at 1080p using over 68,000 human preference votes across 1,000 prompts, and the author states Wan 3.0 leads 10 of 20 category boards.

    Why it matters: The leaderboard breaks results down by use case, capability, and style, and reports price per minute, letting readers compare where each model leads and at what cost.

  17. AI SupremacyBlogAI score50

    Google announces Gemini 4 Argon, its first frontier model since February

    AIGoogle announced Gemini 4 Argon, a model it says is built to sustain deep reasoning across complex, long-horizon workflows, roughly seven months after its last flagship release in February. The article says cybersecurity testing will be completed after October 1, with no benchmark scores, pricing, or availability details provided.

  18. Ai2 (Allen Institute for AI)OfficialAI score62

    Ai2 releases Olmo-core 3, an open framework for training large MoE models

    AIAi2 released Olmo-core 3, an open training framework redesigned to scale mixture-of-experts models into the trillion-parameter range. In one benchmark, expert count rose from 8 to 128 with about 3.2B active parameters per token, total capacity grew from 4.6B to 47B, and throughput fell by less than 5%. The framework is fully open, so researchers can train their own MoEs and experiment with routing and parallelism.

    Why it matters: The release documents concrete MoE scaling results and reported failure modes, useful for teams weighing training-stack tradeoffs before adopting an open framework.

  19. Anthropic ResearchOfficialAI score60

    Matthew Schwartz on finding Claude-shaped science problems with BootLoops

    AIPhysicist Matthew Schwartz describes building BootLoops, an open-source harness for exact quantitative calculations, after choosing problems suited to Claude's strengths. He reports that Claude solved long-standing integrals and found connections across ecology, population genetics, economics, and linguistics, with domain experts steering results toward questions those fields care about. The post states that the approach required constant human oversight, since Claude often overstated results and misjudged time.

    Why it matters: The guest post explains why scientists often find current AI tools frustrating and offers a method for finding problems where AI and researchers match, backed by concrete projects.

  20. Manus BlogOfficialAI score45

    Manus 2.0 Adds Video Editor for Creating and Editing Publishable Videos

    AIManus 2.0 introduces Video Editor, which lets users refine videos Manus generates, including changes to music, captions, and cut timing, without regenerating the entire video. The article describes Manus creating explainers, launch films, and animations from a single prompt, drawing on web search, video models such as Seedance 2.5, and code for motion graphics.

  21. Anthropic NewsroomOfficialAI score38

    Barclays expands Claude across operations, targeting 50% developer adoption by end-2026

    AIBarclays is expanding its collaboration with Anthropic to roll Claude out across its global operations, with Claude Code expected to reach 50% of its developer population by the end of 2026. Its Colleague Knowledge Assistant, powered by Claude through retrieval-augmented generation, has been used by more than 16,000 colleagues and handled over one million searches. In Global Markets, Claude models classify and route roughly 120,000 client emails daily.

  22. LangChain BlogOfficialAI score58

    LangChain shows how to build a model router in its Open SWE coding agent

    AILangChain built a model router inside its open source coding agent Open SWE that picks one of three models for each thread. In an A/B test against always using GPT-6 Astra, the median cost per thread fell 64% with no measurable change in merged PR rate. The router runs on the thread's first message, using a base prompt, per-tier criteria, and a classifier model, and the post lists next steps including subagent routing and mid-thread re-routing.