Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Oct 8

Oct 8Thu
  1. PandailyNewsAI score57

    Shanghai AI Lab Open-Sources Intern-Decision Small Models for Structured Decisions

    AIShanghai AI Lab has open-sourced Intern-Decision, a family of 0.8B, 2B and 4B parameter models that return structured decisions with probabilities instead of free text. The developers self-report that the 4B model averages 90.02% accuracy across seven test suites, ahead of a commercial reference model at 88.74%, with about 44 milliseconds of local latency on a single RTX 4090. Weights are on Hugging Face, and MetaX says the models run on its hardware from launch.

  2. SiliconANGLE · AINewsAI score47

    Kore.ai launches Autoloop to tune enterprise AI agents after deployment

    AIKore.ai launched Autoloop, an optimization engine that automatically adjusts AI agents built on its Kore.ai Agent Platform to meet business-set goals, including after deployment. The engine scores each proposed change against goals such as task completion, business-rule adherence, accuracy and cost. Autoloop is available now to all customers on the Artemis edition of the Kore.ai Agent Platform.

  3. TechCrunch · AINewsAI score62

    Goodfire launches inside-out monitors to catch rogue AI agents at lower cost

    AIGoodfire has launched monitors that read a model's internal signals during agent work instead of reviewing its written output. The monitors are available to Baseten customers, who can choose risks to watch and set automated responses. In Goodfire's tests on Kimi K3, monitoring about 1,500 sessions cost roughly $51 versus about $10,000 for a top-tier AI judge, while catching 94% of malicious hacking sessions.

  4. TechCrunch · AINewsAI score43

    Natura's $99 Interface smart ring lets users control AI agents with a finger press

    AINatura's Interface is a $99 smart ring that lets users ask AI agents to complete tasks with a finger press and also tracks heart rate, HRV, sleep, and activity. At launch it connects with Meta's Muse, Instinct, Grok Bot, Claude, ChatGPT, and more, with preorders expected next month and shipping slated for December or January. After a three- to six-month free period, Natura plans to charge a $9 monthly subscription.

  5. Jerry LiuXAI score38

    LightOn OCR-3 now on OpenDocRouter, near Gemini 3.8 Flash at lower cost

    AILightOn OCR-3 is now available on OpenDocRouter at $0.28 per 1M input tokens and $1.40 per 1M output tokens, about $3.19 per 1k pages on ParseBench. On ParseBench, the author says it sits on the Pareto frontier for open-weight OCR models, with performance similar to Gemini 3.8 Flash low at roughly 45% lower price. It is described as decent at tables, workable for charts, and quite good at grounding.

    Image from @jerryjliu0's post
  6. SantiagoXAI score34

    Atomic Agent Desktop launches with Linux support on day one

    AIAtomic Agent Desktop is a local AI-first agent available for Mac, Windows, and Linux. It offers a 4x larger context window on local models using TurboQuant, and Atomic Fusion orchestrates cloud and local models to reduce costs.

  7. 🚨 AI News | TestingCatalogXAI score62

    Atomic Agent Desktop, an open-source local AI agent app, is now available

    AIAtomic Agent Desktop is a free open-source app for macOS, Windows, and Linux that runs open models like Qwen and Gemma locally without an account. It connects to a cloud model only when selected, and its Fusion feature lets a cloud model plan a task while up to 8 local agents carry it out. The post's own text adds a setup wizard that checks RAM and suggests suitable models, and import from Claude Code, Codex, Hermes, and OpenClaw.

    Video from @testingcatalog's post
  8. Artificial AnalysisOfficialAI score7

    Artificial Analysis publishes AA-Video-T2V v2.0 prompt for snowy cabin scene

    AIArtificial Analysis shares the second part of an AA-Video-T2V v2.0 prompt describing a four-shot documentary-style handheld video of a glass cabin in falling snow. The shots follow a caretaker sweeping snow from the deck, empty snow-covered windows, an empty interior, and the same caretaker stamping snow off his boots at the door, with hard cuts between shots.

    Video from @ArtificialAnlys's post
  9. Artificial AnalysisOfficialAI score11

    Artificial Analysis releases AA-Video-T2V v2.0 video generation prompt

    AIArtificial Analysis shares an AA-Video-T2V v2.0 text-to-video prompt, a 10-second product-page scene of a presenter revealing a volcano-shaped mist diffuser. The prompt specifies a shift from warm daylight to dim evening lamplight, fixed close-up camera angles, and consistent presenter and product appearance across shots.

    Video from @ArtificialAnlys's post
  10. Matt ShumerXAI score18

    Spawn brings playable games directly into X posts

    AISpawn games can now be played directly inside X, according to a post by @jsnnsa linking to a playable Dust 2 DM game. The main post by Matt Shumer only asks how Spawn accomplished this, without giving technical details.

  11. Higgsfield AI 🧩OfficialAI score10

    Higgsfield launches a tool to create custom AI influencers

    AIHiggsfield AI promotes a tool on its platform for users to create their own AI influencer, with a link to the katana page. The post gives no details on features, pricing, or models.

  12. Sara HookerXAI score10

    Adaption AI releases blog on adaptive checklists for AI evaluation

    AISara Hooker announced a new blog post on adaptive checklists from Adaption AI, which she says capture current model capabilities and domain-specific requirements better than traditional checklists. She claims these adaptive checklists achieve this without requiring human supervision.

    Image from @sarahookr's post
  13. Higgsfield AI 🧩OfficialAI score34

    Higgsfield launches Katana AI video editing tool inside Claude

    AIHiggsfield introduced Katana, its most powerful AI video editing tool, powered by Claude Motion and available now inside Claude via Higgsfield MCP. Users can upload a reference to create editable motion graphics, product launch videos, or aura-farming edits.

    Video from @higgsfield's post
  14. OpenRouterOfficialAI score29

    Mercury Decide now available with zero data retention on OpenRouter

    AIInception's Mercury Decide, the fastest-growing decision model on OpenRouter last week, now has a paid zero data retention (ZDR) endpoint alongside the free one. It is priced at $0.02/M input tokens, half the $0.04 list price, with output and cached input free and a 66K context window.

    Video from @OpenRouter's post
  15. Artificial AnalysisOfficialAI score38

    GPT-6 Sol (Daybreak Blue) tops Artificial Analysis Cyber Index

    AIGPT-6 Sol (Daybreak Blue, max) has been added to the Artificial Analysis Cyber Index as a trusted-access model and ranks #1 on the Index. Compared with the publicly available GPT-6 Sol, it shows its largest gains on CyberGym-E2E, the benchmark where the most safety refusals are observed.

    Image from @ArtificialAnlys's post
  16. Boris PowerXAI score46

    OpenAI's GPT-6.1-Sol leads new Arena Alignment Index for agents

    AIThe Arena Alignment Index, built from over 90K real-world agent sessions across 27 models, ranks OpenAI's GPT-6.1-Sol first with a score of 87.9, ahead of Claude-Opus-5.5 at 83.2 and Grok-4.7 at 82.7. GPT-6.1-Sol also posted the lowest observed rates across the index's three signals: 0.89% Unauthorized Action, 1.98% False Attribution, and 2.34% Deceptive Completion. The index's authors report that newer models consistently outperform their predecessors across all four labs, suggesting broad progress in agent safety.

  17. elvisXAI score42

    Voyager: an open harness for creative AI work across video and games

    AIElvis Saravia argues that creative work needs domain-specific agent harnesses rather than coding-oriented ones, and he highlights Voyager as an open harness for video, graphics, and games. According to the quoted post, Voyager lets agents work with local files and drive apps such as Blender, DaVinci Resolve, and Unity, and it is designed to work with models like Opus, Astra, and DeepSeek.

    Video from @omarsar0's post
  18. Tessl BlogOfficialAI score38

    Mozilla.ai's cq Aims to Give Agents a Shared, Reviewable Knowledge Commons

    AIMozilla.ai's cq project proposes a shared knowledge layer where AI agents capture lessons from non-obvious fixes as structured knowledge units that other agents can later query. The default setup is local-first, using a local SQLite database so nothing leaves the machine, with an option to connect to a remote team server that adds review.

  19. LiveKitOfficialAI score22

    LiveKit Simulations lets teams test voice agents before customers do

    AILiveKit is offering free access to its Simulations product through October, letting teams check what their agent can do and find gaps before deployment. The product also lets teams test any model against their own scenarios before switching models.

    Video from @livekit's post
  20. 🚨 AI News | TestingCatalogXAI score49

    Voyager desktop app lets AI agents work inside creative tools on Mac

    AIVoyager has launched a Mac desktop app that lets AI agents read project files and operate creative tools such as After Effects, DaVinci Resolve, Blender, and Unity. The agents produce editable results for video edits, motion graphics, color grading, 3D scenes, and game prototypes. Built-in and custom skills, plus a memory that learns each user's workflow, are included.

    Video from @testingcatalog's post
  21. laurenXAI score29

    Omarchy seeks feedback on Grok Bot plugins and integrations

    AILauren Tan invites users of Grok Bot on Omarchy and developers building plugins for it to share feedback and feature requests. The post points to the Omarchy plugin catalog and asks what integrations could be supported. Background from DHH says SpaceXAI joined the Omacom Foundation as a Founding Corporate Patron, contributing $1,500,000 in Grok tokens for Omarchy's maintenance and development.

  22. SantiagoXAI score42

    Voyager: open harness connecting AI models to creative apps like Blender

    AIVoyager is an open harness for creative work that connects models with applications to build videos, graphics, and games. It works with Blender, DaVinci Resolve, After Effects, Ableton, and Unity, operating similarly to Codex or Claude Code. The harness is designed to get strong creative results from models such as Opus, Astra, and DeepSeek.

    Video from @svpino's post
  23. OpenRouterOfficialAI score24

    Rasp AI sales agent saves 600 monthly hours using OpenRouter

    AIA five-person sales team built Rasp, an AI sales agent on OpenRouter's Ori, to research every inbound lead. Many leads reportedly get a response in under 60 seconds, saving about 600 hours per month at roughly $30 per day in cost.

  24. Jerry LiuXAI score41

    OpenDocRouter offers one API for many document OCR models

    AIOpenDocRouter is a unified API and billing interface for document OCR models, ranging from lightweight open-source options like MinerU to frontier VLMs like Opus 5.5. Per the linked post, models are served at cost with a small transaction cut, rate limits are handled, and bounding boxes and layout are offered as a service.

    Video from @jerryjliu0's post
  25. Alexander DoriaXAI score46

    LightOnOCR-3 claims state-of-the-art OCR performance under 1B parameters

    AILightOn has released LightOnOCR-3, a family of OCR models in 0.8B and 4B versions that it says lead benchmarks including OlmOCR-Bench and ParseBench, with the 0.8B model positioned as the sub-1B option. The models recognize text, handwriting, images, charts and document structure in one pass, process documents up to twice as fast as LightOnOCR-2, and are released under the Apache 2.0 license.

    Image from @Dorialexander's post
  26. 🚨 AI News | TestingCatalogXAI score49

    Odyssey launches Odyssey-3 world model with public research preview

    AIOdyssey has launched Odyssey-3, its most powerful foundation world model, with a public research preview. Odyssey-3 Pro scored 66.1 on Physics-IQ Verified video-to-video with best-of-8 sampling, the highest reported result. The model generates environments from prompts and predicts changes in real time as users move through scenes.

    Image from @testingcatalog's post
  27. elvisXAI score42

    Odyssey-3 Pro tops Physics-IQ Verified and shows robotic error recovery.

    AIOdyssey released Odyssey-3 Pro, which sets a new top score on Physics-IQ Verified, a benchmark where models continue videos of real physics experiments. In robotics, a robot arm with tens of hours of demonstrations recovered from a missed grasp, a behavior absent from those demos.

    Video from @omarsar0's post
  28. GeneralistOfficialAI score28

    Generalist releases GEN-1.5, a foundation model for physical-world robotics

    AIGeneralist has announced GEN-1.5, its latest foundation model for the physical world. The post provides only a link to the company's blog for further details, so no specifications, benchmarks, or availability information can be confirmed from this source.

  29. SantiagoXAI score46

    Odyssey 3 Pro world model tops Physics-IQ and goes live

    AIOdyssey 3 Pro, a world model, is now live as a research preview and ranks first on the Physics-IQ Verified video-to-video benchmark. The post says it can learn from visual observations and map that knowledge to physical controls for robots, cars, video games, and drones. Odyssey-3, the model launched alongside it, is described as free to try.

    Image from @svpino's post
  30. OdysseyOfficialAI score31

    Odyssey-3 world model debuts for physical AI and training environments

    AIOdyssey has released Odyssey-3, which it describes as a major leap toward world models that power physical AI, generate training environments, and enable new human experiences. The post invites readers to try Odyssey-3 at the company's website but gives no specific benchmarks, parameter counts, or pricing.

  31. OdysseyOfficialAI score34

    Odyssey-3 world knowledge can be applied to physical AI systems

    AIOdyssey says its Odyssey-3 model's learned world knowledge can be adapted by physical AI developers to control robots, power humanoids, drive cars, and fly drones. The post describes this as a capability for autonomous machines generally, without providing benchmarks, specifications, or availability details.

    Video from @odysseyml's post
  32. OdysseyOfficialAI score38

    Odyssey-3 is a foundation world model for physical AI and agents

    AIOdyssey announced Odyssey-3, a foundation world model it says enables applications in physical AI, human experiences, and training intelligences. The company highlights agents learning from experience inside Odyssey-3 while working toward objectives.

    Video from @odysseyml's post
  33. OdysseyOfficialAI score22

    Odyssey-3 Pro sets new Physics-IQ video-to-video benchmark record

    AIOdyssey-3 Pro achieved a score of 66.1 on Physics-IQ Verified's video-to-video benchmark, the highest reported score so far. Physics-IQ evaluates physical behavior across fluid dynamics, optics, solid mechanics, magnetism, and thermodynamics.

    Image from @odysseyml's post