Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Oct 8

Oct 8Thu
  1. elvisXAI score42

    Voyager: an open harness for creative AI work across video and games

    AIElvis Saravia argues that creative work needs domain-specific agent harnesses rather than coding-oriented ones, and he highlights Voyager as an open harness for video, graphics, and games. According to the quoted post, Voyager lets agents work with local files and drive apps such as Blender, DaVinci Resolve, and Unity, and it is designed to work with models like Opus, Astra, and DeepSeek.

    Video from @omarsar0's post
  2. Google ResearchOfficialAI score22

    Google Research livestreams EmbeddingGemma 2 demo at COLM 2026 today

    AIGoogle Research is hosting a live demonstration of EmbeddingGemma 2 at its COLM booth #107 today at 12:00pm. The open multimodal model unifies text, image, audio, and video representations, with Sahil Dua available to connect with attendees.

    Video from @GoogleResearch's post
  3. Tessl BlogOfficialAI score38

    Mozilla.ai's cq Aims to Give Agents a Shared, Reviewable Knowledge Commons

    AIMozilla.ai's cq project proposes a shared knowledge layer where AI agents capture lessons from non-obvious fixes as structured knowledge units that other agents can later query. The default setup is local-first, using a local SQLite database so nothing leaves the machine, with an option to connect to a remote team server that adds review.

  4. Tessl BlogOfficialAI score52

    Cisco engineer argues agent skills need a context pipeline with evals

    AIJohn Groetzinger, writing in a personal capacity rather than for Cisco, argues that enterprise skills need packaging, evaluation, syncing, and distribution rather than scattered markdown files. He describes using skills to make cheaper models viable, converting curated TAC knowledge-base articles into maintained skills, and rolling out an eval framework across teams. He also describes syncing a repository README to Confluence with a deterministic script.

  5. LiveKitOfficialAI score22

    LiveKit Simulations lets teams test voice agents before customers do

    AILiveKit is offering free access to its Simulations product through October, letting teams check what their agent can do and find gaps before deployment. The product also lets teams test any model against their own scenarios before switching models.

    Video from @livekit's post
  6. 🚨 AI News | TestingCatalogXAI score50

    OpenAI rolls out GPT-6.1 Sol Ultrafast at 8x standard speed

    AIOpenAI is rolling out GPT-6.1 Sol Ultrafast on ChatGPT Work, Codex, and the API. The Ultrafast mode is priced at $12 per million input tokens and $60 per million output tokens, and it runs 8x faster than Sol Standard.

    Video from @testingcatalog's post
  7. Vaibhav (VB) SrivastavOfficialAI score46

    GPT-6.1 Sol Ultrafast rolls out with up to 8x faster token generation

    AIOpenAI is rolling out GPT-6.1 Sol Ultrafast, which generates tokens up to 8x faster than Sol Standard. On the API it is priced at $12 per 1M input tokens and $60 per 1M output tokens. The mode is also available today in Codex and ChatGPT Work.

  8. Artificial AnalysisOfficialAI score34

    Harvey LAB-AA: Artificial Analysis benchmark for legal AI agents

    AIArtificial Analysis has released Harvey LAB-AA, an evaluation built on Harvey's LAB dataset and developed in collaboration with Harvey. Full results are published on the Artificial Analysis evaluations page, alongside Harvey's commentary on the benchmark and human expert preferences.

  9. Artificial AnalysisOfficialAI score42

    More output tokens don't guarantee higher scores in AI benchmarks

    AIArtificial Analysis reports that generating more output tokens does not necessarily yield a higher score. GPT-6 Astra (max) scored 8.6% using about 81k output tokens per task, while Grok 4.7 (xhigh) used roughly 180k yet scored lower. Three Claude models produced the most output tokens, about 202k to 562k per task, but scored between 2.8% and 6.4%.

    Image from @ArtificialAnlys's post
  10. Artificial AnalysisOfficialAI score28

    Artificial Analysis Pareto frontier: GPT-6 Luna cheapest per task at $0.22

    AIAmong models with a Hallucination-Gated All-Pass Rate above 0%, GPT-6 Luna (max), GPT-6.1 Sol (max), Muse Spark 1.3 (max), and Grok 4.7 (xhigh) set the Pareto frontier for score versus cost per task. GPT-6 Luna (max) is the cheapest at about $0.22 per task, scoring 3.3%, while Grok 4.7 (xhigh) leads at about $9.50 per task and Muse Spark 1.3 (max) costs about $4.20. The three Claude models cost about $18 to $22 per task.

    Image from @ArtificialAnlys's post
  11. Artificial AnalysisOfficialAI score44

    Hallucination gating reshuffles AI model rankings, favoring Grok 4.7 over Muse Spark

    AIOnce hallucinations are accounted for, Muse Spark 1.3 (max) drops from 26.7% to 8.9%, leaving Grok 4.7 (xhigh) first on the headline metric at 9.4%. Kimi K3 (max) falls from 16.7% to 5.3%, and Claude Sonnet 5.5 (max with fallback) falls from 11.7% to 2.8%. GPT-6.1 Sol (max) declines least, from 7.5% to 6.9%.

    Image from @ArtificialAnlys's post
  12. Artificial AnalysisOfficialAI score36

    Kimi K3 and Muse Spark 1.3 trade task completion against hallucinations

    AIArtificial Analysis reports that Kimi K3 (max) achieves a 93.0% Criterion Pass Rate but averages 2.09 material hallucinations per task. Muse Spark 1.3 (max) scores higher at 96.0% while averaging 1.68 material hallucinations per task, showing that completing criteria and avoiding hallucinations are distinct skills.

    Image from @ArtificialAnlys's post
  13. Artificial AnalysisOfficialAI score34

    Artificial Analysis compares six hallucination checkers on 20 shared tasks

    AIArtificial Analysis compared six hallucination checkers on the same deliverables from 20 tasks across eight models. GPT-6 Sol and GPT-6 Luna generally flagged the most material hallucinations, while Claude Sonnet 5.5 and Gemini 3.8 Flash flagged far fewer, with Claude Opus 5.5 falling between Grok 4.7 and Sonnet. The counts reflect checker behavior rather than establishing accuracy or ruling out self-preference.

    Image from @ArtificialAnlys's post
  14. AWS Machine Learning BlogOfficialAI score46

    AWS Pays Per Inference for AI Agents with BlockRun and Incarna

    AIAmazon Bedrock AgentCore payments lets AI agents pay for model inference one request at a time, using x402 with USDC on the Base network. Incarna used the service to connect its agents to BlockRun, a pay-as-you-go router serving more than 90 models from more than 15 providers. Spending limits are enforced at the infrastructure layer, outside the model.

  15. 🚨 AI News | TestingCatalogXAI score49

    Voyager desktop app lets AI agents work inside creative tools on Mac

    AIVoyager has launched a Mac desktop app that lets AI agents read project files and operate creative tools such as After Effects, DaVinci Resolve, Blender, and Unity. The agents produce editable results for video edits, motion graphics, color grading, 3D scenes, and game prototypes. Built-in and custom skills, plus a memory that learns each user's workflow, are included.

    Video from @testingcatalog's post
  16. NVIDIA Technical BlogOfficialAI score29

    NVIDIA KGMON Places Second in KDD Cup 2026 Data Agents Competition

    AIThe NVIDIA KGMON team placed second in the KDD Cup 2026 Data Agents competition with a system built around a smaller, clearer, and easier-to-verify agent harness. The competition required agents to answer natural-language questions over heterogeneous sources, including databases, CSV and JSON files, prose documents, PDFs, and briefing videos.

  17. OpenAI DevelopersOfficialAI score47

    OpenAI expands GPT-6.1 Sol Ultrafast access and EU data residency

    AIOpenAI has made Ultrafast mode for GPT-6.1 Sol available in all supported regions, including US and EU data residency. EU data residency has also been added for GPT-6.1 Sol Fast and GPT-6 Luna Fast. Access to Codex and ChatGPT Work is offered on Pro 500, eligible usage-based Enterprise, and credit-based Edu plans, with Enterprise admins required to enable it.

  18. OpenAI DevelopersOfficialAI score38

    GPT-6.1 Sol Ultrafast mode priced at $12 and $60 per million tokens

    AIOpenAI's API pricing for GPT-6.1 Sol in Ultrafast mode is $12 per million input tokens and $60 per million output tokens. The mode is built for time-sensitive work such as debugging outages, agents navigating apps, and live experiences where speed matters.

    Image from @OpenAIDevs's post
  19. OpenAI DevelopersOfficialAI score62

    OpenAI rolls out Ultrafast for GPT-6.1 Sol in API, Codex, and ChatGPT Work

    AIOpenAI says Ultrafast is rolling out today for GPT-6.1 Sol in the API, Codex, and ChatGPT Work. The company describes it as near-Astra intelligence at up to 8x faster speeds than Sol Standard.

    Why it matters: The post names the access points and a speed comparison to the Sol Standard tier, which helps developers judge whether the faster option fits their workflow.

    Video from @OpenAIDevs's post
  20. DatabricksOfficialAI score32

    Databricks' Vibe Data Modeling builds business-specific data models with an agent

    AIDatabricks introduced Vibe Data Modeling, an open-source agent that helps teams build, validate, and evolve business-specific data models. It applies roughly 250 modeling rules while keeping data modelers and business stakeholders involved. Teams can start from 40 industry models as a baseline and iterate toward models that reflect how their business operates.

    Video from @databricks's post
  21. Satya NadellaXAI score11

    Satya Nadella thanks Trump for honor, pledges tech collaboration

    AIMicrosoft CEO Satya Nadella thanked President Trump for an honor bestowed at an event alongside prominent American innovators. He said he looks forward to continuing work together to advance technology and drive American prosperity. The quoted remarks from Trump credit Nadella with decades of transforming Microsoft.

  22. laurenXAI score29

    Omarchy seeks feedback on Grok Bot plugins and integrations

    AILauren Tan invites users of Grok Bot on Omarchy and developers building plugins for it to share feedback and feature requests. The post points to the Omarchy plugin catalog and asks what integrations could be supported. Background from DHH says SpaceXAI joined the Omacom Foundation as a Founding Corporate Patron, contributing $1,500,000 in Grok tokens for Omarchy's maintenance and development.

  23. daniel_tXAI score34

    Photon launches A2A agent discovery in iMessage group chats

    AIPhoton says its iMessage agents can now discover other agents, look up their phone numbers, and add them to a group chat with the user. The company says the feature opens agent access to every merchant and platform, and it is available in beta at photon.codes.

    Video from @danieltian's post
  24. Tessl BlogOfficialAI score42

    Tessl Proposes Executable Specs to Verify AI Coding Agent Output

    AITessl argues AI code review is slow because generated code outpaces trust, and proposes executable specs that let agents check preview environments against product intent. Its spec reviewer splits work between a planner agent that extracts requirements and parallel verifier agents that test each one against the code and base branch.

  25. TechCrunch · AINewsAI score36

    Ben Affleck's AI expertise goes viral as he explains neural networks and fine-tuning

    AIActor Ben Affleck drew attention this week for explaining machine learning concepts, including convolutional neural networks, tensors, and transformers, in several recent interviews. He said he fine-tuned open video models by unfreezing weights and training only the last cinematic layer, using a dataset he built over about eight months for his startup. Affleck said he worries about students and learned helplessness more than Skynet, and predicted AI will be additive to the movie business.

  26. TechCrunch · AINewsAI score46

    Arena raises $200M at $3.1B valuation, nearly doubling in 10 months

    AIArena, the crowdsourced AI model leaderboard that started as a UC Berkeley research project, raised a $200 million Series B at a $3.1 billion valuation, led by Lightspeed Venture Partners and Khosla Ventures. The company said it reached $100 million in annualized run-rate revenue in June, up from $30 million when it raised its $150 million Series A in January at a $1.7 billion post-money valuation.

  27. TechCrunch · AINewsAI score58

    OpenAI's annualized revenue reportedly about $20 billion below earlier estimates

    AIOpenAI has reportedly told investors its annualized revenue is approaching $50 billion, about $20 billion below a previously reported $70 billion figure. The Financial Times reports the earlier number came from investor attempts to compare OpenAI with Anthropic, which counts cloud partners' sales differently. OpenAI's IPO has reportedly been pushed to early 2027.

  28. TechCrunch · AINewsAI score72

    Google launches unified Gemini agent for businesses, consumers to follow

    AIGoogle announced at a Google Cloud event a unified Gemini agent that can plan and complete tasks from a single interface, starting with businesses. The agent has its own Workspace account, connects to systems including Google Workspace, Microsoft 365, Slack, and Jira through MCP, and writes an audit trail attributed to the agent. Google said consumers will get access later, after it addresses security, scale, and performance.

    This story has a top pick“Google Cloud launches Gemini agent as single universal work agent”

  29. The DecoderNewsAI score80

    Mathematicians call for OpenAI boycott after AI-generated proofs flood the field

    AIThe Association of Historical Mathematicians (AHM) has called for a boycott of OpenAI after the company released more than 700 AI-generated proof files at once. Fields Medalist Terence Tao, who chairs the group, argues that AI solving open problems autonomously reduces seminars, collaborations, and fertile research directions, and that the field should shift its measure of progress toward explanation and community-building.

    Why it matters: The article links the AHM boycott call to Tao's argument that AI-driven proof volume is changing how mathematicians measure progress and whether solutions remain useful.

  30. CNBC · TechnologyNewsAI score38

    Trump's August Disclosure Shows Up to $25 Million in Meta and $5 Million in SpaceX Debt

    AITrump disclosed more than 500 securities transactions in August, including a purchase of up to $25 million in Meta stock and up to $5 million in SpaceX senior unsecured notes. The filing, which reports trades in value ranges, shows total activity of roughly $74.3 million to $273.3 million according to a CNBC analysis. The SpaceX notes were bought two days before Trump signed a national space transportation policy, and the White House says the portfolio is independently managed.

  31. Google Cloud TechOfficialAI score40

    Google Cloud's borderless Lakehouse lets Gemini query multicloud data directly

    AIGoogle Cloud's borderless Lakehouse lets Gemini query data on AWS and Azure without variable egress fees. It reads directly from Salesforce Data 360, SAP, ServiceNow, and Workday without copying data. It also federates open Apache Iceberg tables across Databricks Unity, Snowflake Horizon, and AWS Glue.

    Video from @GoogleCloudTech's post
  32. Google Cloud TechOfficialAI score12

    Google Cloud Smart Storage enriches unstructured data in place for Gemini

    AIGoogle Cloud's Smart Storage enriches files in place and writes metadata directly onto source objects, giving Gemini instant context. The post says this keeps security ACLs from drifting, which matters for dark, unstructured data.

    Image from @GoogleCloudTech's post
  33. PyTorch BlogOfficialAI score62

    NVIDIA Dynamo adds session-level IDs to route and cache agentic inference

    AINVIDIA Dynamo uses a unified session-level identifier to make its inference stack aware of agent sessions, subagents, and their KV cache across turns and tool calls. On SWE-bench, two TP4 MiniMax-M2 replicas on one 8xH100 node gained roughly 12-16% throughput from program-aware scheduling over KV-aware routing alone. The post also describes experimental shared-pool indexing and a proposed KvHint interface for session-aware cache policies in vLLM and SGLang.

    Why it matters: The post explains how session identifiers let an inference stack track agent working sets, with measured throughput gains on SWE-bench and agentic RL rollouts.

  34. KalaXAI score34

    Mistral Large 4 and Reflection Beam promise open weights this month

    AIMistral Large 4 and Reflection Beam are previewed now, with Mistral saying weights drop at the end of October and Reflection promising Apache 2.0 weights this month. The post argues that these announced future weights should be treated as a conditional migration dependency, not a current self-hosting option. API previews can be trialed immediately, but they do not prove an unreleased checkpoint will behave the same when downloaded.

  35. Hacker News · Show HN, AI (20+ points)BlogAI score43

    Pocketty is an iPhone SSH terminal that alerts you when an agent is blocked

    AIPocketty is a $99 iPhone and iPad SSH terminal, with a 14-day free trial, that notifies you when an herdr-managed agent is blocked or done. The alert is sealed on your computer for your phone only, and tapping it opens that exact Pane over SSH so you can answer in a real terminal. The source says the relay forwards only sealed bytes and that terminal traffic goes directly between the app and your computers.