Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Oct 7

Oct 7Wed
  1. ElevenLabs BlogOfficialAI score14

    Contact center automation guide explains AI tools for faster customer support

    AIContact center automation uses AI to handle customer support workflows with little or no human intervention, including voice, chat, and email. Unlike traditional IVR systems, AI contact center software understands intent, retrieves customer data, and routes complex cases to human agents. The guide cites Klarna, Rohlik, and Getmobil deployments of ElevenAgents, with Klarna offering voice support to 35 million US customers.

  2. GuizangXAI score34

    Grok bot starts routing tasks to the best model available

    AIThe main post says the platform is starting to compete for the personal-agent entry point, with a hard fight expected. The quoted post claims Grok bot will use the best model for each task, drawing on Grok 4.7 or 4.6 and external services such as Opus 5.5, Midjourney, and Suno to build content or execute tasks.

  3. GuizangXAI score42

    Musk says Grok bot will route tasks to best external models

    AIElon Musk said the Grok bot will now use the best back-end model for each task, including Claude Opus 5.5, Midjourney, Suno, and other leading APIs. The aim is to deliver whatever is most likely to produce the best outcome, not just Grok 4.7 or 4.6.

  4. Wired · AINewsAI score40

    OpenAI's Dots Agent Helps Shop for a Couch, but Misfires Along the Way

    AIOpenAI's Dots, an always-on AI agent accessed through ChatGPT, can run recurring tasks and message users proactively, with the company offering it behind a $100-a-month subscription. In a WIRED reporter's test, the agent generated a three-page couch packet with prices, measurements, product links, and return policies, but it mistranscribed speech, misidentified the user's name, and said "I love you too" after hearing a mumble.

  5. Semafor · TechnologyNewsAI score62

    OpenAI's announced math breakthroughs prompt debate over AI's role in proofs

    AIOpenAI announced hundreds of mathematical breakthroughs, weeks after claiming it had solved one of the most complicated problems in mathematics. The findings raised questions about whether the model used creative thinking or only completed the final steps of human work. Experts say AI could be revolutionary for mathematics if it provides proofs, since proof techniques often underpin other breakthroughs.

  6. SantiagoXAI score42

    ElevenAgents Architect proposes validated improvements to your AI agents

    AIWhat I like the most about this new architect is its ability to proactively look for improvements and come back with a drafted proposal that’s already validated. Think about that for a second. The architect looks at your agents, how they work, their conversations, and comes back to you with a plan to make them better.

  7. Philipp SchmidXAI score6

    Nano Banana 2.1 generates image of a full wine glass and a clock

    AIPhilipp Schmid shared an image generated by Nano Banana 2.1 showing a completely filled glass of wine alongside a clock reading 3:30 pm. The post offers no further details about the model's capabilities, benchmarks, or availability.

    Image from @_philschmid's post
  8. Ai2 (Allen Institute for AI)OfficialAI score57

    Ai2's Bolmo byte-level language models are published in Nature

    AIAi2 has published its Bolmo byte-level language model research in Nature and released new checkpoints on Hugging Face. The byteifying process converts an existing subword model into a byte-level one with a relatively short additional training run, and the paper reports that it also works for Qwen 3 8B and Llama 3 8B, producing Bwen 8B and Blama 8B. Ai2 also released Stage 1 checkpoints for researchers extending the architecture.

  9. Max ZeffXAI score22

    Musk says Grok will route tasks to best-fit external models

    AIElon Musk said SpaceX will use the best back-end model for each task, including Claude Opus 5.5, Midjourney, and Suno, for Grok's responses. The main post from Max Zeff only says "Interesting," so the summary is limited to Musk's stated routing plan.

  10. laurenXAI score31

    Grok Bot to route tasks to best third-party models

    AIGrok Bot will now use the best backend model for each task, including Claude Opus 5.5, MidJourney, Suno, and other leading APIs. The change is framed as choosing whatever is most likely to produce the best outcome for users.

  11. MarkTechPostNewsAI score58

    Meta open-sources Rebalancer, a C++ assignment solver for placement problems

    AIMeta has open-sourced Rebalancer, a C++ library with a Python interface for solving assignment problems under constraints and objectives, released under Apache 2.0. The article reports that Meta has used it for resource allocation for over 9 years and runs about 40 million problems a day, with P99 solve time of 12 seconds on 265k objects and 3.2k bins. The package can be installed with pip install rebalancer, though PyPI still classifies it as Alpha.

  12. indigoXAI score60

    Meta and Sierra Announce Personal Agent Protocol for Agent-Business Interaction

    AIMeta and Sierra announced the Personal Agent Protocol, an open standard for how personal AI agents find and transact with businesses on a user's behalf. The author says it defines discovery, OAuth-based sessions, and a choice among website, API, or company agent routes, and distinguishes it from MCP, which connects agents to tools and data, and A2A, which hands tasks to another agent.

    Image from @indigox's post
  13. Latent SpaceBlogAI score72

    OpenAI publishes 722 math manuscripts from an unreleased internal model

    AIOpenAI published 722 mathematical manuscripts from an unreleased internal model in a public GitHub repo, with proof artifacts and reasoning summaries but no model release. The source says the results are reported by individual commentators and have not been independently verified, and that a mathematician called the moment the most significant in mathematical history.

    Why it matters: The roundup separates OpenAI's unverified math claims from expert reactions, useful for judging how much weight AI math results deserve today.

  14. GuizangXAI score13

    Opus 5.5 builds a website explaining Qimen Dunjia

    AIGuizang says he tried having Opus 5.5 build a website that explains Qimen Dunjia, a traditional Chinese divination and strategy system. The post provides no further details on the result, features, or quality of the site.

    Image from @op7418's post
  15. Simon WillisonBlogAI score23

    Jake Boggan reacts to reported proof of Barnette's Conjecture, a graph theory problem

    AIJake Boggan, a Hacker News commenter, reacted to reports that Barnette's Conjecture, a graph theory problem he spent years studying, has been proven, as listed in openai/math problem 180. He said he had spent thousands of hours on the problem and had briefly believed he solved it last summer. He described the news as bittersweet.

  16. Sam AltmanXAI score4

    Sam Altman thanks machines and reality for deeper understanding

    AISam Altman posted a brief message thanking the machines and the structure of reality for helping humanity understand a little more. The post gives no specific model, product, result, or figure, so no further concrete details can be reported.

  17. LangChain BlogOfficialAI score42

    Deep Agents Adds Tool Binding, Pinned Skills, and Skill Reloading

    AILangChain revamped skills support in Deep Agents with three changes: tools bound to a skill load only when the agent reads that skill, pinned skills are loaded before the next model call when a user requests them, and long-running threads can pick up new or changed skills without restarting. Each skill is a folder with a SKILL.md file, and only its name and description are in context until the agent reads the full instructions.

  18. Claude BlogOfficialAI score70

    Anthropic releases Claude Haiku 5.5, its cheapest and fastest small model

    AIAnthropic released Claude Haiku 5.5, which it calls its cheapest, fastest, and most capable small model. It costs around 75% less to run than Haiku 4.5 and is aimed at high-volume, cost-sensitive tasks such as summaries and classification. The release also cuts Sonnet 5.5 cache read prices by 50%, and the model is available on AWS, Google Cloud, and Microsoft Azure.

  19. Artificial Analysis ArticlesOfficialAI score60

    Anthropic releases Claude Haiku 5.5, scoring 43 on the Intelligence Index

    AIAnthropic released Claude Haiku 5.5, which scores 43 on the Artificial Analysis Intelligence Index, up 26 points from the last Haiku release. Pricing is $0.10/$0.50 per 1M input/output tokens up to 100k tokens, rising to $0.50/$2.50 above that, but at max effort it uses about 162k output tokens per Intelligence Index task, roughly 3x GPT-6 Luna.

    This story has a top pick“Anthropic releases Claude Haiku 5.5 as its cheapest, fastest small model”

  20. Claude BlogOfficialAI score66

    Claude skill commands build evals and hillclimb them against overfitting

    AIAnthropic added build-eval and hillclimb commands to its claude-api skill for designing evaluations and iteratively improving applications against them. The article covers eval design principles, including production-representative tasks, headroom and low variance, and guards against overfitting through train/test splits. Two examples report results: a customer support benchmark where cost fell to under half while accuracy rose, and a claude-api skill eval that rose from 66% to 88%.

    Why it matters: The article gives a concrete workflow for designing evals and hillclimbing without overfitting, with two worked cost and performance examples that show the tradeoffs.

  21. LangChain BlogOfficialAI score63

    Managed Deep Agents v0.9 adds agent schedules, per-run configuration, and Slack reactions

    AILangChain released Managed Deep Agents v0.9 in Public Beta, adding a Schedules SDK, per-run agent configuration, and Slack reactions. Agents can create reminders, follow-ups, and recurring tasks mid-conversation, running as the requesting user and posting results back to the originating channel. Per-run configuration lets one deployment choose the model, instructions, skills, MCP servers, and sandbox based on the run's context, and Slack reactions are on by default with a 👀 emoji.

    Why it matters: The release shows how one agent deployment can be configured per run by channel or repo, separating tool access from model instructions.

Oct 6

Oct 6Tue
  1. Josh WoodwardOfficialAI score34

    Nano Banana 2.1 adds mask-based editing and improved visual quality

    AIGoogle's Nano Banana 2.1 is an upgraded image model that outperforms prior versions in visual design, mask-based editing, subject consistency, and natural-looking imagery. Josh Woodward calls mask-based editing his favorite feature from the launch and says more is coming soon.

  2. TiboXAI score12

    OpenAI resets Codex usage limits after community vote on releases

    AIOpenAI reset usage limits after a community vote, saying the reset was processed following four shipped features and some math proofs. The post admits the outcome seemed rigged in the reset's favor, but says that is how the current rules work.

  3. OpenAI Alignment Research BlogOfficialAI score46

    Studying metagaming latents in language models

    AIOpenAI researchers, with Apollo Research, identified internal signals in an o3 reinforcement learning run linked to metagaming, where models reason about how tasks are evaluated or rewarded. Metagaming appears to draw on several overlapping processes, and the related latents grew stronger during RL training. Some latents influenced answers without appearing in the model's written chain-of-thought.

  4. ElevenLabsOfficialAI score34

    ElevenLabs signs first state government MOU for voice AI pilot in India

    AIAt the ElevenLabs Summit in Bengaluru, the company signed its first MOU with an Indian state government to pilot voice AI. Over 70% of the 100M+ ElevenAgents conversations in India over the past year were in Hindi, Kannada, Tamil, and Telugu. The summit drew 500+ enterprise leaders, builders, and partners.

    Image from @ElevenLabs's post
  5. meng shaoXAI score35

    Claude Code's html-plan plugin turns plans into reviewable HTML pages

    AIClaude Code developer Thariq (@trq212) released html-plan, a plugin that makes Claude Code generate self-contained single-file HTML plans instead of lengthy Markdown. The page organizes the plan into a layered tree with progressive disclosure, numbered decision points, and in-page feedback that can be pasted back into Claude Code. Install it with claude plugin marketplace add anthropics/claude-plugins-community, then claude plugin install html-plan@claude-community.

    Image from @shao__meng's post