Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Oct 9

Oct 9Fri
  1. Artificial AnalysisOfficialAI score22

    HeyGen Voice ranks first in Artificial Analysis assistant and customer service arenas

    AIHeyGen Voice ranks first in the Artificial Analysis Controlled Voice Arena for Assistants at 1,214 and Customer Service at 1,212. It ranks second in Knowledge Sharing at 1,158 and fourth in Entertainment at 1,178. By accent, it ranks first in English (UK) at 1,190 and second in English (US) at 1,206, behind Qwen-Audio-3.1-TTS-Plus at 1,220.

    Image from @ArtificialAnlys's post
  2. Artificial AnalysisOfficialAI score32

    HeyGen Voice scores 83.1% on pronunciation robustness benchmark

    AIHeyGen Voice scores 83.1% overall on Artificial Analysis's Pronunciation Robustness benchmark, ranking #10 of 29 models. The benchmark has human reviewers judge whether text-to-speech models pronounce challenging text correctly against pre-agreed accepted pronunciations. HeyGen Voice places #2 for preserving exact sequences at 83.3%, behind SpaceXAI TTS at 85.7%, and scores 77.0% on expanding shorthand, while Eleven v4 leads that category at 94.1%.

    Image from @ArtificialAnlys's post
  3. publikphigorXAI score7

    Auritrack launches AI tool that tracks where monthly salary goes

    AISolo founder @auritrack says its AI-powered app reads bank statements, answers money questions, and logs expenses from a text. It also emails users a spending report. The product is available at auritrack.com.

    Video from @publikphigor's post
  4. elvisXAI score62

    StepFun's Step 5 Preview targets long coding agent runs

    AIElvis Saravia says he has tested StepFun's Step 5 Preview as a coding agent since early access and found that it checks its own work and stops when tasks are done. The post says the model is built for engineering tasks such as bug fixing, multi-file features, and refactoring, plus frontend generation and financial report output.

    Image from @omarsar0's post
  5. Sophia YangXAI score22

    Open-weight models are net positive for AI infrastructure demand

    AISophia Yang, who runs Mistral's account, says open-weight models are net positive for AI infrastructure demand. Gavin Baker argues that an open-weight token consumes roughly the same compute as a frontier token of a similar-sized model, even though open-weight models compress margins at the model layer.

  6. Yuchen JinXAI score23

    Yuchen Jin says AI found a more elegant math proof

    AIYuchen Jin describes a video of a math PhD and her advisor who spent six months proving a theorem, then found ChatGPT produced a more elegant proof using an approach they had not considered. He argues mathematicians should not despair, pointing to coding, where developers now depend on AI, and says deep domain expertise combined with knowing how to use AI is a strong advantage.

  7. Alex BraginXAI score29

    Solo CTO rebuilds AQUA wallet in React Native with Claude Code

    AIJAN3 CTO Jan Ceuleers says he rebuilt the AQUA wallet from Flutter into React Native, using Claude Code for most of the coding. He started by binding GDK to a React Native app and showing a wallet balance in about half an hour, then ported the Aqua UI component library from Flutter over roughly a year of prior work.

  8. Boris PowerXAI score6

    Boris Power praises an article on population decline and reversal

    AIOpenAI's Boris Power calls an article on population decline "very insightful" and "clearly stated." The post shares the article by Fabrice Grinda titled "Population Decline Is Real. So Is the Reversal," which Sahil Pinker (@sapinker) describes as insightful and original.

  9. Artificial AnalysisOfficialAI score50

    Ideogram 4.5 keeps edited photos intact over 30 consecutive edits

    AIArtificial Analysis ran four image editing models through 30 consecutive edits of the same real estate photo, with each model editing its previous output. Ideogram 4.5 and FLUX 3 left 95% or more of the image essentially untouched on small edits, while GPT Image 2.5 Sunburst re-rendered most of the image each time, leaving only about a fifth unchanged. Nano Banana 2.1 edited locally but shifted and gradually darkened the rest of the image.

    Video from @ArtificialAnlys's post
  10. RadixArkOfficialAI score28

    Miles v0.1.2 adds Kubernetes support and torchtitan training backend

    AIRadixArk releases Miles v0.1.2, adding score centering for stable async RL and an experimental native Kubernetes backend. RL jobs can run as ordinary cluster workloads, and orchestration can restart while training continues. The release also adds torchtitan as a third training backend alongside Megatron and FSDP, plus support for DeepSeek-V4.1-Flash and MiMo-V2.6-Flash-RL.

    Image from @radixark's post
  11. Microsoft CopilotOfficialAI score30

    Microsoft unveils new Copilot Home combining Chat and Cowork

    AIMicrosoft says the new Copilot Home brings Chat and Cowork into one experience, so users can move from thinking to doing without losing context. Microsoft Copilot EVP Jacob Andreou explains why Home is the new starting point for work.

    Video from @MSFTCopilot's post
  12. TechCrunch · AINewsAI score36

    People instinctively treat AI and robots as human, experts warn

    AIAmanda Silberling describes how she greeted a Unitree humanoid robot at MIT's CSAIL as a person, and how MIT researchers Sherry Turkle and Pat Pataranutaporn say people instinctively treat chatbots as caring companions. Turkle writes that people are "wired to care for" relational artifacts, and a study found about 70% of people are polite to AI. Pataranutaporn, who served as an expert in a wrongful death lawsuit against Character.AI, warns that people may favor chatbots over other humans.

  13. TechCrunch · AINewsAI score22

    Danu Robotics builds H.E.R.O., a recycling robot that sorts waste with a pincer claw

    AIDanu Robotics, an Edinburgh startup founded by Amy Ma, is bringing its H.E.R.O. recycling-sorting robot to market, using a pincer claw instead of suction. The company has letters of interest from two large customers, $500,000 in signed contracts, and more than 200 prospective customers in its sales pipeline. Ma estimates a working site would earn $485,000 in added revenue against a $160,000 initial investment and $24,000 in annual maintenance fees.

  14. TechCrunch · AINewsAI score36

    Amazon drops data center NDAs as AI agents seek consumer credit cards

    AIAmazon says it will stop using NDAs when negotiating data center deals with local governments, following a similar move from Microsoft earlier this year. Secrecy has fueled community backlash against AI infrastructure, with hundreds of proposed and enacted moratoriums from New York to San Francisco. A wave of startups is betting consumers will give AI agents access to their inboxes, files, and credit cards.

  15. LangChain BlogOfficialAI score40

    LangChain adds emoji reactions to Managed Deep Agents Slack channels

    AILangChain's Managed Deep Agents v0.9 adds a reactions attribute for Slack channels that accepts either an emoji string or a callable returning one. The article shows a function that returns a bug emoji when a message contains "broken" and eyes otherwise. It also shows a TypeSafe Classifier that picks from a seven-emoji vocabulary and falls back to eyes below 25% confidence.

  16. NewcomerBlogAI score42

    OpenAI's ARR estimate falls $20 billion short of reported $70 billion figure

    AIThe Financial Times reported that OpenAI's annual recurring revenue was $20 billion below a widely cited $70 billion figure, a gap the paper tied to gross versus net revenue accounting. OpenAI has used a net figure while Anthropic reports gross, which includes partner sales by Amazon and Microsoft. Bloomberg reports OpenAI expects to reach $70 billion annualized revenue by year end on a net basis.

  17. a16z NewsBlogAI score40

    a16z leads investment in TypeSafe AI, maker of Jev System One model

    AIa16z says it is leading an investment in TypeSafe AI, whose Jev model hands decisions to code as typed values and reached 1 trillion tokens generated three days after launch. The company says Jev costs roughly 1/100 to 1/500 of frontier models and runs 100x faster on classification tasks at comparable accuracy. TypeSafe says 25% of the Fortune 500 have integrated Jev.

  18. a16z NewsBlogAI score33

    Prediction markets show no partisan bias in election pricing, NBER study finds

    AIA preliminary NBER working paper by Prof. Zitzewitz, covering over 100 years of prediction markets, finds no statistically significant bias by political affiliation, gender, race, or age. The only exception is non-US elections, where markets appear to overrate right-leaning candidates, but that result is not statistically significant. Separately, prediction markets had Flávio Bolsonaro's Brazilian presidential rise about three weeks before his first-round win.

  19. O'Reilly RadarBlogAI score40

    US AI oversight debate, OpenAI Dots, and Gemini 4 Argon featured in This Week in AI

    AIThe Trump administration announced a voluntary agreement with major AI companies calling for internal safety monitoring, external audits, and independent board reviews, and the Federal Trade Commission launched an investigation into OpenAI, Anthropic, and other AI companies over potential consumer risks. OpenAI released Dots, a proactive assistant that retains context, works across applications, and acts without waiting for prompts. Google says Gemini 4 Argon can generate up to a million output tokens in a single response.

  20. 🚨 AI News | TestingCatalogXAI score41

    Pine AI launches Pine Computer, a cloud runtime for agentic tasks

    AIPine AI launched Pine Computer, a cloud computer, harness, and runtime layer built for agentic tasks. On the publisher's SaaS-Bench v1.1, it posts a 78.3% checkpoint score against 74.3% for Opus 5 with Claude Code, but completes fewer whole tasks, 27.4% against 31.1%. Instead of simulating clicks and screenshots, it reads web pages as structured data, and access is through a private beta waitlist.

    Image from @testingcatalog's post
  21. 🚨 AI News | TestingCatalogXAI score62

    Anthropic moves dynamic workflows in Claude Managed Agents into public beta

    AIAnthropic has expanded dynamic workflows in Claude Managed Agents into a public beta, according to Testing Catalog. Users can configure their agents for multiagent orchestration, with Claude planning and operating a fleet of agents to achieve a goal. The post also links a video from Anthropic's ClaudeDevs account, which the author describes as a new SWE norm.

    Video from @testingcatalog's post
  22. SantiagoXAI score44

    Pine launches agentic cloud computers with built-in AI agents

    AIPine has released a cloud computer service with a built-in AI agent that applications can control through its SDK. Developers give the agent a plain-English task, and it can use a browser, files, and a shell while the app receives notifications and final outputs. Pine's Stanley Wei says the computer is built for AI rather than humans.

    Video from @svpino's post
  23. Vaibhav (VB) SrivastavXAI score4

    OpenAI rolls out invites for DevDay Exchange Berlin, Paris, and London

    AIOpenAI says invites for its DevDay Exchange events in Berlin, Paris, and London are rolling out now. Recipients are asked to register as soon as possible to secure a spot. People still waiting can reply with their city, what they are building or exploring, and why they want to attend.

    Image from @reach_vb's post
  24. Alexandr WangXAI score12

    Alexandr Wang discusses AI safety in a conversation with Cleo Abram

    AIAlexandr Wang calls Cleo Abram brilliant and says their conversation on AI safety is one of his favorites. He also says he included a subtle dig at some well-known people during the interview. The interview covers what Muse can do, when it can be trusted, and how to make sure the technology goes right.