Skip to contentSkip to stories

Updated

All AI news

Showing low-relevance items too. Hide low-relevance items

Oct 9

Oct 9Fri
  1. Don't Worry About the Vase (Zvi Mowshowitz)BlogAI score73

    OpenAI releases 719 AI-generated math manuscripts, splitting the mathematics community

    AIZvi Mowshowitz reports that OpenAI released 722 math manuscripts from an internal frontier model on GitHub, later reduced to 719 after three withdrawals, covering 90 of the top 500 open problems. He says the work came mostly from a single prompt, with an average of three hours of compute per solution. Mathematicians reacted with mixed feelings, and the post highlights concerns about unread papers, cryptography implications, and the role of Lean verification.

    Why it matters: The post traces how OpenAI's release of 719 math manuscripts divided mathematicians and reshaped verification, credit, and publication norms in the field.

  2. Simon WillisonBlogAI score27

    Simon Willison builds a new blog feature largely by voice with Codex

    AISimon Willison says he built a Newsletters index for his blog almost entirely by voice, using the ChatGPT desktop app's Codex voice mode while cooking dinner. The feature imports weekly Substack posts via RSS and undocumented API, monthly newsletters from a GitHub archive repository, and a private sponsors-only newsletter. He says he switched back to typing for review and fixes before deploying the pull request.

  3. Rohan PaulXAI score46

    Microsoft's TeleTune evolves agent skills from raw usage logs

    AIMicrosoft researchers present TeleTune, which lets agents learn software skills from raw usage logs by keeping only skill edits that better predict users' next actions. The method needs no live test environment, because next-action accuracy on held-out logs tracked live success. Unlike earlier methods such as Agent Workflow Memory, which need goal-labeled examples or a live environment, TeleTune guesses each session's goal and uses wrong guesses to suggest edits to a text skill library.

    Image from @rohanpaul_ai's post
  4. Sakana AIOfficialAI score37

    Sakana AI paper uses LLMs to catch errors in research papers

    AISakana AI researchers introduce a benchmark that plants contradictions in papers to test whether LLM reviewers can detect errors, and propose Multi-Layered Review, modeled on the Three-Pass Approach to reading. Their system detected more errors than the other review systems tested, including in papers withdrawn for real mistakes, while its paper-quality assessments stayed broadly consistent with human judgments. The work, accepted at TMLR, is framed as support for human reviewers rather than a replacement.

    Video from @SakanaAILabs's post
  5. Google WorkspaceOfficialAI score12

    RigStrips uses Google Workspace with Gemini to draft customer replies faster

    AIRigStrips co-founder Zhach Pham says he uses Google Workspace with Gemini to draft customer support replies instantly, saving hours that would otherwise go to outdoor gear testing. The post highlights the company's more than 150,000 units shipped, though it gives no specific figures for time saved.

    Video from @GoogleWorkspace's post
  6. 🚨 AI News | TestingCatalogXAI score22

    Google spotted testing an Ultra mode in AI Studio Build

    AITesting Catalog reports that a new Ultra mode is in development in Google AI Studio Build, described as building with advanced skills and tools, alongside Plan, Build, and a previously spotted Security review mode. The post gives no details on which tools or skills it uses, and it does not say whether Ultra mode will require an Ultra subscription. The author suggests these modes are likely being developed to work with Gemini 4 Argon.

    Image from @testingcatalog's post
  7. LangChainOfficialAI score34

    Snyk's Assist support agent handles 60k queries with 85% resolution

    AISnyk's Assist, a customer support agent built on LangChain and LangGraph with observability in LangSmith, has handled over 60,000 queries for more than 500 customer accounts. Over 85% of sessions are resolved without a support ticket, and more than 250 cases were automatically detected and escalated to the right team.

    Image from @LangChain's post
  8. LeiphoneNewsAI score42

    Credo moves into optical chips with DSP, PIC and diagnostics in one 1.6T module

    AICredo's latest full-DSP module, Cardinal 802, uses a 4×200G design aimed at both 800G and 1.6T, after the company expanded its ZeroFlap optical module line from 800G to 1.6T over the past year. The company also added Kfir200 silicon photonics PIC from its DustPhotonics acquisition and a PILOT diagnostics platform to the module.

  9. TechRadar · AINewsAI score36

    Google Playground turns plain-language prompts into playable AI-generated games

    AIGoogle's Playground experiment lets users describe a game in ordinary language and have generative AI build a playable browser-based result that can be revised through further prompts. TechRadar's reviewer built a dragon platformer, Mystic Dragon Glide, from a couple of sentences and a satirical puzzle RPG, Red Tape Hero, from a longer prompt. Playground produced working controls, objectives and music, but the reviewer found the results impressive as prototypes rather than games they would want to play for dozens of hours.

  10. X.PINXAI score36

    Tencent considers up to $5 billion offshore bond sale for AI

    AITencent is reportedly considering an offshore bond sale of up to $5 billion in US dollars and offshore yuan, possibly as early as this month, according to Bloomberg. The move follows its $4.66 billion bond offering in June, its largest debt deal since 2020, with proceeds earmarked for general corporate purposes including AI development. The post notes that major tech firms are increasingly turning to debt to fund AI infrastructure spending beyond their existing cash flows.

    Image from @thexpin's post
  11. QbitAINewsAI score67

    TRAE merges Code and Work into one platform with Agent and IDE modes

    AITRAE has merged its TraeCode and TraeWork products into a unified new TRAE with an Agent mode and an IDE mode. In hands-on tests, multiple agents handled planning, design, coding, testing, and fixes within one project, with outputs saved in a shared 'My Artifacts' area. The tests also found that agents working in parallel produced conflicting specifications, so someone had to coordinate them.

    Why it matters: The hands-on tests show how parallel agents split planning, design, coding, testing, and fixing inside one project, and where their outputs conflicted.

  12. Rest of WorldNewsAI score40

    Insta360 opens its first U.S. flagship store as DJI faces FCC drone and camera sales ban

    AIShenzhen-based Insta360 opened its first U.S. flagship store in New York City's Times Square in September, marketing its pocket-sized cameras to content creators, travelers, and sports enthusiasts. Co-founder Max Richter said the U.S. accounts for 30% to 40% of the company's revenue, while the FCC has blocked DJI from selling new drones and cameras in the country. Insta360 also launched drone brand Antigravity, whose A1 drone received FCC approval just before a new foreign drone rule took effect.

  13. Semafor · TechnologyNewsAI score44

    OpenAI reportedly tells investors it expects $70 billion annualized revenue this year

    AIOpenAI reportedly assured investors it expects to reach a $70 billion annualized revenue target this year, an apparent effort to ease concerns that it is missing earlier projections. The figures matter because they are seen as a gauge of underlying demand for cutting-edge AI models and a justification for huge spending on tech infrastructure. Some analysts, however, say the annualized revenue metric itself is flawed.

  14. Amazon Web ServicesOfficialAI score10

    Meliá Hotels International adopts AI approach, per AWS post

    AIMeliá Hotels International, a hotel chain with more than 380 hotels across four continents, is cited in an AWS post as having taken a step the post describes as the approach needed. The source gives no further details on the specific technology, implementation, or results.

  15. Amazon Web ServicesOfficialAI score33

    AWS helps migrate COBOL reservation system to microservices

    AIA company used AWS to migrate its entire COBOL-based central reservation system to microservices. The migration handles 50 million daily availability requests, up from 26 million, with faster response times.

  16. Amazon Web ServicesOfficialAI score14

    AWS customer cuts feature delivery to one month and compute costs 60%

    AIAn AWS customer reports that new features now ship in one month instead of four, with 60% compute cost savings worth seven figures. The post also cites a 75% faster time to market, near 99.99% availability, and a four-year project finished in two years.

  17. QbitAINewsAI score38

    Lenovo's TianxiCode Agent Tops SWE-bench-Live Lite Leaderboard at 71%

    AILenovo's TianxiCode, paired with DeepSeek-v4.1-Flash, ranked first on the SWE-bench-Live Lite leaderboard with a 71% issue resolution rate and passed official Verified review. The framework combines multi-hop retrieval, autonomous planning with multi-turn tool calling, and test-driven self-correction, and will be applied to Lenovo AI hardware products.

  18. The DecoderNewsAI score54

    Anthropic's Claude Science maps the full sky in ultraviolet light

    AIAnthropic's Claude Science has produced what the source describes as the first complete ultraviolet map of the sky. AI agents downloaded data from multiple space missions, calibrated and merged it, and used inpainting to fill gaps left by NASA's GALEX mission, which skipped bright star-forming regions. In tests, predictions averaged about ten percent deviation from actual measurements, and the map is intended as teaching material.

  19. The Verge · AINewsAI score58

    OpenAI defends firing three AI safety researchers over information handling

    AIOpenAI says an internal investigation found Jasmine Wang, Tomek Korbak and Mikita Balesni committed a significant breach of trust by violating policies on handling sensitive information. The company denies the dismissals were about the researchers speaking out on AI safety, responding to an open letter in which the group said it was fired for raising safety concerns. OpenAI says the investigation found breaches beyond those in the letter but has not provided details.

  20. Gemini API ChangelogOfficialAI score22

    Gemini 3.7 Flash and 3.5 Flash are deprecated and rerouted to newer models

    AIGoogle says gemini-3.7-flash is deprecated and replaced by gemini-3.8-flash, and gemini-3.5-flash is deprecated and replaced by gemini-3.6-flash, with requests to the old strings automatically routed to the new ones. Developers should update their model strings, with gemini-3.8-flash recommended for the 3.5 Flash replacement as well. Google also says the deep-research-pro-preview-12-2025 agent will shut down on October 23, 2026, and that developers should migrate to deep-research-preview-04-2026 or deep-research-max-preview-04-2026.

  21. OpenAI · YouTubeOfficialAI score36

    Sophos Cuts Threat Response Time by 96% With OpenAI Daybreak Agents

    AISophos says agents built through OpenAI Daybreak, combined with its cybersecurity expertise, cut average response time from 38 minutes to 89 seconds for cases handled by those agents. The company says the agents help its MDR team investigate threats faster and protect customers at scale while keeping human judgment central.

  22. X.PINXAI score41

    Biren Technology raises HK$4.04 billion in second share placement this year

    AIChinese AI chipmaker Biren Technology is raising HK$4.04 billion ($520 million) through a share placement of 130 million shares at HK$31.08 each, a 33% discount to its July placement price. The proceeds will mainly fund supply-chain purchases and production preparation for its next-generation BR20X chip, with about 70% going to procurement and commercialization. Its shares fell 11.85% on October 8 after the announcement.

    Image from @thexpin's post
  23. MarkTechPostNewsAI score67

    Google Cloud launches Gemini agent, a single universal agent for enterprise work

    AIGoogle Cloud has launched the Gemini agent, a cloud-hosted agent that answers questions, does knowledge work, creates media, and writes and runs code from one prompt box and one API. It orchestrates across Gemini and Anthropic Claude models today, and its project-level spend caps pause an agent when a limit is reached. The article reports no benchmarks, pricing, or general availability date yet.

  24. The Guardian · AINewsAI score62

    OpenAI projects $50bn revenue, $20bn below its earlier investor signal

    AIOpenAI told investors it expects $50bn in revenue this year, about $20bn less than the $70bn it had signalled last month. The gap stems partly from comparing with Anthropic, which counts revenue sold through cloud partners such as AWS and Google Cloud, while OpenAI does not. The news weighed on US tech stocks, and OpenAI is in early talks to raise $30bn at a valuation of about $1.4tn.

  25. The DecoderNewsAI score61

    OpenAI bans Russian and Iranian influence ops that planted fake stories in real outlets

    AIOpenAI exposed a Russian and an Iranian influence operation and banned the ChatGPT accounts involved, both of which planted content in legitimate media using fake identities. The Iranian operation, "Bogus Bylines," used seven fake journalists to place nearly 100 articles about the US-Iran conflict, while the Russian "Dark Clark" operation triggered fact-checks and official denials in Ecuador and Peru. Both operations used AI mainly for internal reporting and adapting propaganda to different languages.

  26. Tencent · new models on Hugging FaceOfficialAI score41

    Tencent Releases Youtu-Parsing-Omni, a 5B Omni-Modal Document and Media Parsing Model

    AITencent has open-sourced Youtu-Parsing-Omni, a 5B-parameter omni-modal model that outputs a single structured JSON covering layout, text, tables, formulas, ASR, OCR, and video segments. It scores 96.96 Overall on OmniDocBench, the highest among the compared models, and ships with weights on Hugging Face, a vLLM plugin, and inference examples.

  27. NVIDIA · new models on Hugging FaceOfficialAI score16

    NVIDIA releases Agile One S SSD Pick GR00T N1.7 checkpoint 40000 model on Hugging Face

    AINVIDIA published the Agile One S SSD Pick deployment model, GR00T N1.7 checkpoint 40000, on Hugging Face for SSD pickup tasks. The repository includes five ONNX graphs with external tensor files and two existing TensorRT BF16 engines, with original configurations and build metadata, but no retraining or re-export was performed. The files are not a robot deployment or safety qualification, and engine compatibility depends on the target GPU and TensorRT environment.

  28. NVIDIA · new models on Hugging FaceOfficialAI score23

    NVIDIA publishes Agile One S SSD pick model, GR00T N1.7 checkpoint 58000, on Hugging Face

    AINVIDIA has released a deployment model for Agile One S SSD pickup, based on GR00T N1.7 checkpoint 58000 and using three cameras: ego, left wrist, and right wrist. The repository republishes ONNX graphs, external tensor files, and two existing TensorRT BF16 engines without retraining or re-export, and the original export reported a numerical warning that full FP32, node, and BF16 parity did not pass all tolerances. The files are not a certified robot deployment or safety qualification.

  29. NVIDIA · new models on Hugging FaceOfficialAI score25

    NVIDIA releases Agile One S Walk GR00T N2 checkpoint 1680 on Hugging Face

    AINVIDIA published the Agile One S Walk GR00T N2 checkpoint 1680, a walking deployment model with four cameras, on Hugging Face. The repository includes ONNX graphs, TensorRT BF16 plans/engines, and the original checkpoint files, republished without retraining or re-export. The shared Cosmos-Reason1-7B dependency and the Isaac/GR00T runtime must be set up separately, and the files are not a robot safety qualification.

  30. NVIDIA · new models on Hugging FaceOfficialAI score14

    NVIDIA Releases Agile One S SSD Place GR00T N1.7 Deployment Model on Hugging Face

    AINVIDIA published the nvidia/agile_one_s_place_ssd_n17_24050 repository on Hugging Face, containing a GR00T N1.7 checkpoint 24050 model for placing an SSD with three cameras. The repository includes five ONNX graphs with external tensor files and two existing TensorRT BF16 engines, republished without retraining, re-export, or engine rebuild. Engine compatibility depends on the target GPU and TensorRT environment, and the files are not a robot deployment or safety qualification.

  31. QbitAINewsAI score62

    Google's AMIE Chatbot Tested in Real Pre-Visit Clinical Study Published in The Lancet

    AIA study led by Google and BIDMC tested Google's diagnostic AI chatbot AMIE with 98 outpatients before emergency visits, with a supervising doctor monitoring every exchange. No conversation needed interruption under the predefined safety criteria, and clinicians said AI summaries helped them prepare for 75% of visits. AMIE's differential diagnoses matched final diagnoses 90% of the time, but the authors say larger trials are needed.