Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Sep 30

Sep 30Wed
  1. Liquid AIOfficialAI score42

    LongevityBench: Liquid AI's compact LFMs beat frontier models on aging tasks

    AILiquid AI and InSilicoMeds released LongevityBench, an aging benchmark with 17 tasks spanning clinical records, DNA methylation, transcriptomics, proteomics, and genetics. On several tasks, Liquid AI's compact LFMs outperformed every frontier model the team evaluated. The team plans to present the work to the longevity research community at ARDD this week.

    Video from @liquidai's post
  2. Google DeepMind · The KeywordOfficialAI score46

    Google DeepMind introduces SynthID Bio to watermark AI-designed proteins

    AIGoogle DeepMind has introduced SynthID Bio, a technology that embeds an imperceptible, verifiable watermark into AI-designed protein sequences and predicted 3D structures. In laboratory tests across target proteins, watermarked designs matched the performance and natural diversity of unwatermarked versions. The company says the watermark provides a provenance layer intended to strengthen biosecurity and preserve the integrity of open scientific databases.

  3. Azure BlogOfficialAI score36

    Azure Circular Centers recover value from retired hyperscale hardware

    AIMicrosoft says its Circular Centers now operate eight facilities across North America, Europe, and Asia Pacific to decide the next life of decommissioned Azure hardware. Last year, Microsoft achieved a 92% reuse and recycling rate for decommissioned servers and components. Since 2014, Azure cores per rack have increased about 13-fold while power for the same task fell roughly 90%.

  4. Baidu Inc.OfficialAI score23

    Baidu says full-stack AI integration drives value across chips, cloud, and models

    AIBaidu argues its full-stack AI architecture, spanning Kunlunxin chips, Baidu AI Cloud, ERNIE models, and applications, adds value when layers are optimized together. The post says AI-powered business reached 50% of General Business revenue in Q2 and cites Gartner's forecast that inference will account for 55% of AI-optimized IaaS spending in 2026.

  5. Aidan GomezXAI score38

    Cohere launches Embed 5 Pro and Embed 5 Fast embedding models

    AICohere has introduced Embed 5, a new family of state-of-the-art embedding models, with Embed 5 Pro for frontier capabilities and Embed 5 Fast for low-latency performance. The post says the models are extremely scalable, offer SOTA accuracy, and can be deployed privately and accessed through Model Vault.

  6. Tencent HyOfficialAI score62

    Tencent Hunyuan releases ExplorationBench to test how AI systems discover rules

    AIResearchers from Tencent Hy, Fudan University, and Tsinghua University released ExplorationBench, a benchmark that tests whether AI systems can discover hidden rules in executable Alien World sandboxes. Across 10 frontier systems, getting feedback from experiments outperformed thinking alone, with the best run reaching 89.0% after four rounds. The authors note that rankings barely transfer between the two worlds, and the code is listed as coming soon.

    Why it matters: The benchmark separates feedback-driven discovery from recall by testing systems on rules that conflict with familiar knowledge, with answers checked by an interpreter or proof checker.

    Image from @TencentHunyuan's post
  7. Cloudflare Blog · AIOfficialAI score72

    Cloudflare launches Auto Router in AI Gateway to cut AI token spend

    AICloudflare has released Auto Router in public beta through AI Gateway, where setting the model to cloudflare/auto routes each request to a model judged capable enough for the task. Internal tests showed up to 30% cost savings against frontier models, and on a 97-task internal benchmark cloudflare/auto scored 86.6% at $0.0084 per success versus 96.6% at $0.0210 for Claude Opus 5.5. The router is free during beta.

    Why it matters: The source gives a benchmark table of success rates and costs per trial, showing how routing trades quality against price for a gateway deployment.

  8. SenseTimeOfficialAI score23

    SenseTime previews Dynamic Design, animating static images with SenseNova 6.8 Flash

    AISenseTime previewed Dynamic Design, powered by SenseNova 6.8 Flash, which turns static images into animated visuals. The system decides which elements stay static and which to animate, chooses HTML/CSS, SVG, transparent images, or video for each element, and choreographs text reveals and subject motion. SenseNova 6.8 Flash is coming soon, and SenseTime is offering a limited beta.

    Video from @SenseTime_AI's post
  9. Kling AIOfficialAI score35

    Kling 4.0 full-powered version showcased in a short film demo

    AIKling AI showcased a short film generated by the full-powered KLING 4.0, which the source describes as the full-powered version. The background post from @hq4ai says the film was made with all-round reference generation, runs a native 30 seconds in 21:9 cinematic format, and offers clearer visuals and sound with more precise lip-sync than Flash. The same post states Kling 4.0 will launch in October.

  10. Hamel HusainBlogAI score42

    Hamel Husain Tests Anthropic's Claude Eval Plugin on Leasing Assistant Traces

    AIHamel Husain reviewed Anthropic's new build_eval and hill-climb commands in the claude-api plugin for Claude Code, finding it useful for discovering issues like human handoff, formatting, and voice agent problems. He criticized it for pushing evaluator creation before data review, asking for label validation in Markdown files, and bundling four failure checks into one broad call-transfer evaluator. Husain says he would hold off on using it for now.

  11. X.PINXAI score72

    DeepSeek releases Ascend versions of its core kernel toolkit

    AIDeepSeek has released an Ascend toolkit that mirrors its Nvidia components, including TileLang, DeepGEMM, DeepEP, TileKernels, FlashMLA and DeepSelect. It says every TileLang kernel used in its training now has a high-performance Ascend implementation. The post also reports that a 128-card Ascend 950 supernode, jointly optimized with Huawei, has key compute and communication tests approaching hardware limits.

    Why it matters: The release shows DeepSeek moving its core kernel stack onto Huawei Ascend, which matters for readers tracking China's alternatives to Nvidia-based AI training.

    Image from @thexpin's post
  12. EveryBlogAI score40

    Sam Altman Says OpenAI's Dot Agent Gives Him Time Back

    AIOpenAI CEO Sam Altman says Dot, the company's new always-on agent, runs his day and gives him time back, according to an interview with Dan Shipper for The Every Podcast. He also says he can't quit Astra's new Ultrafast mode and that AI will bring on a new Renaissance. The interview was recorded at OpenAI's DevDay, where the company shipped twenty-two products and features.

  13. Kling AI BlogOfficialAI score49

    Kling 4.0 Extends Native Video to 30 Seconds With Up to 10 Keyframes

    AIKling 4.0 extends native single-pass video generation from 15 to 30 seconds and adds Multiple Keyframes supporting up to 10 keyframe images, versus Start & End Frames in Kling 3.0. It also expands reference inputs to up to 15 combined assets, including up to 5 videos totaling 30 seconds, and adds 10-bit HDR at 1080p and 4K. The all-new Kling 4.0 will officially launch in October, and Kling 4.0 Flash became available to a limited group of early-access users on September 28.

  14. Artificial Analysis ArticlesOfficialAI score39

    Upstage Releases Solar Mini 4 Reasoning Model, Scoring 24 on Intelligence Index

    AIKorean AI lab Upstage has released Solar Mini 4, a proprietary reasoning model that scores 24 on the Artificial Analysis Intelligence Index with 35B total and 3B active parameters. It is priced at $0.10/$0.40 per 1M input/output tokens and has a 1M-token context window, but averages 7.1 minutes per task due to heavy output token use. Its weights are not released, and its size cannot be independently verified.

  15. Artificial Analysis ArticlesOfficialAI score75

    Gemini 4 Argon matches GPT-6 Astra on intelligence index at lower cost

    AIArtificial Analysis reports that Google's Gemini 4 Argon scores 53 on its Intelligence Index with high reasoning, matching GPT-6 Astra (max) and one point ahead of GPT-6.1 Sol (max). At the current 50% launch discount, its cost per task is $1.99, about 60% of GPT-6 Astra's $3.26, but the discount's end date is unconfirmed and standard pricing would raise it to $3.98. The model is being rolled out to selected users and is not publicly available.

    Why it matters: The benchmark compares Gemini 4 Argon's cost per task and hallucination rate with GPT-6 Astra, showing where its value depends on a temporary 50% discount.

  16. Kling AI BlogOfficialAI score62

    Kling 4.0 enters early access with 30-second native video generation

    AIKling 4.0 is entering early access, with a wider rollout planned for October, and Kling 4.0 Flash opens to Ultra Yearly subscribers on September 28. The update generates videos up to 30 seconds in a single pass, accepts up to 15 reference assets, and supports up to 10 keyframe images. Upcoming features include 10-bit HDR output at 4K and 1080p and video extension up to 2 minutes.

    Why it matters: The post specifies concrete capability limits such as 30-second native generation, up to 15 references, and 10 keyframes, which help users judge fit for production workflows.

  17. Anthropic ResearchOfficialAI score62

    Anthropic study finds robots can do most physical tasks but rarely cost-effectively

    AIAnthropic's research rates how well present-day robots can perform US job tasks, finding they can do 74% of physical tasks, or 34% of working hours, mostly in limited settings. Robots are cost-competitive for only 0.3% of job tasks, and at a 3% annual price decline it would take about 40 years to reach 10%. The report also finds robot-exposed jobs tend to pay less and be more physically demanding than LLM-exposed jobs.

    Why it matters: The report separates current robot capability from cost, showing that physical automation is technically broad but economically narrow for now.

Sep 29

Sep 29Tue
  1. TechNode · AINewsAI score54

    ByteDance's Doubao reportedly preparing personal AI agent codenamed Spell

    AIByteDance's Doubao is reportedly accelerating work on a personal AI agent codenamed Spell, which entered small-scale internal testing in April. According to Sina Tech, the project is being combined with core capabilities from Doubao's conversational AI team, with a public launch expected in the near future. The report places it alongside Doubao Work, an enterprise agent launched August 25, as a sign Doubao is pursuing parallel enterprise and consumer agent tracks.

  2. Jerry LiuXAI score23

    Jerry Liu says OpenAI's Dots feels like a ChatGPT feature

    AIJerry Liu wishes OpenAI's Dots were a standalone app rather than part of ChatGPT, since it feels like a product feature similar to GPTs rather than a major release. He suggests OpenAI keeps it inside ChatGPT to protect the brand, which is its canonical application-layer product. He argues that focused rivals launching dedicated agent products, such as Instinct or Muse, may gain an advantage in mindshare, reflecting an innovator's dilemma.

  3. Sundar PichaiXAI score50

    Google's Pichai signs White House Accord on Super Intelligence with US leaders

    AISundar Pichai said Google signed the White House Accord on Super Intelligence after a meeting with President Trump, Vice President Vance, Speaker Johnson, and administration and tech leaders. He said Google has invested hundreds of billions of dollars over the past two years and will commit more, and that it will release models or products only after thorough review, testing, and safeguards against misuse and misalignment.

    Image from @sundarpichai's post
  4. PlatformerBlogAI score49

    OpenAI's Dots agent is a paid, messaging-based assistant for ChatGPT users

    AIOpenAI has launched Dots, an AI agent that a Platformer columnist tested and found highly capable and focused on work tasks. Dots is available only to paid ChatGPT users for now, and it runs as a running chat inside ChatGPT. The columnist used it to decline a radio appearance, draft a company vacation policy email to a lawyer, and answer bookkeeper questions.

  5. Allie K. MillerXAI score62

    OpenAI launches Dots, a proactive always-on agent with a dedicated VM per Dot

    AIOpenAI launched Dots, and the author argues its always-on design and dedicated virtual machine for each Dot make it feel more like a persistent teammate. The post says the product is currently limited to one primary Dot, with a team of Dots promised later, and that early reviewers report bugs the author expects to be fixed over the next few weeks.

  6. ChatGPTOfficialAI score22

    ChatGPT adds shareable profiles for Sites and plugins

    AIChatGPT now lets users publish shareable profiles that bring their Sites and plugins together in one place for others to find and reuse. Teammates can discover shared skills within their workspace, and the feature is available to Free, Go, Plus, Pro, and Business users. ChatGPT Enterprise, Edu, and Healthcare plans will get it soon.

    Image from @ChatGPT's post
  7. Jerry LiuXAI score22

    GPT-6.1 Sol Improves Table Parsing and Reading Order in OCR Benchmarks

    AIJerry Liu benchmarked gpt-6.1 sol on document OCR tasks and found a sizable increase in table parsing and reading order over gpt-6 sol from a week earlier. Its table parsing is similar to gpt-6 astra. He noted frontier models still cost roughly an order of magnitude more than cost-effective document parsing solutions, leaving room to improve the premium end above 1c per page.

    Image from @jerryjliu0's post
  8. Factory NewsOfficialAI score42

    Factory Launches Generally Available Automations to Run Recurring Engineering Workflows

    AIFactory's Automations, now generally available, let users describe a recurring workflow, set a schedule or event trigger, and have its Droid run it, with the model chosen per task. Templates cover ticket-to-PR, code review, security audits, PR babysitting, and morning Slack briefs. Among enterprise organizations using Automations in the past 30 days, 52% used automated code review, 48% used security review, and 35% used AutoWiki.

  9. Hugging Face BlogOfficialAI score46

    Open TTS Leaderboard ranks multilingual and voice cloning models using objective metrics

    AIHugging Face released the Open TTS Leaderboard, which evaluates open-source text-to-speech models using objective metrics instead of arena-style human votes. It measures intelligibility via WER and CER using Qwen3 ASR, speed via RTFx and time-to-first-audio on an H200 GPU, and speaker similarity via WavLM embeddings. The leaderboard covers multilingual results and voice cloning, and it is intended to complement, not replace, human preference rankings.

  10. Fireworks AI BlogOfficialAI score51

    Fireworks explains how numerical mismatch and MoE routing can derail RL training

    AINumerical differences between a rollout engine and a trainer can make reinforcement learning collapse even when algorithm and data stay identical. In a GLM 5.2 experiment, reward fell from about 0.9 to under 0.2 around step 20 without alignment, while aligned numerics kept reward stable over 25 steps. A Qwen3.5-MoE investigation traced a significant mismatch to how expert outputs were combined, and router replay alone was judged insufficient.

  11. Apple Machine Learning ResearchOfficialAI score38

    LLM Conditioning Study Finds Steering Methods Trade Fluency for Effectiveness

    AIApple researchers systematically tested LLM conditioning methods and found efficient activation steering often degrades fluency. Steering is far less effective on instruction-tuned models than base models, while prompting and full supervised fine-tuning work for concept injection but are weaker at concept removal. Cheap textual metrics correlate highly with costly LLM-as-judge scores.

  12. PromptArmor Threat IntelligenceOfficialAI score54

    Malicious Copilot Cowork skill hijacked AI gateway to exfiltrate files

    AIPromptArmor disclosed that a malicious Skill could hijack Copilot Cowork's AI gateway to spawn cloud agents that exfiltrate a victim's files to an attacker's server. No human approval was required, and any data Copilot could access was exposed. The vulnerability was reported to Microsoft on July 14, 2026, and Microsoft confirmed a fix on September 2, 2026.

  13. Google Developers BlogOfficialAI score47

    Google Details Sparse Attention Speedup for Video Diffusion on TPUs

    AIGoogle Developers Blog describes how Sparse VideoGen (SVG) routes video diffusion attention heads into spatial or temporal sparse masks and implements them as custom JAX and Pallas Splash Attention kernels on TPU v6e. In isolated single-chip tests with 75.6K tokens and 10 heads, the sparse variants retain about 38.87% of query-key pairs. The article argues that theoretical sparsity must be converted into hardware tile skipping to yield real speedups.

  14. v0OfficialAI score42

    GPT-6.1 Sol now available in v0

    AIGPT-6.1 Sol is now live in v0, with access via the v0 app link provided in the post. The quoted Vercel post says it is also on AI Gateway and improves on GPT-6 Sol for coding, computer use, multi-step workflows, and complex document analysis.

  15. Prime IntellectOfficialAI score20

    Prime Intellect to deploy on NVIDIA Vera CPU for agentic workloads

    AIPrime Intellect says it will be among the first to deploy on NVIDIA's new Vera CPU, after earlier access to benchmark sandboxes on Vera in March. The company plans to use Vera's dynamic memory latency for agentic workloads, which it describes as ideal for Prime Sandboxes.

    Image from @PrimeIntellect's post