Skip to contentSkip to stories

Updated

All AI news

Items with an AI score under 20 are hidden. Show low-relevance items

Oct 7

Oct 7Wed
  1. Google Developers BlogAI score62

    Google open-sources ML Drift, a cross-platform GPU engine for on-device AI

    AIGoogle's AI Edge Team open-sourced ML Drift under Apache 2.0, a GPU compute engine for on-device AI inference across OpenGL ES, OpenCL, Metal, and WebGPU. It serves as the core GPU acceleration engine within LiteRT and succeeds the legacy TFLite GPU delegate, which will no longer receive new features. The post cites benchmarks showing up to 40% lower frame latency in YouTube Shorts and up to 30% faster on-device performance in Adobe Lightroom and Photoshop.

    Why it matters: The post explains how ML Drift unifies GPU shaders across platforms and replaces the TFLite GPU delegate, which matters for developers deploying on-device models.

  2. Apple Machine Learning ResearchAI score42

    Apple's Normalizing Trajectory Models generate images in four steps with exact likelihood

    AIApple researchers introduced Normalizing Trajectory Models (NTM), which model each reverse diffusion step as a conditional normalizing flow trained with exact likelihood. The model matches or outperforms strong image generation baselines on text-to-image benchmarks in just four sampling steps while retaining exact likelihood over the generative trajectory.

  3. Waymo BlogAI score42

    Sober Drivers Still Face Nearly 4x Nighttime Fatal Crash Risk, Waymo Study Finds

    AIWaymo research found that even fully sober human drivers face nighttime fatal crash risk 3.1 to 3.9 times higher than daytime risk, pointing to systemic hazards beyond impairment. The study used an exposure reconstruction model across the 50 most populous U.S. urban areas, showing removing alcohol-involved drivers lowers the average urban fatal crash rate by 23%, from 1.42 to 1.10 per 100 million miles.

  4. Waymo BlogAI score46

    Waymo Closes $5 Billion Debt Financing to Fund Expansion

    AIWaymo closed a $5 billion term loan, its first debt financing, with PIMCO, Blackstone, and Sixth Street as lead syndicated lenders and Goldman Sachs as sole lead bookrunner. The debt complements a $16 billion equity investment earlier this year and gives the company added financial flexibility to expand its fully autonomous ride-hailing service across the United States and internationally.

  5. Epoch AIAI score67

    Epoch tests six AI models on real Epoch work and finds they cannot yet fully automate it

    AIEpoch gave six models 11 real work tasks from its own operations, including graphic design, data insights, and research design, and graded outputs against employee standards. Fable 5.1 and GPT-6 Astra led on average task performance, reliably handling well-defined work such as coding and computational analysis. The report finds that all models still fail on open-ended judgment, including matching Epoch's standards, designing informative experiments, and generating diverse ideas, so the authors conclude AI cannot yet replace workers at Epoch.

    Why it matters: The report separates well-defined task reliability from open-ended judgment failures, which benchmark scores on easily verifiable tasks would miss.

  6. Google Developers BlogAI score62

    Google's AQuA agent diagnoses production failures in a multi-agent travel concierge

    AIGoogle Developers Blog introduces AQuA, an ambient quality agent that runs in a customer's Google Cloud project and samples production sessions to find recurring agent failures. In a 32-session travel-concierge sweep, it verified six issues and traced two of them to specific prompt lines, and a replay after the fixes raised full-session passes from 5/32 to 13/32. The post notes that verification and diagnosis are model-based, and that the tool proposes edits without applying them.

    Why it matters: The post walks through a concrete production workflow, from sweep and verification to a code-anchored fix and replay, that shows how to diagnose silent agent failures.

  7. Hugging Face BlogAI score66

    How one developer built six custom models with ML-Intern for about USD 103

    AIA Hugging Face blog author used the ML-Intern agent in HuggingChat to build six small models by writing detailed prompts that specify datasets, base models, baselines, smoke tests, and spending limits. The projects include a citrus disease vision-language model, a Huggy character LoRA, a camera-angle LoRA, a doodle-to-object LoRA, a 0.8B prompt rewriter, and a 4-step distilled Agate model, with total compute cost of about USD 103. Each project's prompts and public models are linked from the post.

    Why it matters: The author shows how prompt structure, baselines, smoke tests, and budget caps shape an agent-driven training workflow, with per-project costs given.

  8. DatabricksAI score34

    Databricks adds Workday Data Connect federation to Unity Catalog in Beta

    AIDatabricks has put Workday Data Connect federation into Beta in Unity Catalog, letting teams query Workday HR and finance data without copying it. Workday Data Cloud customers get zero-copy, read-only access to the shared tables, with Databricks running queries and Unity Catalog governing access, lineage, and auditing. Teams can combine current people and financial data with other enterprise data for analytics and AI, including Genie-powered natural-language exploration.

    Image from @databricks's post
  9. Meta NewsroomAI score28

    Meta's Head of Infrastructure Explains Why Data Centers Are Central to Its AI Strategy

    AIMeta's Head of Infrastructure, Santosh Janardhan, discusses the company's approach to building infrastructure for AI in a conversation with Tom Shaw. The discussion covers why Meta views itself as more than a software company, why AI differs from other technologies, and why data centers are essential to AI development. It also addresses power for Meta's AI infrastructure, gigawatt-scale energy needs, chip selection, and the benefits of building its own data centers.

  10. Ethan MollickAI score60

    Mathematicians react to hundreds of AI-generated proofs released by OpenAI

    AIEthan Mollick shares early first-hand accounts from mathematicians grappling with hundreds of AI proofs released by OpenAI. He highlights problems solved in ways no human has yet understood, raising questions about what it means to know something. The linked Scott Aaronson post quotes a researcher, Dana, describing the proofs as unclear and hard to read without AI help, with some possibly verified by a Lean certificate.

    Image from @emollick's post
  11. Google ResearchAI score23

    Google Research invites COLM visitors to ContinuousBench walkthrough on DP synthetic data

    AIGoogle Research is hosting a walkthrough at its COLM booth #107 today at 5:00 PM of ContinuousBench, a standardized benchmark for measuring knowledge transfer in differentially private synthetic data. The session, led by Alex Bie, asks whether DP synthetic data preserve actual information or only style. A paper is linked on arXiv.

    Image from @GoogleResearch's post
  12. IThome · AIAI score72

    Anthropic releases Claude Haiku 5.5, cutting run costs about 75% from Haiku 4.5

    AIAnthropic released Claude Haiku 5.5, which it calls the fastest, cheapest, and most capable Haiku model so far. On average it costs about 75% less to run than Haiku 4.5, with input at $0.10 and output at $0.50 per million tokens for requests up to 100,000 tokens. Anthropic also cut Sonnet 5.5's cache read price from $0.20 to $0.10 per million tokens, which it says lowers run costs by about 20% on many agent tasks.

  13. TypeSafe AIAI score25

    Jev-killer OpenAI Decisions API benchmarked against Jev for HiringCafe

    AIThe main post is a short reply saying reports of a company's death have been greatly exaggerated, with no details about products or figures. The background post from @h_nilforoshan reports that OpenAI's Decisions API, billed as a "Jev-killer," was benchmarked against Jev for HiringCafe, which serves 2.5 million users. On the task of scoring job-description relevance from 1 to 10, the author reports OpenAI costing 2x more and performing 5-10% worse.

  14. IThome · AIAI score44

    Economist Acemoglu estimates AI will automate only about 5% of jobs within 10 years

    AINobel economist Daron Acemoglu estimates AI could technically automate about 20% of work, but adoption limits will cut actual automation to roughly 5% within 10 years. He said the figure is admittedly only an estimate, noting AI models excel in lab settings but underperform when enterprises deploy them in real environments. Microsoft AI CEO Mustafa Suleyman shared the forecast on X on October 6.

  15. ThariqAI score20

    Anthropic's Thariq says game dev content creation is booming

    AIThariq says this is an incredible time to be a game dev content creator because many people make games for fun and will pay for it. Background context from Tim Sweeney notes that Fab seller revenue rose significantly in September, as AI acceleration increased content demand and developers used AI assistants to build scenes with Fab assets.

  16. DatabricksAI score36

    Claude Haiku 5.5 launches on Databricks as a Day 0 release

    AIAnthropic's Claude Haiku 5.5 is available on Databricks from day zero, which Databricks calls its cheapest, fastest, and most capable small model. On Databricks' OfficeQA Pro V1 benchmark, it delivers about 15% higher quality than Haiku 4.5 at a fraction of the cost. Users can run it alongside 60+ other models on data already in Databricks, with Unity Gateway handling governance, monitoring, and security.

    Video from @databricks's post
  17. Ars Technica · AIAI score46

    Artcraft releases open source clones of Adobe Photoshop, Premiere and other apps built with Claude

    AIDeveloper Brandon Thomas's Artcraft has launched seven open source apps in Rust that aim to replicate the interfaces and tools of Adobe Photoshop, Illustrator, Premiere, Lightroom, After Effects, InDesign, and Acrobat Pro. Thomas said he used Anthropic's Claude Opus 5.5 to build the clean-room replacements, with WebAssembly versions available for browser use. The apps remain in a "super early alpha" state, and commenters have pointed out many current shortcomings.

  18. ZDNet · AIAI score36

    Managers using AI for performance reviews draws mixed employee reactions, survey finds

    AIA Highwire survey of 1,034 corporate employees found 78% of managers used AI to help draft, edit, summarize, or inform performance reviews in the past year. Fifty-four percent of employees said feedback became more specific and actionable, while 34% found it more generic and 32% said it was less useful. Nearly one in four employees rehearse difficult workplace conversations with an AI tool, according to the survey.