Skip to contentSkip to stories

Updated

#Google

Oct 9

TodayOct 9Fri2 items
  1. QbitAIAI score62

    Google's AMIE Chatbot Tested in Real Pre-Visit Clinical Study Published in The Lancet

    AIA study led by Google and BIDMC tested Google's diagnostic AI chatbot AMIE with 98 outpatients before emergency visits, with a supervising doctor monitoring every exchange. No conversation needed interruption under the predefined safety criteria, and clinicians said AI summaries helped them prepare for 75% of visits. AMIE's differential diagnoses matched final diagnoses 90% of the time, but the authors say larger trials are needed.

Oct 8

Oct 8Thu
  1. Sundar PichaiAI score65

    Google's AMIE Chat System Is Tested With Real Urgent Care Patients in The Lancet

    AIGoogle published a prospective study of AMIE, a research conversational system that patients chat with before doctor appointments, in The Lancet with Beth Israel Deaconess Medical Center. Clinicians reported the summaries helped them prepare for visits in 75% of cases and influenced their approach to care in more than half. AMIE's differential diagnoses matched the doctors' final diagnoses 90% of the time.

    Why it matters: The study tests a patient-facing diagnostic chat system in a real urgent care clinic, a setting that goes beyond lab evaluation and is useful for judging clinical readiness.

  2. GoogleAI score62

    Google's AMIE Diagnostic Chatbot Studied Prospectively in Real Clinical Setting

    AIGoogle reports that its medical research system AMIE, described as the first patient-facing conversational diagnostic tool of its kind studied prospectively in a real-world clinical setting, was evaluated in a study published in The Lancet. Patients who chatted with AMIE before in-person appointments reported stronger confidence and better organized thoughts. Physicians reviewing the pre-visit conversation information gained more time for collaborative care and shared decision-making instead of digging through data.

  3. Elvis SaraviaAI score48

    Google's FlowAgent auto-repairs failing tests inside code review

    AIGoogle proposed FlowAgent, a ReAct-style agent that generates and validates fixes for pre-submit test failures and shows them in its code review tools. Two abstention filters, before and after execution, suppress weak suggestions; in a manual review of 195 real failures, 67.18% of fixes were correct. After the Google-wide launch, it suggested fixes on 295,508 changes, with developers previewing 65,069 and applying 28,554.

  4. Artificial AnalysisAI score34

    Artificial Analysis compares six hallucination checkers on 20 shared tasks

    AIArtificial Analysis compared six hallucination checkers on the same deliverables from 20 tasks across eight models. GPT-6 Sol and GPT-6 Luna generally flagged the most material hallucinations, while Claude Sonnet 5.5 and Gemini 3.8 Flash flagged far fewer, with Claude Opus 5.5 falling between Grok 4.7 and Sonnet. The counts reflect checker behavior rather than establishing accuracy or ruling out self-preference.

  5. Google ResearchAI score14

    Google Research demos EnvHarness for co-evolving LLM agents and environments at COLM 2026

    AIGoogle Research is presenting EnvHarness, a flexible framework that enables co-evolution between LLM agents and their training environments, at the #COLM2026 Google booth #107 today at 11:00 AM PT. The post notes that static environments limit agent growth, and EnvHarness is described as a plug-in architecture that dynamically reshapes environment behaviors to improve reinforcement learning and adaptability.

  6. The Next PlatformAI score43

    How Distributed AI Training Changes the Network Between Datacenters

    AILarge-scale AI training is spreading across multiple datacenters, with Google, Microsoft, AWS, Meta, and CoreWeave cited as examples. Because synchronized GPU clusters must exchange data in bursts, inter-site links can become a bottleneck, which Cisco estimates may require aggregate bandwidth about 14x a conventional DCI baseline.

  7. Artificial Analysis ArticlesAI score50

    Harvey LAB-AA v1.1 adds hallucination checks to legal AI benchmark

    AIHarvey LAB-AA v1.1 adds hallucination checks that audit every model deliverable against task source documents, with material hallucinations zeroing a task's score. GPT-6 Astra averaged 0.03 material hallucinations per task across 120 tasks, while Gemini 3.8 Flash averaged 13.96. Harvey uses GPT-6 Sol (high) as the hallucination checker, separate from its three-judge rubric panel.

Oct 7

Oct 7Wed
  1. Elvis SaraviaAI score67

    Tool-using multimodal models refuse harmful requests less often, NVIDIA study finds

    AIA NVIDIA study accepted at NeurIPS 2026 reports that multimodal models refuse harmful requests less reliably when they call tools. Refusal failures rise by up to 68.7% relative and by 17.7% on average across the models tested, including Claude Opus 4.6 and 4.7 and Gemini Agentic Vision. The authors attribute this to tool outputs crowding out the original harmful intent and to attention shifting toward describing tool results. Re-inserting the original request and image before the final response restores part of the lost refusals.

  2. Google ResearchAI score23

    Google Research invites COLM visitors to ContinuousBench walkthrough on DP synthetic data

    AIGoogle Research is hosting a walkthrough at its COLM booth #107 today at 5:00 PM of ContinuousBench, a standardized benchmark for measuring knowledge transfer in differentially private synthetic data. The session, led by Alex Bie, asks whether DP synthetic data preserve actual information or only style. A paper is linked on arXiv.

  3. Google ResearchAI score62

    Google Research finds AI boosts patent drafting but junior lawyers' gains vanish without it

    AIA Google Research field experiment with 133 patent lawyers found AI tool access raised drafting scores by 0.34 to 0.38 standard deviations over three months. When the tool was removed for a redlining task, only senior lawyers kept an advantage of 0.45 SD, while junior lawyers showed no discernible improvement. The authors argue that tools which boost current output must not stop junior professionals from building the judgment that senior experts rely on.

    Why it matters: The field experiment separates AI's short-term productivity gains from skill retained after the tool is removed, which matters for training junior professionals.

Oct 6

Oct 6Tue
  1. GoogleAI score30

    Google Earth AI helps forecast disease outbreak spread faster

    AIGoogle Earth AI, according to new research, can help communities respond to public health crises more quickly and proactively. The post says it combines behavioral trends, geospatial AI models, and other insights beyond simple statistics to help public health teams understand complex issues and bridge reporting gaps. The aim is to shift emergency response from reactive management toward proactive prevention.

  2. Google · Innovation & AIAI score42

    Google Study Tests AI-Guided Blind Sweep Ultrasounds for Pregnant Women in Kenya and Chicago

    AIGoogle researchers, working with Northwestern Medicine and Jacaranda Health, trained healthcare workers to perform "blind sweep" ultrasounds analyzed by machine learning models. The models estimated gestational age and fetal presentation as accurately as a trained sonographer in a study of 1,000 mothers each in Nairobi and Chicago. The AI processes results on the device, so it needs no electricity supply or Wi-Fi.

  3. Google ResearchAI score51

    Google's PDFM location embeddings improve five global public health tasks

    AIGoogle Research reports that Population Dynamics Foundation Model (PDFM) embeddings, built from search trends, mobility, built environment, and weather signals, were tested by partners across five public health tasks. The embeddings improved results in cross-border MMR vaccination coverage, dengue forecasting, postpartum depression screening, and cholera outbreak prediction, and matched census inputs for cardiovascular mortality nowcasting.

Oct 5

Oct 5Mon

Oct 2

Oct 2Fri
  1. Google ResearchAI score60

    Google's TEE-based federated learning system adds verifiable privacy guarantees

    AIGoogle announces a next-generation federated learning system that uses Trusted Execution Environments to provide verifiable, auditable data anonymization. The system publishes access policies to a public transparency log and is deployed in Gboard, which has launched English and Japanese next-word prediction models with stronger privacy guarantees and improved accuracy. Training time has also sped up significantly because computation moved to the server and is parallelized across many machines.

    Why it matters: The post shows how Trusted Execution Environments make federated learning's privacy claims externally verifiable, rather than relying on trust in the server operator.

Sep 30

Sep 30Wed

Sep 29

Sep 29Tue
  1. Google Developers BlogAI score47

    Google Details Sparse Attention Speedup for Video Diffusion on TPUs

    AIGoogle Developers Blog describes how Sparse VideoGen (SVG) routes video diffusion attention heads into spatial or temporal sparse masks and implements them as custom JAX and Pallas Splash Attention kernels on TPU v6e. In isolated single-chip tests with 75.6K tokens and 10 heads, the sparse variants retain about 38.87% of query-key pairs. The article argues that theoretical sparsity must be converted into hardware tile skipping to yield real speedups.

  2. Google ResearchAI score35

    Google Research unveils Diffusion Controller for steering AI image generation

    AIGoogle Research introduced Diffusion Controller, a framework that treats image generation as a continuous control problem rather than separate inference-time guidance and fine-tuning fixes. Its lightweight add-on "steering damper" network keeps the base model frozen and works on black-box or gray-box models, and it outperformed the industry standard on human preference matching. In a Stable Diffusion v1.4 test, the fully unlocked version achieved a 90% win rate over the baseline.

Sep 24

Sep 24Thu
  1. Google ResearchAI score60

    Google Research details four agentic frameworks for coherent long-form video generation

    AIGoogle Research introduces four multi-agent frameworks for generating minutes-long videos with consistent characters and environments across shots. The frameworks include AI video co-director, CANVAS, A²RD, and VQQA, which are built as orchestration layers on Gemini and Veo and use SynthID watermarking. The post reports measured gains on benchmarks such as GenAD-Bench, HardContinuityBench, and LVBench-C, with the full architectures described in the linked papers.

    Why it matters: The post links four frameworks to specific failure modes in long video generation, such as semantic drift and cascading errors, making the design choices easier to compare.

Sep 23

Sep 23Wed
  1. Google Developers BlogAI score62

    Google reproduces Olmo 3 7B pre-training in MaxText on TPUs

    AIGoogle Developers reproduced Ai2's Olmo 3 7B from scratch in MaxText on Google Cloud TPUs, covering both the stage-1 pre-training run and the stage-2 mid-training anneal. The match was checked on held-out C4 loss, an 8-task accuracy suite, multi-domain perplexity, and token-level KL, not just the training loss curve. The post also describes a data-loader bug that made training loss look better than the reference while held-out metrics did not move.

    Why it matters: The post documents how a faithful reproduction was verified on held-out metrics, including a data bug that training loss alone would have hidden.

Sep 18

Sep 18Fri
  1. Google ResearchAI score22

    Google Research releases MilleMiglia, a public middle-mile logistics benchmark

    AIGoogle Research has introduced MilleMiglia, a standardized benchmark for optimizing middle-mile logistics, the segment that moves goods across hundreds of miles overnight. The benchmark uses spatial clustering and gravity models to simulate realistic middle-mile delivery scenarios. It addresses the difficulty of optimizing these networks without public data.

Sep 17

Sep 17Thu
  1. Google ResearchAI score52

    Google Research enables teachers to create generative UI learning interactives

    AIGoogle Research is sharing an experiment that lets educators generate custom, guided STEM simulations tailored to their curriculum using generative UI. It is releasing a sample library of over 30 English interactives for physics, chemistry, biology, and math, all AI-generated and reviewed by teachers. Schools using Google Workspace for Education can sign up through the Google for Education Pilot Program to give feedback.