AI News from Google Research Blog
Latest coverage from Google Research Blog, summarized and scored for signal.
- ToolGrad: Efficient tool-use dataset generation with textual "gradients" — Google Research's ToolGrad inverts the standard tool-use dataset pipeline — generating tool-call chains first, then deriving user queries — cutting annotation to a single LLM step instead of a full search loop. Models trained on ToolGrad data outperform baseline methods and match state-of-the-art proprietary LLMs on out-of-distribution benchmarks with unseen tools.
- Transfer learning for genomic prediction in underrepresented populations — Google Research finds that transfer learning from European genome databases (UK Biobank) improves polygenic risk score accuracy in underrepresented Japanese populations only when target cohort size is small—once local samples grow, European-derived models degrade performance, especially for population-specific traits.
- A connectomics milestone: Mapping the complete male fruit fly brain — Google Research and HHMI Janelia published the complete wiring map of the male fruit fly brain and central nervous system in Cell — 166,000 neurons and 125 million synaptic connections, the largest neural map by neuron count ever assembled, built over a decade with AI-assisted reconstruction.
- Mapping global methane emissions from space with deep learning — Google Research's MAPL-EMIT uses deep learning on NASA's EMIT hyperspectral satellite to automatically detect and quantify methane plumes globally, achieving 84% recall on expert-annotated data — turning satellite imagery into actionable emissions monitoring for the Global Methane Pledge's 30% cut target.
- TimesFM-3: A zero-shot foundation model for multivariate forecasting — Google Research released TimesFM-3, a 330M-parameter time series foundation model that extends zero-shot forecasting to full multivariate scenarios — the first in the TimesFM family to handle multiple correlated series, past covariates, and known future events simultaneously without task-specific fine-tuning.
- Planetary prediction engine: Automating global models via Earth AI — Google Research's Planetary Prediction Engine (PPE) autonomously builds geospatial prediction models from natural-language queries — cutting what previously required weeks of expert data curation to minutes, with improvements demonstrated across public health, food security, and environmental risk benchmarks.
- GlucoFM: Foundation model for continuous glucose monitoring — Google Research's GlucoFM is a self-supervised dual-stream foundation model for continuous glucose monitoring that outperforms prior CGM models by 5.8 percentage points PR-AUC, enabling diabetes risk assessment, insulin resistance detection, and glycemic forecasting from wearable data with minimal labeled samples.
- AgentHands: Generating interactive hand gestures for spatially grounded agent conversations in XR — Google Research's AgentHands prototype, published at CHI 2026, gives XR AI agents synchronized, expressive hand gestures that spatially anchor verbal instructions in 3D space—pointing, tracing, and grip-mimicking to make mixed-reality task guidance more intuitive than flat 2D bounding boxes.
- An AI tool for prioritizing candidate biomarkers from wearable sensor data — Google Research's Biomarker Discovery Framework uses six coordinated AI agents with adversarial validation to turn wearable sensor streams into clinically meaningful biomarkers, recovering known clinical signals across 9,279 participants.
- How mobility gives language models a deeper understanding of place — Google Research introduces ME-POIs, a framework enriching LLM place representations with anonymized mobility patterns — arrival times, stay durations, movement flows — achieving up to 81.9% relative gain in visit intent prediction, 75.1% improvement in price level classification, and 24.7% better busyness estimation.
- Seeing beyond BMI: Estimating cardiometabolic risk with smartphone imagery — Google's PhotoScan deep learning model estimates body composition and predicts insulin resistance from standard smartphone photos, achieving accuracy comparable to expensive DXA clinical scans — potentially enabling metabolic risk screening without clinical infrastructure.
- Empty shelves or lost keys? Recall is the bottleneck for parametric factuality — Google Research's knowledge profiling framework finds frontier LLMs (Gemini 3, GPT-5) encode nearly all facts correctly but fail to recall many of them — most factual errors are retrieval failures, not knowledge gaps, shifting the optimal fix toward post-training and inference-time methods rather than more pretraining data.
- Advancing AMIE towards expert-level audio-visual clinical consultations — Google's AMIE medical AI now conducts real-time video clinical consultations, achieving expert-level performance in a first-of-its-kind randomized controlled study—interpreting visual cues like patient gait and breathing that text-based systems inherently discard.
- Science One Framework: A verifiable autonomous research framework via Chain-of-Evidence — Google Research's Science One Framework introduces Chain-of-Evidence (CoE), a verifiability standard for autonomous AI research that eliminates hallucinated references entirely and produces fully reproducible experimental scores, while achieving state-of-the-art on MLE-Bench and Parameter-Golf.
- SymptomAI: Towards a conversational AI agent for everyday symptom assessment — Google Research's SymptomAI used Gemini Flash 2.0 to conduct end-to-end symptom interviews for 13,917 real patients in a randomized national study, with diagnostic accuracy validated against clinician outcomes two weeks later — the first AI symptom-checker evaluated in true real-world conditions rather than synthetic case studies.
- Towards a quantum computer that learns from its errors — Google Quantum AI published in Nature a reinforcement learning framework that lets a quantum computer continuously recalibrate its control parameters mid-computation — eliminating mandatory calibration halts that have capped quantum algorithm runtimes to minutes.
- Towards demystifying the creativity of diffusion models — Google Research's ICLR 2026 paper proves that diffusion models' ability to generate novel images is a mathematical consequence of neural networks learning a 'smoothed' score function — forcing interpolation between training examples along the data manifold rather than memorizing any single one.
- SensorFM: Towards a general intelligence and interface for wearable health data — Google Research released SensorFM, a foundation model pre-trained on over one trillion minutes of wearable sensor data from 5 million people — the largest wearable health dataset ever used — achieving state-of-the-art transfer to 35 health prediction tasks including cardiovascular, sleep, and mental health.
- The power of collaboration: How we can reduce traffic congestion — A six-month switchback experiment across 10 US cities shows that coordinating a small fraction of Google Maps users onto alternative routes reduces overall network congestion for all drivers — published in Nature Cities.
- Expanding our Heat Resilience data to 50+ global cities — Google Research expanded its AI-driven heat resilience dataset from 14 to 50+ global cities, mapping building-level rooftop reflectivity using Sentinel-2 fused with 30cm Airbus imagery — published in Nature Communications and accessible via a new Earth Engine App.
- Introducing TabFM: A zero-shot foundation model for tabular data — Google launches TabFM, a zero-shot foundation model for tabular classification and regression integrated directly into BigQuery ML, eliminating per-dataset model training and feature engineering with a single forward pass.
- Accelerating Gemini Nano models on Pixel with frozen Multi-Token Prediction — Google retrofitted Multi-Token Prediction onto frozen Gemini Nano v3 models for Pixel 9 and 10 using a late-exit strategy, generating multiple tokens per forward pass without a separate drafter model — speeding up AI Notification Summaries and Proofread while cutting energy use.
- Optimizing cloud economics with linear elastic caching — Google Research published a CIDR paper on linear elastic caching—a technique that frames cache eviction as a ski rental problem, using lightweight ML to dynamically resize cache allocation, targeting serverless environments where memory costs up to $3/day per GiB.
- Thinking to recall: How reasoning unlocks parametric knowledge in LLMs — Google Research finds chain-of-thought reasoning unlocks correct factual recall in LLMs even for simple single-hop questions — operating via two mechanisms: a computational buffer effect where intermediate tokens do latent work, and factual priming where related facts trigger the target answer.
- From pixels to planning: Earth AI for nature restoration — Google Research releases Vectorized Farmscapes 2020 — a dataset converting high-resolution satellite maps into actionable vector inventories of England's hedgerows, copses, and stone walls, enabling carbon accounting for fine-scale vegetation features that are too small for standard satellite detection.