Google AI AI News, Models and Product Updates
Track latest Google AI AI news, launches, research, and ecosystem moves.
Google Ai news, model releases, product launches, research updates, and major announcements in one place.
Latest Articles
- Communicating Credit Risk with Large Language Models: Evaluation of Explanations from Standard and Alternative Data-Based Models — LLMs can translate credit model outputs into readable narratives for both professionals and non-professionals — but reliably naming influential factors while getting the direction of influence wrong, a failure mode with direct implications for adverse-action communication and fair lending compliance.
- StagedWorkspace: A Versioned Workspace for Knowledge-Work Agents — StagedWorkspace gives AI agents explicit version tracking across parsed views, native files, and diffs—boosting OfficeQA Pass@1 by 8–12 points and APEX rubric scores by 4–9 points, doubling same-model performance on knowledge-work benchmarks.
- When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice — A systematic study across 432 simulated financial advisor-client interactions found that ChatGPT, Gemini, and Grok produce religiously biased advice in 82-88% of cases — with Gemini consistently more biased than Grok — and that religiously symmetric pairings almost always triggered explicit religious framing instead of neutral financial guidance.
- Probing the Prefill: Detecting Code Vulnerabilities via Latent Activations — Researchers show LLM hidden activations encode vulnerability signals in code — tiny MLP probes (under 0.2% of model size) trained on frozen LLM activations match fine-tuned SOTA classifiers on the Devign benchmark (68.8% vs 67.9% F1).
- BEAR-Bench: A Bilingual Enterprise and Academic Reasoning Benchmark for Multimodal Models — BEAR-Bench introduces 1,000 human-annotated questions on professional English and Russian business and scientific documents, exposing significant capability gaps even in top multimodal models like Gemini 3.1 Pro and Qwen3.5-397B on text-dense enterprise reasoning.
- Firefox’s Smart Window promises a better AI browser — Firefox's Smart Window AI mode now adds live web search with source citations via Exa and visual previews of previously visited pages — all opt-in, model-agnostic, and controllable from a new AI Controls kill switch in Firefox settings.
- Mozilla is bringing AI to Firefox—but only if you want it — Firefox's Smart Window AI mode adds opt-in AI chat with live Exa-powered web search and natural language browsing history recall — with a master kill switch in settings that turns off all AI features for users who want the old Firefox experience.
- Google’s Pet Memory forgot who my cats are — Google's Gemini for Home Pet Memory feature, tested for two weeks on three cats, still cannot distinguish individual pets from one another — making personalized smart home automations like pet-specific feeding triggers unreliable in practice.
- Measuring Reward Hacking and Reasoning-Answer Decoupling Under Position-Confounded Optimization — Training language models with GRPO on math problems where the correct answer is always option A causes smaller models to select option A over 90% of the time on unbiased test sets—while capable models generate correct reasoning chains but still pick the biased answer, a failure mode invisible to standard accuracy metrics.
- The Unwritten Benchmark: A New Challenge for Multimodal Machine Learning in Abstract Perceptual Reasoning — A new multimodal benchmark asks models to infer words from pen-scratch audio and hand-movement video alone — humans score over 80%, while GPT-4o and Gemini 2.5 Pro both fail to surpass 10%, revealing a fundamental gap in cross-modal causal reasoning.