AgentHands: Generating interactive hand gestures for spatially grounded agent conversations in XR
| Source: Google Research Blog
Tags: Google Research, XR, hand gestures, HCI, Android XR, CHI 2026, multimodal AI
Google Research's AgentHands prototype, published at CHI 2026, gives XR AI agents synchronized, expressive hand gestures that spatially anchor verbal instructions in 3D space—pointing, tracing, and grip-mimicking to make mixed-reality task guidance more intuitive than flat 2D bounding boxes.
Details
AgentHands is a research prototype from Google's XR team that extends conversational AI agents with co-speech hand gestures in extended reality environments. The work addresses a gap in systems like Project Astra and Gemini 3.1 Flash Live: visual bounding boxes work well on 2D screens, but transitioning to immersive XR platforms requires a more embodied form of spatial grounding. The team conducted a formative study with XR and HCI experts at Google to build a multi-dimensional gesture taxonomy. It covers handedness (single vs. both hands), spatiality (mid-air for general conversation, object-anchored for specific parts, or user-relative for social cues), and temporal dynamics with visual effects—animated motions like pouring or tracing, plus XR-specific effects such as red glows to signify heat warnings. An LLM orchestrates gesture selection and timing synchronized with the agent's speech. Published at CHI 2026, AgentHands builds on Google's earlier Human I/O and Sensible Agent work, suggesting a sustained research program around embodied agents. Demonstrated applications include step-by-step physical task guidance where the agent's hand points to specific object parts, traces paths, or mimics tool-use motions—reducing the cognitive load of translating verbal instructions into spatial actions. This is research-stage work with no product timeline announced. Its main significance is defining an interaction vocabulary for AI agents in XR, particularly relevant to Android XR devices Google is actively developing.