ReWeight: Leveraging Human Data for VLA Post-Training via Demonstration Retrieval and Sample Weighting

| Source: arXiv AI

Tags: robotics, VLA, human-demonstrations, embodiment, post-training, optimal-transport

ReWeight boosts VLA robot policy performance by selectively incorporating egocentric human demonstrations via optimal transport retrieval and cross-embodiment sample weighting, improving π0.5's real-world success rate to 68.8%—a 28.8 percentage point gain over robot-only training.

Details

VLA (vision-language-action) models like π0.5 need in-domain robot demonstrations for post-training, but collecting diverse robot data is expensive. Human egocentric video is abundant but naively mixing human and robot data degrades performance—different viewpoints, kinematics, and action distributions create cross-embodiment discrepancies that confuse the policy. ReWeight solves this with a two-stage approach: first, a cross-embodiment visuomotor representation is learned that combines visual observations with future actions to measure behavioral similarity between human and robot demonstrations. Then, using optimal transport, the system retrieves human demonstrations most relevant to the target robot task and assigns higher weights to samples with smaller cross-embodiment discrepancies. Across eight simulation tasks and four real-world tasks, ReWeight improves π0.5: from 39% success (robot-only) and 44% (random mixing) to 57% in simulation. In real-world settings, the gains are larger—68.8% versus baselines of 40% and 55% (improvements of 28.8 and 13.8 percentage points). The method is directly applicable to robotics labs that have human demonstration datasets but face the high cost of collecting diverse robot data.