REFACTOR-VLA: Unsupervised Library Learning of Typed Motor Programs
| Source: Apple ML Research
Tags: REFACTOR-VLA, Apple, robotics, VLA, motor programs, LIBERO
Apple researchers introduce REFACTOR-VLA, a robot policy system that learns reusable motor skills via a wake/sleep architecture, finding that InfoNCE contrastive loss during world-model training dramatically improves skill clustering on the LIBERO benchmark.
Details
REFACTOR-VLA addresses a core limitation of current vision-language-action models (OpenVLA, π0, RT-2, RDT-1B): they generate raw motor commands without organizing behavior into reusable abstractions, which limits performance on multi-step tasks and makes learned behaviors hard to interpret. The system uses a wake/sleep cycle. In the sleep phase, a Behavioral-Equivalence Kernel clusters motor program segments based on outcomes in a learned latent world model. In the wake phase, typed lambda terms (Hindley-Milner style) are generated from the skill library. Skills are retained only if they pass a Minimum Description Length criterion and a return-preservation gate. Two clean ablation findings stand out. First, scaling the world model from 188M to 430M parameters worsened performance on all 4 LIBERO suites — bigger is not always better for world models. Second, adding an InfoNCE contrastive loss during world-model warmup significantly improved clustering quality, with NMI rising to 0.915 on the Goal suite. Apple ML Research publishing robotics work suggests the company is investing in robot foundation models beyond on-device AI.