Steering Generative Robot Policies with Lexicographic Preferences
| Source: arXiv AI
Tags: robotics, diffusion policy, flow matching, inference steering, deployment constraints, LIBERO
Frozen diffusion and flow-matching robot policies can be steered at inference time to respect priority-ordered deployment requirements — no weight updates, no retraining — using dynamic-barrier guidance and cascade sample filtering.
Details
Pretrained generative robot policies are increasingly capable across diverse environments, but deployment often introduces requirements that were not part of training: hardware-specific feasibility constraints, operator safety preferences, or user-specific behavioral priorities. Modifying weights for each new deployment is impractical.\n\nJia and How (MIT) show that frozen generative policies — whether diffusion-based or flow-matching-based — can be steered at inference time to satisfy lexicographically ordered objectives. Their method introduces two modifications to the sampler: dynamic-barrier guidance constrains updates so that higher-priority costs do not increase, and cascade sample selection filters candidates in priority order.\n\nOn a navigation benchmark, the method improves success rate, traversability, and preference compliance over the frozen baseline, beating tuned weighted-sum baselines on compliance. It transfers directly to a flow-matching manipulation policy on LIBERO without task success degradation.\n\nThis matters for industrial robotics deployment: the same base policy can be adapted to different site-specific constraints without retraining.