LePlanner: An Iterative Amortized Controller For World Models

| Source: arXiv AI

Tags: robotics, world-models, LePlanner, planning, reinforcement-learning, manipulation

LePlanner, an amortized iterative controller for world model-based robot planning, matches search planners (CEM, MPPI) at 98% success on PushT and 100% on Reacher while requiring 3-49x less wall-clock time per decision—bringing fast goal-aware robot control closer to deployment.

Details

Planning in latent world model spaces splits between two approaches: search-based methods (CEM, MPPI, iCEM) that find good action sequences through many rollouts—high quality but slow—and policy-based methods that amortize inference into a single forward pass but struggle on contact-rich, multimodal tasks. LePlanner occupies a middle ground: an iterative amortized controller that refines latent action sequences through a frozen world-model predictor. Two training objectives define the approach. The arrival-and-hold objective encourages the controller to reach the goal at the earliest feasible horizon and stay there, addressing 'horizon-reset procrastination'—where repeated receding-horizon replanning continually postpones goal arrival rather than committing to a solution. An action-Gaussian loss keeps generated actions near the offline dataset's support to prevent out-of-distribution actions. Results across four environments are strong: 98% on PushT, 100% on Reacher, 100% on TwoRooms, and 92% on OGBench Cube (contact-rich manipulation). The computational advantage is 3-49x fewer predictor evaluations and correspondingly lower wall-clock time versus CEM-class planners. Working with a frozen world model means LePlanner is compatible with existing pretrained predictors without retraining them.