LPA-CWM: A Learned Physical Adjudicator for Motion Reasoning with Counterfactual World Models
| Source: arXiv AI
Tags: video understanding, motion tracking, counterfactual world models, computer vision, video motion estimation
LPA-CWM adds a lightweight 3M-parameter adjudicator to counterfactual world models for video motion tracking, predicting response reliability rather than using uniform aggregation—improving tracking accuracy by 60% on DAVIS and 29% on Kinetics without modifying the underlying video model.
Details
Counterfactual world models (CWMs) extract motion from video by comparing factual and intervened predictions under different target-frame masks, but responses under different masks vary in reliability. Uniform aggregation ignores this variance, treating all candidate responses equally.\n\nLPA-CWM addresses this with a Learned Physical Adjudicator (LPA): a 3M-parameter network trained on MOVi-F dense trajectories that predicts relative reliability weights for candidate responses based on visual context and response structure. The CWM predictor and intervention generator remain entirely frozen.\n\nThe paper also introduces Completeness-aware Motion Correspondence (CMC), an evaluation protocol that counts missing predictions as failures on visible dynamic points—a stricter and more honest measure than standard metrics. On DAVIS and Kinetics subsets, LPA-CWM improves DCA_avg by 60.0% and 29.0% over Uniform CWM, with additional improvements under TAP-Vid First. The adjudicator is lightweight enough to add at inference time without significant overhead.