HarnessBandit: Joint Learnability-Transferability Scheduling for Multi-Harness Agentic Reinforcement Learning
| Source: arXiv AI
Tags: reinforcement learning, agentic AI, HarnessBandit, Qwen, GRPO, multi-harness training
HarnessBandit improves multi-harness agentic RL training by using an online bandit scheduler that balances per-step learnability and cross-harness transferability, outperforming mixed-batch training of Qwen3.5-2B on held-out tasks and harnesses.
Details
As language model agents are deployed across varied interfaces — different system prompts, tool schemas, control loops, and trajectory formats — a single trained policy must generalize across these harness variations. Training on a mix of harnesses simultaneously helps, but introduces a scheduling question: which harness should each gradient step use? HarnessBandit addresses this with an online bandit scheduler. After each GRPO (Group-Relative Policy Optimization) update, it observes two signals: learnability (mean absolute advantage on the batch, indicating how much the current harness is teaching the model) and transferability (cosine similarity between a low-dimensional gradient sketch of the current harness and EMAs of the other harnesses, indicating whether this update benefits all harnesses). Both signals are fused after normalization, with an exploration bonus to prevent harness starvation. Training Qwen3.5-2B across six harnesses on ClawGym and evaluating on PinchBench (held-out tasks, in-distribution harness) and ClawEval (held-out tasks and harness), HarnessBandit outperforms mixed-batch multi-harness training on both. Diagnostic analysis confirms learnability and transferability carry distinct, evolving signals throughout training. The work is from Harbin Institute of Technology and Alibaba Cloud, suggesting practical motivation from large-scale deployment. The framework is model-agnostic and applies to any RL training setup with multiple harness variants.