Async GRPO with LoRA across HF Jobs: a bucket, a proxy, and no NCCL
| Source: Hugging Face Blog
Tags: TRL, LoRA, GRPO, vLLM, Hugging Face, RLHF, AsyncGRPO
TRL v1.14's AsyncGRPOTrainer now supports LoRA adapters, enabling distributed GRPO training across separate HF Jobs without NCCL — a rank-1 adapter syncs via shared storage bucket instead of full model weights, cutting a 500-step run from 3h 27min to 53min.
Details
TRL v1.14 ships LoRA support in AsyncGRPOTrainer (PR #7017), allowing the trainer to fine-tune a LoRA adapter instead of the full model and sync only that adapter to vLLM inference workers. This decouples training from inference at the infrastructure level: the two can now run as separate Hugging Face Jobs on separate VMs with no NCCL group between them. The key insight is size. A rank-1 LoRA adapter for a 1.5B model is a few megabytes; the full model is ~3 GB. That gap makes a storage bucket a viable sync channel where full-weight NCCL would be impractical across cloud VMs. Research from Thinking Machines ('LoRA Without Regret') provides the theoretical backing: the policy-gradient advantage function carries only ~O(1) bits of information per episode, so a rank-1 adapter has enough capacity to absorb every update without the quality loss you'd expect from such extreme compression. The system architecture places a small proxy in front of the vLLM replicas. It injects auth headers, routes each rollout to the replica that already holds the relevant KV prefix (reducing redundant prefill), and broadcasts adapter updates so all replicas stay in sync. vLLM can hold multiple adapters simultaneously, meaning in-flight rollouts finish with the policy they started under — no mid-run consistency issues. Real-world benchmarks on five runs show the same recipe going from 3 h 27 min to 53 min for 500 steps — a 3.9× speedup. The feature ships in TRL v1.14 and is ready to use today.