LoRA, a Storage Bucket, and a Proxy: How HuggingFace Jobs Got Async GRPO Working Across Machines
HuggingFace shows how to split RL training and inference across separate Jobs with no shared NCCL—just a rank-1 adapter, a FUSE-mounted bucket, and smart routing. Training time drops from 3h27m to 53m.