Training a coding model to paint watercolours with TRL and OpenEnv

| Source: Hugging Face Blog

Tags: GRPO, reinforcement learning, Qwen, TRL, HuggingFace, code generation, creative AI, LoRA

Hugging Face engineer Sergio Paniego open-sources the full GRPO reinforcement learning pipeline that trains Qwen3.5-35B with LoRA to generate p5.brush JavaScript watercolor paintings—reproducing a viral video (1.5M views) with every artifact public: training scripts, RL environment, scorer model, and trained weights on the Hub.

Details

In late August 2026, a viral video (1.5M views at posting) showed a language model generating watercolor paintings by writing JavaScript through the p5.brush library. The original project by Surya Narreddi demonstrated the concept but lacked open artifacts. Hugging Face engineer Sergio Paniego reproduced the full training pipeline and published everything openly. The technical approach uses GRPO (Group Relative Policy Optimization) reinforcement learning via TRL (Transformer Reinforcement Learning library). The model—Qwen3.5-35B-A3B with LoRA applied to all linear layers—generates JavaScript code that produces watercolor-style paintings using p5.brush. An RL environment hosted as a HuggingFace Space runs the generated code and a scorer model evaluates aesthetic quality. Three reward mix variants were trained and compared in parallel, providing ablation data on which scoring signals matter for artistic quality. The full pipeline runs end-to-end on HuggingFace infrastructure: training via HF Jobs on H200 hardware, with the environment and scorer as Spaces and all models on the Hub. A single command launches training. The post is notable for full reproducibility: the hand-rated reference pool dataset, RL environment, scorer model, and trained weights are all public. It also demonstrates GRPO as a general-purpose technique for creative code generation, not just math and reasoning tasks.