100 GRPO steps on a free Colab GPU lift a 350M model's structured output score by 7 points
Hugging Face fine-tuned LiquidAI/LFM2.5-350M with GRPO in TRL on ~500 samples and three reward functions: IFStruct rose from 22.6% to 29.7% overall and from 18.0% to 31.9% on JSON, on a 16 GB free-tier Colab or Kaggle GPU with LoRA r=16.