A continual learning loop combining SFT, GRPO and prompt compression on PyTorch and vLLM let Shopify match frontier-model quality at a fraction of the cost.