Safety as a Constraint: Fine-Tuning a LLM Recommender to Explain Itself
| Source: arXiv AI
Tags: RLHF, GRPO, recommender-systems, fine-tuning, AI-safety, LLM
Researchers at a large video streaming service fine-tuned an LLM recommender to generate personalized, faithful, and non-harmful explanations using constrained GRPO — raising the all-three-criteria pass rate from 65% to 96% without degrading recommendation quality.
Details
Recommender systems traditionally predict what users will watch next, but explaining why in a personalized, faithful, and non-harmful way requires a different approach. This paper describes a production-tested technique for adding explanation generation to an LLM-based recommender at a large video streaming service, with strict safety and faithfulness requirements.\n\nThe approach trains two LLM-judge reward models covering three specific criteria (faithfulness, non-harm, and personalization), then uses constrained GRPO to optimize explanation generation. The all-three-criteria PASS rate improves from 0.649 to 0.956 under the team's own judges and from 0.677 to 0.931 under an independent judge. Critically, using a frontier model as a drop-in generator (the baseline comparison) performs similarly to the untuned base model — confirming that fine-tuning is the right path, not just prompting a better model.\n\nThe most important finding for practitioners: fine-tuning for explanation generation leaves recommendation performance unchanged. Language capabilities are also preserved. This supports the broader thesis that single-model architectures can handle both core recommendation and user-facing explanation without capability tradeoffs — an appealing target for reducing inference cost and pipeline complexity in production.\n\nThe constrained GRPO approach is generalizable to other multi-criteria optimization problems in RLHF.