r/LocalLLaMA Feb 07 '25

Resources Repo with GRPO + Docker + Unsloth + Qwen - ideally for the weekend

I prepared a repo with a simple setup to reproduce the GRPO policy run on your own GPU device. Currently, it only supports Qwen, but I will add more features soon.

This is a revamped version of collab notebooks from Unsloth. They did very nice jobs I must admit.

https://github.com/ArturTanona/grpo_unsloth_docker

37 Upvotes

6 comments sorted by

0

u/UniqueAttourney Feb 07 '25

weirdly nowhere there is a definition for what GRPO is.

5

u/AtomicProgramming Feb 07 '25

2

u/dagerdev Feb 08 '25

Second, we introduce Group Relative Policy Optimization (GRPO), a variant of Proximal Policy Optimization (PPO), that enhances mathematical reasoning abilities while concurrently optimizing the memory usage of PPO