announcement aidev SIG 3/5

GSPO: Group Sequence Policy Optimization for stable RL scaling

Qwen published GSPO (Group Sequence Policy Optimization), a new reinforcement learning algorithm designed to address training instability in existing RL methods like GRPO. The paper identifies that existing algorithms exhibit severe instability during long training runs leading to irreversible model collapse, and proposes GSPO as a more stable alternative for scaling RL-based language model training.

PUBLISHED2025-07-27
OBSERVED2026-08-11
AGE1y
SOURCES1
  • Addresses instability issues in existing RL algorithms (e.g., GRPO) during long training runs
  • Existing methods can lead to irreversible model collapse with increased compute
  • GSPO aims to enable stable and robust training dynamics for RL scaling

COMMUNITY

No curated reactions recorded for this event. Facts and takes are kept in separate layers — community context is added by hand, never blended into the record above.