# GSPO: Group Sequence Policy Optimization for stable RL scaling

> Qwen published GSPO (Group Sequence Policy Optimization), a new reinforcement learning algorithm designed to address training instability in existing RL methods like GRPO. The paper identifies that existing algorithms exhibit severe instability during long training runs leading to irreversible model collapse, and proposes GSPO as a more stable alternative for scaling RL-based language model training.

| | |
|---|---|
| **Tool** | Qwen |
| **Version** | — |
| **Kind** | announcement |
| **Published** | 2025-07-27 |
| **Observed** | 2026-08-11 |
| **Significance** | 3/5 |
| **Breaking** | no |
| **Categories** | feature, capability |

## What changed


- Addresses instability issues in existing RL algorithms (e.g., GRPO) during long training runs
- Existing methods can lead to irreversible model collapse with increased compute
- GSPO aims to enable stable and robust training dynamics for RL scaling


## Sources

- [blog_rss](https://qwenlm.github.io/blog/gspo/) — retrieved 2026-08-11


## Community

_No curated reactions recorded. Facts and community takes are kept in separate layers
and never blended._

---
Canonical: https://changelogs.info/qwen/gspo-group-sequence-policy-optimization-for-stable-rl-scaling
Entity: https://changelogs.info/qwen
Event ID: `evt_2025-07-27_qwen_gspo-towards-scalable-reinforcement-learning-for-language-models`
Licence: event synthesis © changelogs.info, CC BY 4.0. Linked sources belong to their vendors.
