# Qwen publishes QwQ-32B with RL scaling research

> Qwen released a research blog post on QwQ-32B, a model whose development centered on scaling reinforcement learning rather than conventional pretraining and post-training alone. The post discusses RL as a path to improved reasoning, cites DeepSeek R1's cold-start data and multi-stage training approach, and states the work explores RL scalability. Platform badges in the source point to availability on Qwen Chat, Hugging Face, and ModelScope, but the raw text contains no benchmark results, model architecture details, or capability claims.

| | |
|---|---|
| **Tool** | Qwen |
| **Version** | — |
| **Kind** | announcement |
| **Published** | 2025-03-05 |
| **Observed** | 2026-08-11 |
| **Significance** | 3/5 |
| **Breaking** | no |
| **Categories** | new-model, capability |

> **Note:** this classification is below our confidence threshold and is pending human review.

## What changed


- Research blog post exploring RL scalability for improving LLM intelligence
- Cites DeepSeek R1 as prior work using cold-start data and multi-stage training
- Platform badges indicate availability on Qwen Chat, Hugging Face, and ModelScope
- No benchmark numbers, architecture details, or concrete capability claims in the source text


## Sources

- [blog_rss](https://qwenlm.github.io/blog/qwq-32b/) — retrieved 2026-08-11


## Community

_No curated reactions recorded. Facts and community takes are kept in separate layers
and never blended._

---
Canonical: https://changelogs.info/qwen/qwen-publishes-qwq-32b-with-rl-scaling-research
Entity: https://changelogs.info/qwen
Event ID: `evt_2025-03-05_qwen_qwq-32b-embracing-the-power-of-reinforcement-learning`
Licence: event synthesis © changelogs.info, CC BY 4.0. Linked sources belong to their vendors.
