原題 JAJapanese title
GRPOはなぜ長時間学習で崩壊するのか――Qwenが出した「系列単位」の答え、GSPOGRPOはなぜ長時間学習で崩壊するのか――Qwenが出した「系列単位」の答え、GSPO
この記事は参考になりましたか?Was this article useful?
匿名の公開いいねです。記事の保存・お気に入りではなく、Featured、Top 3、重要度、掲載順位には影響しません。仕組みとプライバシーAnonymous public likes are reactions, not saved articles or bookmarks. They do not affect Featured, Top 3, importance, or listing order.How it works and privacy
AI2 点サマリSummary highlight
- 推論モデルのRL手法GRPOがトークン単位の重要度比のばらつきで長時間学習時に崩壊する問題を、一次情報(arXiv 2507.18071とQwen公式)から解説。
- Qwenが提案した系列単位で最適化するGSPOがこれをどう安定化させるかを読み解く。
Explains why the GRPO reinforcement-learning method collapses during long training due to noisy token-level importance ratios, and how Qwen's sequence-level GSPO stabilises optimisation for reasoning models.
本ページの要約は AI による自動生成です。日本語版と英語版は言語ごとに独立して生成されるため、表現や詳しさが異なる場合があります。正確性は元記事 (zenn.dev) をご確認ください。The summaries are AI-generated independently for each language, so wording and detail may differ. Verify accuracy at the original source (zenn.dev).




