新世代Qwen3.6を4070で測ったら「前世代の半速」に見えた。だが犯人はモデルではなかったOn an RTX 4070, Qwen first looked half as fast as the prior generation, but…
匿名の公開いいねです。記事の保存・お気に入りではなく、Featured、Top 3、重要度、掲載順位には影響しません。仕組みとプライバシーAnonymous public likes are reactions, not saved articles or bookmarks. They do not affect Featured, Top 3, importance, or listing order.How it works and privacy
RTX 4070でQwenのMoEモデルを測ると当初は前世代の半速に見えたが、原因はモデルでなく構成にあり、エキスパートをCPUにオフロードすると2.8倍速い34.6 tok/sを達成し、標準7問で全問正解した。
On an RTX 4070, Qwen first looked half as fast as the prior generation, but offloading MoE experts to the CPU gave a 2.8x speedup to 34.6 tok/s with 7/7 correct, proving the config—not the model—was the culprit.
本ページの要約は AI による自動生成です。日本語版と英語版は言語ごとに独立して生成されるため、表現や詳しさが異なる場合があります。正確性は元記事 (qiita.com) をご確認ください。The summaries are AI-generated independently for each language, so wording and detail may differ. Verify accuracy at the original source (qiita.com).




