HomeTags#reasoning

Tag timeline

#reasoning21 total

同じキーワードで束ねられた更新を確認できます。カテゴリをまたいだ関連ニュースや実装トピックの追跡に使えます。

Total21#reasoning の全掲載記事All listed entries tagged #reasoning
Showing21このページの表示件数Entries on this page
Page1/1静的ページ位置Static page position
Updated公開index snapshotPublished index snapshot

Entriespage 1/1 · 21 total

YESTERDAY1 entries
新規収集INDEXEDコミュニティCommunityLocal Models·Zenn AI

Qwen3.8 27B に Reasoning Effort を実装してみるThe author resolved Qwen3.8 27B's tendency to over-think on ambiguous tasks by…

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約Qwen3.8 27Bで思考が長引き生成上限に達する問題を、llama.cppのPer-request reasoning budgetで強制打ち切りすることで解消し、曖昧なタスクでも自律的に完走できるようになった。

AI SUMMARYThe author resolved Qwen3.8 27B's tendency to over-think on ambiguous tasks by enabling per-request reasoning budget in llama.cpp, allowing the model to complete complex tasks like Minecraft clone creation autonomously without hitting generation limits.

Qwen3.8 27B に Reasoning Effort を実装してみるog
Thu, Aug 62 entries
🔥 HOT公式OfficialCline/Roo·Cline Releases

原題 ENEnglish titleCline CLI v3.0.51Cline CLI v3.0.51

重要度 HighHigh priority公式リリース · Cline / Rooofficial release · Cline / Roo

AI要約推論エフォートがOllamaを含む全プロバイダーで一貫して適用されるようになり、推論オフの設定も正しく尊重される。また、meta/muse-spark-1.2-contributorがClineプロバイダーで選択可能になった。

AI SUMMARYReasoning effort now applies uniformly across all providers including Ollama, with opt-out respected everywhere, and meta/muse-spark-1.2-contributor is newly selectable on the Cline provider.

Cline CLI v3.0.51media
公式OfficialCline/Roo·Cline Releases

原題 ENEnglish titleCline SDK v0.0.71Cline SDK v0.0.71

重要度 MediumMedium priority公式リリース · Cline / Rooofficial release · Cline / Roo

AI要約推論設定がAI SDKプロバイダー間で統一的に解決されるようになり、努力レベルや有効/無効フラグがOllamaを含むネイティブ設定にマッピングされた。明示的な推論無効化リクエストが最優先されるよう改善された。

AI SUMMARYReasoning settings now resolve portably across AI SDK providers, mapping effort levels and enable/disable flags to each provider's native setting including Ollama, with explicit disable requests taking top priority.

Cline SDK v0.0.71media
Tue, Aug 41 entries
公式OfficialCopilot·GitHub Changelog

Copilot クラウドエージェントの推論レベルをカスタマイズ可能にCustomize the reasoning level for Copilot cloud agent

重要度 MediumMedium priority変更履歴 · GitHub Copilotchangelog · GitHub Copilot

AI要約GitHub Copilot クラウドエージェントにタスクを委任する際、対応モデルで推論レベルを設定できるようになった。これにより処理の深さをユーザーが制御できる。

AI SUMMARYGitHub Copilot cloud agent now lets users configure the reasoning level for supported models when delegating tasks, giving finer control over how deeply the agent thinks through problems.

Customize the reasoning level for Copilot cloud agentog
Mon, Jul 271 entries
コミュニティCommunityLocal Models·Zenn LLM

ローカルLLMにThoughtsStoreを搭載させてみた(実装応用編)This article demonstrates how to integrate a ThoughtsStore into a local LLM…

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約ローカルLLMにThoughtsStoreを組み込む応用実装を解説した記事で、思考履歴を永続化することでLLMの推論品質と文脈保持能力を向上させる手法を紹介している。

AI SUMMARYThis article demonstrates how to integrate a ThoughtsStore into a local LLM setup, enabling persistent storage of reasoning traces to improve inference quality and context retention.

Wed, Jul 221 entries
公式OfficialAgent Frameworks·AWS Machine Learning Blog

Amazon Novaによる教師あり微調整のための自己蒸留推論の探求Exploring self-distilled reasoning for supervised fine-tuning with Amazon Nova

重要度 MediumMedium priority技術記事 · Agent Frameworkstechnical post · Agent Frameworks

AI要約Amazon Novaモデルを用いて、モデル自身の推論プロセスをデータとして活用する自己蒸留手法でSFTの品質を向上させる方法を解説。外部アノテーションなしで高品質な学習データを生成できる点が重要。

AI SUMMARYThis article explores using self-distilled reasoning traces from Amazon Nova models to improve supervised fine-tuning quality, enabling higher-quality training data without external annotation.

Tue, Jul 211 entries
コミュニティCommunityLocal Models·Qiita LLM

Claude Fable 5 を9Bモデルに蒸留? 100万トークンの超長文推理モデル「Qwythos-9B」を4GBのVRAMで動かすQwythos-9B is a purported Claude Fable 5 distillation that supports 1M-token…

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約Qwythos-9BはClaude Fable 5からの蒸留とされる9Bパラメータの推論モデルで、100万トークンのコンテキストを持ちながら4GB VRAMで動作する点が注目される。

AI SUMMARYQwythos-9B is a purported Claude Fable 5 distillation that supports 1M-token context while running on just 4 GB of VRAM, making long-context reasoning accessible on consumer hardware.

Mon, Jul 201 entries
コミュニティCommunityLocal Models·Zenn LLM

BIRD:ブートストラップ自己蒸留で推論CoTを64%圧縮しつつ精度も向上BIRD is a bootstrap self-distillation method that compresses chain-of-thought…

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約BIRDはモデル自身の推論チェーンをブートストラップ自己蒸留で圧縮する手法で、CoTトークン数を最大64%削減しながら精度を維持・向上させる。推論コストの削減と性能の両立を示した点で注目に値する。

AI SUMMARYBIRD is a bootstrap self-distillation method that compresses chain-of-thought reasoning traces by up to 64% while maintaining or improving accuracy. This matters because it offers a practical path to reducing inference costs without sacrificing model performance.

Sat, Jul 181 entries
コミュニティCommunityClaude Code·Qiita Claude

Kimi K3 と Claude Fable 5 を実測比較:差が出たのは推論力より出力予算と検証性A hands-on benchmark comparing Kimi K3 and Claude Fable 5 found that practical…

重要度 MediumMedium priority技術記事 · Claude / Claude Codetechnical post · Claude / Claude Code

AI要約Kimi K3 と Claude Fable 5 を実際のタスクで比較した結果、純粋な推論精度よりも出力トークン予算の柔軟性と回答の検証しやすさに実用上の差が現れた。モデル選定の判断軸を見直す上で参考になる知見を提供している。

AI SUMMARYA hands-on benchmark comparing Kimi K3 and Claude Fable 5 found that practical differences stem less from raw reasoning ability and more from output budget flexibility and answer verifiability, offering a useful framework for model selection.

Tue, Jul 145 entries
論文PaperPapers/Benchmarks·arXiv cs.LG

知識グラフとグラフニューラルネットワークの融合:包括的サーベイKnowledge Graphs Meet Graph Neural Networks: A Comprehensive Survey

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約知識グラフとGNNを組み合わせた研究領域を体系的に整理し、知識グラフ補完・推論・質問応答などへの応用を網羅的に調査した論文。両技術の相互強化の可能性と今後の課題を明示している。

AI SUMMARYThis survey systematically reviews how knowledge graphs and graph neural networks reinforce each other across tasks like KG completion, reasoning, and QA, offering a unified taxonomy and highlighting open research challenges.

論文PaperPapers/Benchmarks·arXiv cs.LG

訓練不要なLLM推論のための深度エントロピー誘導サンプリングDepth-Entropy Guided Sampling for Training-Free LLM Reasoning

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約各トークン生成時にTransformerの層ごとのエントロピーを活用してサンプリングを動的に調整する手法を提案し、追加学習なしに推論精度を向上させる。計算コストを抑えながら複雑な推論タスクの性能を高められる点が重要。

AI SUMMARYThis paper proposes a training-free sampling method that uses per-layer entropy signals from Transformer depth to guide token generation, improving LLM reasoning without any fine-tuning. It offers a practical way to boost performance on complex reasoning benchmarks at low additional cost.

論文PaperPapers/Benchmarks·arXiv cs.AI

潜在的CoT推論を動的システムとして解釈するInterpreting Latent CoT Reasoning as Dynamical Systems

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本研究はLLMの潜在空間におけるChain-of-Thought推論を動的システムの観点から分析し、推論過程の内部構造を数理的に解明する手法を提案する。推論メカニズムの解釈可能性向上に貢献する。

AI SUMMARYThis paper proposes a dynamical-systems framework for analyzing latent Chain-of-Thought reasoning in LLMs, offering a principled mathematical lens to understand how intermediate reasoning steps evolve in hidden states. It advances interpretability of complex multi-step reasoning.

論文PaperPapers/Benchmarks·arXiv cs.CL

量子化LLM推論におけるサイレント障害:「中空収束」と障害モードシフトの分類論的分析Silent Failures in Quantized LLM Reasoning: A Taxonomy-Based Analysis of Hollow Convergence and Failure Mode Shifts

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約量子化されたLLMが表面上は正しく見える回答を生成しながら内部推論が破綻する「中空収束」現象を分類・分析し、量子化が引き起こす障害モードの質的変化を明らかにした研究。

AI SUMMARYThis paper identifies and classifies "hollow convergence" in quantized LLMs—where models produce plausible-looking outputs while reasoning has silently broken down—revealing systematic failure mode shifts that standard benchmarks fail to detect.

論文PaperPapers/Benchmarks·arXiv cs.CL

非英語言語における推論コスト:日本語を事例とした研究Cost of Reasoning in non-English Languages: A Case Study on Japanese

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約推論型LLMが日本語などの非英語言語でタスクを処理する際、英語より大幅に多くのトークンを消費することを実証した研究。多言語展開におけるコストと効率の課題を明らかにしている。

AI SUMMARYThis paper demonstrates that reasoning LLMs consume significantly more tokens when processing non-English languages like Japanese compared to English, revealing hidden cost and efficiency disparities in multilingual deployments.

Fri, Jul 103 entries
論文PaperPapers/Benchmarks·arXiv cs.AI

文脈内探索はいつ有効か?リフレクション駆動推論のサンプリング複雑性理論When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本論文は、LLMが推論中に自己反省・再探索を行う「リフレクション」の有効条件をサンプリング複雑性の観点から理論的に解析し、どのような問題設定で計算コストに見合う恩恵が得られるかを明らかにする。

AI SUMMARYThis paper develops a sampling-complexity theory to formally characterize when in-context search and reflection-driven reasoning improve LLM performance, offering principled guidance on the conditions under which iterative self-reflection is worth its computational cost.

論文PaperPapers/Benchmarks·arXiv cs.AI

エージェントベースモデリングにおけるLLMを活用した推論LLM-powered reasoning in agent-based modeling

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約LLMをエージェントベースモデルに組み込むことで、エージェントの意思決定に高度な推論能力を付与する手法を提案。社会シミュレーションや複雑系研究の表現力向上に貢献する。

AI SUMMARYThis paper proposes integrating large language models into agent-based modeling to enable human-like reasoning in simulated agents, improving the fidelity of social and complex-systems simulations.

論文PaperPapers/Benchmarks·arXiv cs.CL

LLMの論理は信頼できるか?グラフベースのフレームワークで不確実性・一貫性・頑健性を定量化Can We Trust LLM's Logic? Quantifying Uncertainty, Coherence, and Robustness via a Graph-Based Framework

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本研究はグラフ構造を用いてLLMの推論における不確実性、論理的一貫性、入力変動への頑健性を定量的に評価する手法を提案する。LLMの信頼性を客観的に測る基盤として重要な貢献となる。

AI SUMMARYThis paper proposes a graph-based framework to quantify uncertainty, logical coherence, and robustness in LLM reasoning, enabling more objective evaluation of whether LLM outputs can be trusted for critical tasks.

Thu, Jul 91 entries
論文PaperPapers/Benchmarks·arXiv cs.AI

推論一貫性スキャン:AI安全性評価におけるChain-of-Thought妥当性監査フレームワークReasoning Consistency Scanning: A Framework for Auditing Chain-of-Thought Validity in AI Safety Evaluations

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約AIの思考連鎖(Chain-of-Thought)推論の一貫性を体系的に監査するフレームワークを提案し、安全性評価における推論の欠陥や矛盾を検出する手法を示した研究。信頼性の高いAI安全評価の実現に貢献する。

AI SUMMARYThis paper proposes a framework for systematically auditing chain-of-thought reasoning in AI safety evaluations, detecting logical inconsistencies and flawed reasoning steps. It matters because reliable safety assessments depend on valid reasoning chains.

Wed, Jun 171 entries
新規収集INDEXED公式OfficialLocal Models·Hugging Face Blog

GLM-5.2: 長期タスク向けに設計された新モデルGLM-5.2: Built for Long-Horizon Tasks

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約ZAIがGLM-5.2を公開し、長時間にわたる複雑なタスクの処理能力を強化した。ローカルLLMとして実用的な性能を提供する点が注目される。

AI SUMMARYZAI released GLM-5.2, an open model optimized for long-horizon tasks requiring sustained reasoning and multi-step execution, making it a practical choice for local deployment.

Thu, May 281 entries
公式OfficialGemini/Gemma·Google Developers Blog

コミュニティがTunixとTPUを使ってGemmaに「思考」を学ばせた方法How the community trained Gemma to "Think" with Tunix and TPUs

重要度 InfoInformational深掘り候補 · 技術記事 · Gemini / GemmaDeep-dive candidate · technical post · Gemini / Gemma

AI要約KaggleのGoogle Tunixハッカソンで、開発者たちがTPUと限られた計算リソースを使い、小型の非推論ベースモデルを汎用推論エンジンへと変換。Tunixの強化学習・蒸留手法でGemmaに「思考」を教える取り組みを紹介する。

AI SUMMARYThe Google Tunix Hackathon on Kaggle challenged developers to turn small non-reasoning base models into general reasoning engines on TPUs with limited compute, showcasing how Tunix's RL and distillation techniques teach Gemma to reason.

Fri, May 81 entries
新規収集INDEXED公式OfficialClaude Code·YouTube - Anthropic

Anthropic、Claudeの思考を言語化する解釈可能性研究を公開(新しいタブで開きます)Translating Claude’s thoughts into language(opens in a new tab)

重要度 InfoInformational技術記事 · Claude / Claude Codetechnical post · Claude / Claude Code

AI要約Anthropicが、Claudeの内部表現を人間の言語へ翻訳する解釈可能性研究の動画を公開。モデルが推論中に何を考えているかを可視化し、AIの透明性と安全性の向上を目指す取り組みを示した。

AI SUMMARYAnthropic shares interpretability research that translates Claude's internal representations into human language, visualizing what the model thinks during reasoning to advance AI transparency and safety.