HomeTags#arxivPage 3

Tag timeline

#arxivpage 3/3

同じキーワードで束ねられた更新の続きです。カテゴリをまたいだ関連ニュースや実装トピックの追跡に使えます。

Total80#arxiv の全掲載記事All listed entries tagged #arxiv
Showing20このページの表示件数Entries on this page
Page3/3静的ページ位置Static page position
Updated公開index snapshotPublished index snapshot

Entriespage 3/3 · 80 total

Sat, Jul 112 entries
論文PaperPapers/Benchmarks·arXiv cs.AI

LLMが一致するとき、それは正しいのか?自己一貫性とモデル間合意を信頼度シグナルとして検証When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約複数のLLMが同じ答えを出す場合や単一モデルが一貫した回答を示す場合、それが正確さの指標になるかを実証的に検証した研究。合意が信頼度シグナルとして有効かを明らかにし、AI出力の信頼性評価に示唆を与える。

AI SUMMARYThis paper empirically audits whether self-consistency within a single LLM and agreement across multiple LLMs reliably signal factual correctness, finding nuanced limits to using consensus as a confidence proxy.

論文PaperPapers/Benchmarks·arXiv cs.AI

説得攻撃によりCoTモニタリングの有効性が低下する可能性Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約説得的なプロンプト操作がChain-of-Thoughtの監視機構を欺き、AIの安全性検査を回避できることを示した研究。CoTベースの監視手法の脆弱性として重要な警鐘となる。

AI SUMMARYThis research demonstrates that persuasion-based prompt attacks can undermine Chain-of-Thought monitoring, causing safety oversight mechanisms to miss harmful model behavior. The findings highlight a critical vulnerability in CoT-based AI supervision.

Fri, Jul 1015 entries
論文PaperPapers/Benchmarks·arXiv cs.AI

AgentLens: コーディングエージェント評価のための本番環境ベースのトラジェクトリレビューAgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約AgentLensはコーディングエージェントの動作軌跡を本番環境の基準で評価する新しいフレームワークで、エージェントの実用的な性能をより正確に測定できる。

AI SUMMARYAgentLens introduces a trajectory review framework for coding agents that uses production-grounded assessment, enabling more realistic and reliable evaluation of agent behavior beyond traditional benchmarks.

論文PaperPapers/Benchmarks·arXiv cs.AI

文脈内探索はいつ有効か?リフレクション駆動推論のサンプリング複雑性理論When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本論文は、LLMが推論中に自己反省・再探索を行う「リフレクション」の有効条件をサンプリング複雑性の観点から理論的に解析し、どのような問題設定で計算コストに見合う恩恵が得られるかを明らかにする。

AI SUMMARYThis paper develops a sampling-complexity theory to formally characterize when in-context search and reflection-driven reasoning improve LLM performance, offering principled guidance on the conditions under which iterative self-reflection is worth its computational cost.

論文PaperPapers/Benchmarks·arXiv cs.AI

エージェントベースモデリングにおけるLLMを活用した推論LLM-powered reasoning in agent-based modeling

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約LLMをエージェントベースモデルに組み込むことで、エージェントの意思決定に高度な推論能力を付与する手法を提案。社会シミュレーションや複雑系研究の表現力向上に貢献する。

AI SUMMARYThis paper proposes integrating large language models into agent-based modeling to enable human-like reasoning in simulated agents, improving the fidelity of social and complex-systems simulations.

論文PaperPapers/Benchmarks·arXiv cs.AI

計算・実験数学における SageMath 拡張 LLM エージェントの評価Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約SageMath を統合した LLM エージェントが計算数学・実験数学タスクをどの程度解けるかを体系的に評価した研究。数式処理システムとの連携がLLMの数学的推論能力を大幅に向上させることを示した。

AI SUMMARYThis paper systematically benchmarks LLM agents augmented with SageMath on computational and experimental mathematics tasks, showing that tool-integrated agents significantly outperform bare LLMs on complex mathematical problems.

論文PaperPapers/Benchmarks·arXiv cs.AI

「ハーネス効果」:オーケストレーション設計がエンタープライズ向けエージェントAIのトークン経済学を左右するThe Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約エージェントAIシステムにおけるオーケストレーション層の設計が、トークン消費量とコスト構造に直接影響することを実証した研究。企業導入における費用対効果の最適化に重要な示唆を与える。

AI SUMMARYThis paper demonstrates that the design of orchestration harnesses in agentic AI systems directly governs token consumption and cost structures, coining the term "Harness Effect." The findings offer actionable guidance for enterprises seeking to optimize the economics of large-scale AI agent deployments.

論文PaperPapers/Benchmarks·arXiv cs.CL

ソルバーから研究へ:LLM駆動の形式数学が研究フロンティアに到達From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約大規模言語モデルが数学の問題を解くだけでなく、未解決問題の探索や定理証明など本格的な数学研究を支援できる段階に達しつつあることを論じたサーベイ論文。形式数学とLLMの融合が数学研究の在り方を変える可能性を示す。

AI SUMMARYThis survey paper examines how LLMs are advancing beyond competition-style problem solving into genuine mathematical research, including formal theorem proving and exploration of open problems, signaling a shift in how AI can contribute to frontier mathematics.

論文PaperPapers/Benchmarks·arXiv cs.CL

DeepSearch-World: 検証可能な環境における深層検索エージェントの自己蒸留DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約検証可能な環境でエージェントが自己蒸留により深層検索能力を向上させる新手法を提案。強化学習なしに検索精度を高められる点が注目される。

AI SUMMARYDeepSearch-World proposes a self-distillation framework that trains deep search agents in a verifiable environment, improving search accuracy without relying on reinforcement learning.

論文PaperPapers/Benchmarks·arXiv cs.CL

人間とLLMの協働による拡張可能な文化固有のステレオタイプデータセット構築Scalable and Culturally Specific Stereotype Dataset Construction via Human-LLM Collaboration

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約LLMと人間のアノテーターを組み合わせ、文化ごとに異なるステレオタイプを効率的に収集・構築する手法を提案。AIの公平性評価に役立つ多様なデータセット作成を可能にする。

AI SUMMARYThis paper proposes a scalable pipeline combining human annotators and LLMs to build culturally specific stereotype datasets, enabling more representative fairness evaluations for AI systems across diverse cultures.

論文PaperPapers/Benchmarks·arXiv cs.CL

非現実的なトークンが強化されるとき:LLM強化学習のためのテール考慮クレジット調整When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約強化学習でLLMを訓練する際、低確率トークンが誤って強化される問題を指摘し、テール分布を考慮したクレジット調整手法を提案。報酬の帰属精度を高めることでモデルの品質向上を図る。

AI SUMMARYThis paper identifies that low-probability tokens can be incorrectly reinforced during LLM RL training, and proposes a tail-aware credit calibration method to more accurately assign reward signals, improving overall model quality.

論文PaperPapers/Benchmarks·arXiv cs.CL

全二重音声エージェント向けLALMオーディオ審判の信頼性評価A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本研究は、全二重音声エージェントの評価に用いられるLALMベースの自動審判モデルの信頼性を体系的に検証し、その一致度や偏りを明らかにした。自動評価の限界を把握することで、より堅牢な音声AIの品質測定が可能になる。

AI SUMMARYThis study systematically evaluates the reliability of large audio language model (LALM) judges used to assess full-duplex voice agents, revealing consistency and bias issues. The findings help establish more trustworthy automated evaluation pipelines for conversational speech AI.

論文PaperPapers/Benchmarks·arXiv cs.CL

Hallucination Self-Play: 進化した生成器による強化検出器のブートストラップHallucination Self-Play: Bootstrapping Reinforced Detector via Evolved Generator

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約LLMの幻覚検出を改善するため、生成器と検出器が自己対戦的に互いを強化し合うフレームワークを提案。ラベル付きデータなしで高精度な幻覚検出を実現できる点が重要。

AI SUMMARYThis paper proposes a self-play framework where a hallucination generator and detector iteratively improve each other, enabling robust hallucination detection without relying on expensive labeled data.

論文PaperPapers/Benchmarks·arXiv cs.CL

実行から教育へ:LLMにおける教育的制御を測定するBloom準拠フレームワークFrom Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約ブルームの分類法に基づき、LLMが教育的文脈でどの程度学習者の認知レベルを制御できるかを定量評価するフレームワークを提案。AIチュータリングの品質保証に新たな指標をもたらす。

AI SUMMARYThis paper proposes a framework aligned with Bloom's Taxonomy to measure how well LLMs can control educational outputs across cognitive levels, enabling more rigorous evaluation of AI tutoring systems.

論文PaperPapers/Benchmarks·arXiv cs.CL

低遅延システムにおけるツール生成と自己進化型LLMエージェントTool-Making and Self-Evolving LLM Agents in Low-Latency Systems

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約LLMエージェントが自らツールを作成・改良しながら低遅延環境で動作する手法を提案し、エージェントの自律的な能力拡張と応答速度の両立を実現した研究。

AI SUMMARYThis paper proposes a framework where LLM agents autonomously create and refine tools while operating under low-latency constraints, enabling self-improvement without sacrificing response speed.

論文PaperPapers/Benchmarks·arXiv cs.CL

LLMの論理は信頼できるか?グラフベースのフレームワークで不確実性・一貫性・頑健性を定量化Can We Trust LLM's Logic? Quantifying Uncertainty, Coherence, and Robustness via a Graph-Based Framework

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本研究はグラフ構造を用いてLLMの推論における不確実性、論理的一貫性、入力変動への頑健性を定量的に評価する手法を提案する。LLMの信頼性を客観的に測る基盤として重要な貢献となる。

AI SUMMARYThis paper proposes a graph-based framework to quantify uncertainty, logical coherence, and robustness in LLM reasoning, enabling more objective evaluation of whether LLM outputs can be trusted for critical tasks.

論文PaperPapers/Benchmarks·arXiv cs.CL

PLURAL: 価値アラインメントのためのグローバルデータセットPLURAL: A Global Dataset for Value Alignment

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約PLURALは多文化・多言語にわたる価値観の違いを捉えた大規模データセットで、AIの価値アラインメント研究に多様な視点を提供する。単一文化に偏ったベンチマークの限界を補う点で重要。

AI SUMMARYPLURAL is a large-scale multilingual dataset capturing diverse human values across cultures, designed to improve value alignment in AI systems. It addresses the critical gap left by culturally homogeneous benchmarks.

Thu, Jul 93 entries
論文PaperPapers/Benchmarks·arXiv cs.AI

社会規範の学習が動的な人間とAIの協調における互換性を高めるLearning social norms enhances compatibility in dynamic human-AI coordination

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約AIエージェントが社会規範を学習することで、初対面の人間パートナーとの動的な協調タスクにおける適応性と互換性が向上することを示した研究。人間とAIの協働設計に新たな指針を提供する。

AI SUMMARYThis research shows that AI agents trained to learn social norms achieve better compatibility with unfamiliar human partners in dynamic coordination tasks, offering new design principles for human-AI teaming.

論文PaperPapers/Benchmarks·arXiv cs.AI

マルチエージェントLLM安全性における「操作的リフレーミング」と「承認フレーム委譲」Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約複数のLLMエージェントが連携する環境で、タスクの表現を巧みに書き換えたり承認済みとして偽装することで安全制約を回避できる脆弱性を分析した研究。マルチエージェント設計における信頼伝播の危険性を示す点で重要。

AI SUMMARYThis paper analyzes how adversarial prompts can bypass safety constraints in multi-agent LLM systems by reframing tasks or falsely presenting delegated actions as pre-approved, exposing critical trust-propagation risks in agentic pipelines.

論文PaperPapers/Benchmarks·arXiv cs.AI

推論一貫性スキャン:AI安全性評価におけるChain-of-Thought妥当性監査フレームワークReasoning Consistency Scanning: A Framework for Auditing Chain-of-Thought Validity in AI Safety Evaluations

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約AIの思考連鎖(Chain-of-Thought)推論の一貫性を体系的に監査するフレームワークを提案し、安全性評価における推論の欠陥や矛盾を検出する手法を示した研究。信頼性の高いAI安全評価の実現に貢献する。

AI SUMMARYThis paper proposes a framework for systematically auditing chain-of-thought reasoning in AI safety evaluations, detecting logical inconsistencies and flawed reasoning steps. It matters because reliable safety assessments depend on valid reasoning chains.