HomearXiv Papers

Research lane

arXiv Papers80 papers

arXiv 系ソースは通常タイムラインから分離しました。論文だけをまとめて追いたい時は、このページで cs.AI / cs.CL / cs.SE / cs.LG を確認できます。Follow cs.AI, cs.CL, cs.SE, and cs.LG papers in this dedicated lane, separate from the main Timeline.

Papers80現在のarXivレーンCurrent arXiv lane
Last 7d0公開日基準の直近7日Published in the rolling last 7 days
Last 30d6source統計の直近30日Archive-backed source totals, last 30 days
Updated公開index snapshotPublished index snapshot
Sources4登録済みarXiv feedRegistered arXiv feeds
Main timeline1840arXivを除く現在一覧Current main Timeline, excluding arXiv

Filter

View

Latest papers

arXiv 論文一覧All · 80 papers

論文ソースだけを新着順に表示します。source filter で cs.AI / cs.CL / cs.SE / cs.LG を絞り込み、読み込み量に応じて Compact 表示へ切り替えられます。

Tue, Jul 281 papers
論文PaperPapers/Benchmarks·arXiv cs.LG

Semalith v1.4: Llama-Guard-3-8Bの44分の1のパラメータ数で最先端のプロンプトインジェクション検出を実現した184Mキャリブレーション済み安全分類器Semalith v1.4: A Calibrated 184M Safety Classifier Achieving State-of-the-Art Prompt-Injection Detection at 44x Fewer Parameters than Llama-Guard-3-8B

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約Semalith v1.4は1億8400万パラメータの軽量安全分類器で、Llama-Guard-3-8Bの44分の1のサイズながらプロンプトインジェクション検出で同等以上の精度を達成した。小規模モデルでも高精度な安全フィルタリングが可能であることを示し、実用的なデプロイコストの大幅削減につながる。

AI SUMMARYSemalith v1.4 is a 184M-parameter safety classifier that matches or surpasses Llama-Guard-3-8B on prompt-injection detection while using 44x fewer parameters. This demonstrates that highly capable safety filtering can be achieved at a fraction of the computational cost, making deployment far more practical.

Mon, Jul 271 papers
論文PaperPapers/Benchmarks·arXiv cs.LG

時間的介入下におけるパーソナルLLMエージェントのユーザー条件付き評価に向けてToward User-Conditioned Evaluation of Personal LLM Agents under Temporal Interventions

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本論文は、個人用LLMエージェントをユーザーの状況や時間的変化を考慮して評価する新たなフレームワークを提案し、既存ベンチマークでは捉えられなかった現実的な評価軸を提供する。

AI SUMMARYThis paper proposes a framework for evaluating personal LLM agents conditioned on individual user contexts and temporal interventions, addressing gaps in existing benchmarks that overlook real-world variability.

Fri, Jul 243 papers
論文PaperPapers/Benchmarks·arXiv cs.AI

大規模言語モデルにおける不完全プロンプトによるジェイルブレイクIncomplete Prompt Jailbreaks in Large Language Models

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約プロンプトを意図的に未完成にすることでLLMの安全制約を回避できる新たな脆弱性が報告された。この手法はモデルの応答補完メカニズムを悪用するため、既存の防御策では対処が難しい。

AI SUMMARYResearchers demonstrate that deliberately incomplete prompts can bypass safety guardrails in LLMs by exploiting their tendency to complete partial inputs. This reveals a novel attack surface that existing alignment defenses may not adequately address.

論文PaperPapers/Benchmarks·arXiv cs.SE

AIが生成したコードにおけるセキュリティ脆弱性パターン:モデル横断比較研究Security Vulnerability Patterns in AI-Generated Code: A Cross-Model Comparative Study

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約複数のAIコード生成モデルを横断的に比較し、生成コードに共通して現れるセキュリティ脆弱性のパターンを分析した研究。どのモデルがどの種類の脆弱性を生みやすいかを明らかにし、安全なAI活用に向けた知見を提供する。

AI SUMMARYThis study systematically compares security vulnerability patterns across multiple AI code generation models, identifying which weakness types each model tends to introduce. The findings offer actionable guidance for developers and organizations relying on AI-assisted coding.

論文PaperPapers/Benchmarks·arXiv cs.SE

Tencent WorkBuddy Bench: 汚染耐性タスク構築を備えたマルチドメインコーディングエージェントベンチマークTencent WorkBuddy Bench: A Multi-Domain Coding-Agent Benchmark with Contamination-Resistant Task Construction

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約テンセントはコーディングエージェント評価用ベンチマーク「WorkBuddy Bench」を発表。学習データ汚染を防ぐ設計と複数ドメイン対応により、より信頼性の高いエージェント性能評価を実現する。

AI SUMMARYTencent introduces WorkBuddy Bench, a coding-agent benchmark spanning multiple domains with a contamination-resistant task construction method, enabling more reliable and fair evaluation of LLM-based coding agents.

Wed, Jul 221 papers
論文PaperPapers/Benchmarks·arXiv cs.LG

Interactive Training 2: ライブモデル訓練のための監査可能なコントロールプレーンInteractive Training 2: Auditable Control Plane for Live Model Training

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約モデル訓練中にリアルタイムで介入・監査できるコントロールプレーンの設計を提案し、訓練プロセスの透明性と制御性を高める研究。人間がループに参加しながら学習を動的に調整できる点が重要。

AI SUMMARYThis paper proposes an auditable control plane for live model training, enabling real-time human intervention and oversight during the training process. It advances interactive and accountable ML workflows by making training dynamics inspectable and steerable.

Wed, Jul 157 papers
論文PaperPapers/Benchmarks·arXiv cs.LG

報酬はいつ状態を教えるか?隠れオートマトン操作変数と群言語境界When Does Reward Teach State? A Hidden-Automaton Instrument and the Group-Language Boundary

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約強化学習において報酬信号だけで潜在的な環境状態を識別できる条件を理論的に解析し、群言語の境界が識別可能性の鍵となることを示した研究。

AI SUMMARYThis paper establishes theoretical conditions under which reward signals alone can identify latent environment states in RL, introducing a hidden-automaton instrumental framework and showing that the group-language boundary determines identifiability.

When Does Reward Teach State? A Hidden-Automaton Instrument and the Group-Language Boundaryog
論文PaperPapers/Benchmarks·arXiv cs.LG

LiteTopK: 次元の呪いを活用した長文脈スパースアテンション向け融合インデクサー・TopKカーネルLiteTopK: Exploiting the Curse of Dimensionality for a Fused Indexer-TopK Kernel in Long-Context Sparse Attention

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約LiteTopKは高次元空間での距離集中現象を逆手に取り、スパースアテンションのインデクサーとTopK選択を単一カーネルに融合することで、長文脈推論の効率を大幅に改善する手法を提案する。

AI SUMMARYLiteTopK leverages the concentration of distances in high dimensions to fuse the indexer and TopK selection into a single GPU kernel, significantly reducing overhead in long-context sparse attention inference.

論文PaperPapers/Benchmarks·arXiv cs.LG

マージすべきモデルを間違えていないか?LLMのモデルマージにおける専門家の訓練時間の影響Are we Merging the Right Models? Impact of Expert Training Duration on Model Merging for LLMs

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約LLMのモデルマージにおいて、専門家モデルの訓練ステップ数がマージ後の性能に大きく影響することを示した研究。適切な訓練段階のモデルを選ぶことがマージ成功の鍵となる。

AI SUMMARYThis study investigates how the training duration of expert models affects the quality of merged LLMs, finding that selecting models at the right training stage is critical for achieving strong post-merge performance.

論文PaperPapers/Benchmarks·arXiv cs.LG

PFAdapter: 個人化連合MLLMのための階層的LoRA分解PFAdapter: Hierarchical LoRA Decomposition for Personalized Federated MLLMs

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約PFAdapterは階層的LoRA分解により、連合学習環境でマルチモーダル大規模言語モデルを個人化する手法を提案し、プライバシーを保ちながら各クライアントの異質なデータに適応できる点が重要です。

AI SUMMARYPFAdapter proposes a hierarchical LoRA decomposition framework for personalizing multimodal LLMs in federated learning settings, enabling privacy-preserving adaptation to heterogeneous client data without sharing raw information.

論文PaperPapers/Benchmarks·arXiv cs.LG

連合学習におけるMLLMファインチューニングのための弾性正則化と合成リプレイを用いた継続学習Continual Learning with Elastic Regularization and Synthetic Replay for Federated MLLM Fine-Tuning

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約連合学習環境でのMLLMファインチューニング時に生じる破滅的忘却を、弾性正則化と合成データリプレイの組み合わせで緩和する手法を提案。プライバシーを保ちながら継続的なモデル更新を実現できる点が重要。

AI SUMMARYThis paper proposes combining elastic weight regularization with synthetic data replay to combat catastrophic forgetting in federated multimodal LLM fine-tuning, enabling privacy-preserving continual learning across distributed clients.

論文PaperPapers/Benchmarks·arXiv cs.LG

ニューラル演算子の自動発見に向けたエージェント型AI科学コミュニティAn Agentic AI Scientific Community for Automated Neural Operator Discovery

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約複数のAIエージェントが科学者コミュニティとして協調し、ニューラル演算子のアーキテクチャを自動探索・発見する枠組みを提案。人手によるアーキテクチャ設計を大幅に省力化できる点で注目される。

AI SUMMARYResearchers propose a multi-agent AI system that mimics a scientific community to autonomously discover and design neural operator architectures, reducing the need for manual expert design in scientific machine learning.

論文PaperPapers/Benchmarks·arXiv cs.LG

「Speculate with Memory」: LLMエージェントの無損失高速化手法Speculate with Memory: Lossless Acceleration for LLM Agents

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約過去の実行履歴をメモリとして活用する投機的デコード手法を提案し、LLMエージェントの推論を無損失で大幅に高速化することを実現した。繰り返しタスクが多いエージェント環境での実用的な高速化に貢献する。

AI SUMMARYThis paper proposes a speculative decoding method that leverages past execution history as memory to accelerate LLM agents without any output quality loss, offering practical speedups in repetitive agentic workflows.

Tue, Jul 1427 papers
論文PaperPapers/Benchmarks·arXiv cs.LG

知識グラフとグラフニューラルネットワークの融合:包括的サーベイKnowledge Graphs Meet Graph Neural Networks: A Comprehensive Survey

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約知識グラフとGNNを組み合わせた研究領域を体系的に整理し、知識グラフ補完・推論・質問応答などへの応用を網羅的に調査した論文。両技術の相互強化の可能性と今後の課題を明示している。

AI SUMMARYThis survey systematically reviews how knowledge graphs and graph neural networks reinforce each other across tasks like KG completion, reasoning, and QA, offering a unified taxonomy and highlighting open research challenges.

論文PaperPapers/Benchmarks·arXiv cs.LG

AuditWeave: AIアシストおよびデータ変換ワークフロー向けの改ざん防止・監査者対応エビデンス層AuditWeave: A Tamper-Evident, Auditor-Navigable Evidence Layer for AI-Assisted and Data-Transformation Workflows

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI 要約 ENEnglish AI summaryAuditWeave proposes a tamper-evident evidence layer that cryptographically secures processing histories in AI-assisted and data-transformation pipelines, making them navigable for auditors. This addresses growing demands for accountability and regulatory compliance in automated workflows.

AI SUMMARYAuditWeave proposes a tamper-evident evidence layer that cryptographically secures processing histories in AI-assisted and data-transformation pipelines, making them navigable for auditors. This addresses growing demands for accountability and regulatory compliance in automated workflows.

論文PaperPapers/Benchmarks·arXiv cs.LG

MawForge: ローカル環境でのMixture-of-Experts推論向けメモリ制約エキスパート実体化MawForge: Memory-Bounded Expert Materialization for Local Mixture-of-Experts Inference

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約MawForgeは、限られたメモリ環境でMoEモデルをローカル推論する際に、使用頻度の高いエキスパートを事前に実体化してキャッシュする手法を提案する。これによりメモリ効率を維持しながら推論速度を大幅に改善できる。

AI SUMMARYMawForge proposes a memory-bounded strategy for local Mixture-of-Experts inference by selectively materializing frequently activated experts within a fixed memory budget, enabling faster inference on consumer hardware without sacrificing model quality.

論文PaperPapers/Benchmarks·arXiv cs.LG

コーディングエージェントが行動するために実際に必要なコンテキストとは?What Context Does a Coding Agent Actually Need to Act?

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本研究は、コーディングエージェントが効果的にタスクを遂行するために必要なコンテキストの種類と量を体系的に分析し、不要な情報を削減して性能を維持できる条件を明らかにした。

AI SUMMARYThis paper systematically investigates which types and amounts of context are truly necessary for coding agents to act effectively, revealing that many agents can maintain performance with significantly reduced input context.

論文PaperPapers/Benchmarks·arXiv cs.LG

LLMにおける参照ベースの蒸留検出Reference-Based Distillation Detection in LLMs

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約大規模言語モデルが他のモデルから知識蒸留されているかを参照モデルを用いて検出する手法を提案。モデルの知的財産保護やサプライチェーンの透明性確保に貢献する。

AI SUMMARYThis paper proposes a reference-based method to detect whether an LLM has been trained via knowledge distillation from another model, enabling protection of model intellectual property and improving AI supply-chain transparency.

論文PaperPapers/Benchmarks·arXiv cs.LG

訓練不要なLLM推論のための深度エントロピー誘導サンプリングDepth-Entropy Guided Sampling for Training-Free LLM Reasoning

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約各トークン生成時にTransformerの層ごとのエントロピーを活用してサンプリングを動的に調整する手法を提案し、追加学習なしに推論精度を向上させる。計算コストを抑えながら複雑な推論タスクの性能を高められる点が重要。

AI SUMMARYThis paper proposes a training-free sampling method that uses per-layer entropy signals from Transformer depth to guide token generation, improving LLM reasoning without any fine-tuning. It offers a practical way to boost performance on complex reasoning benchmarks at low additional cost.

論文PaperPapers/Benchmarks·arXiv cs.LG

低ランク注意残差(Low-Rank Attention Residuals)Low-Rank Attention Residuals

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約Transformerのアテンション層における残差接続を低ランク近似で置き換える手法を提案し、モデルの表現力を保ちながらパラメータ効率を大幅に改善できることを示した研究。

AI SUMMARYThis paper proposes approximating attention residuals with low-rank structures in Transformers, showing that model expressiveness can be maintained while significantly reducing parameter counts and improving efficiency.

論文PaperPapers/Benchmarks·arXiv cs.LG

安全な応答が重要:MLLMsにおける過剰拒否を軽減する出力認識型セーフティガードレールSafe responses matter: Output-aware safety guardrail mitigate over-refusal in MLLMs

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約マルチモーダル大規模言語モデルが安全なリクエストまで拒否しすぎる「過剰拒否」問題に対し、出力内容を考慮したガードレール手法を提案。有害コンテンツを防ぎつつ正当な要求への応答精度を向上させる。

AI SUMMARYThis paper proposes an output-aware safety guardrail for multimodal LLMs that reduces over-refusal by evaluating the model's generated response, not just the input. This improves usability without compromising safety.

論文PaperPapers/Benchmarks·arXiv cs.LG

EvoClawBench: エージェントは自身の実行履歴から再利用可能なスキルを学習できるか?EvoClawBench: Can Agents Learn Reusable Skills from Their Own Runs?

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約EvoClawBenchは、AIエージェントが過去の実行経験からスキルを抽出・再利用できるかを評価する新ベンチマークで、汎化能力の研究を促進する。

AI SUMMARYEvoClawBench introduces a benchmark for evaluating whether AI agents can extract and reuse skills from their own prior runs, advancing research into agent generalization and continual learning.

論文PaperPapers/Benchmarks·arXiv cs.LG

符号分岐繰り返しペナルティにおけるゲージ依存性と構造化出力の破損:モデル・推論スタック・代替制御手法にわたる測定Gauge dependence and structured-output corruption in sign-branched repetition penalties: measurements across models, inference stacks, and alternative repetition controls

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約繰り返しペナルティの実装において符号の分岐がゲージ依存性を生じさせ、JSON等の構造化出力を破損させることを実験的に示した研究。モデルや推論スタックをまたいだ測定により、代替制御手法の有効性も評価している。

AI SUMMARYThis paper demonstrates that sign-branched repetition penalty implementations introduce gauge dependence that corrupts structured outputs such as JSON across multiple models and inference stacks, and evaluates alternative repetition control strategies to mitigate the problem.

論文PaperPapers/Benchmarks·arXiv cs.AI

Format Sensitivity Index:トークン制御プロンプトラッパーの堅牢性とLLMベンチマークにおけるスキーマ準拠Format Sensitivity Index: Token-Controlled Prompt Wrapper Robustness and Schema Compliance in LLM Benchmarking

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約LLMがプロンプトの書式変更に対してどれだけ出力を安定させられるかを定量化する「Format Sensitivity Index」を提案し、ベンチマーク評価の信頼性向上に貢献する研究。

AI SUMMARYThis paper introduces the Format Sensitivity Index, a metric that quantifies how much LLM outputs shift under token-level prompt wrapper variations, highlighting reliability gaps in current benchmarking practices.

論文PaperPapers/Benchmarks·arXiv cs.AI

忠実であって修正はしない:マルチホップエージェントリレーにおけるメッセージ形式の影響はティアに依存するFaithful, Not Corrective: Message-Format Effects in Multi-Hop Agent Relays Are Tier-Dependent

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約マルチホップエージェントリレーでは、メッセージの形式が下流エージェントの動作に与える影響がエージェントの階層によって異なり、上流エージェントは誤りを修正せず忠実に伝達することが示された。複数エージェント系の設計における信頼性評価に重要な知見を提供する。

AI SUMMARYThis study finds that message-format effects in multi-hop agent relay chains are tier-dependent: agents faithfully propagate upstream content rather than correcting errors, with implications for reliability in multi-agent system design.

論文PaperPapers/Benchmarks·arXiv cs.AI

潜在的CoT推論を動的システムとして解釈するInterpreting Latent CoT Reasoning as Dynamical Systems

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本研究はLLMの潜在空間におけるChain-of-Thought推論を動的システムの観点から分析し、推論過程の内部構造を数理的に解明する手法を提案する。推論メカニズムの解釈可能性向上に貢献する。

AI SUMMARYThis paper proposes a dynamical-systems framework for analyzing latent Chain-of-Thought reasoning in LLMs, offering a principled mathematical lens to understand how intermediate reasoning steps evolve in hidden states. It advances interpretability of complex multi-step reasoning.

論文PaperPapers/Benchmarks·arXiv cs.AI

YUKTI: 自然言語の状況から堅牢で検証可能な意思決定へ――不確実性型命題IR・仮定ロバストパレートフロンティア・後悔証明書YUKTI: From Natural-Language Situations to Robust, Verifiable Decisions An Uncertainty-Typed Proposition IR, Assumption-Robust Pareto Frontiers, and a Regret Certificate

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約YUKTIは自然言語で記述された意思決定状況を不確実性型命題の中間表現に変換し、仮定に対してロバストなパレートフロンティアと後悔証明書を生成することで、検証可能な意思決定を実現するフレームワークである。曖昧な前提を明示的に扱える点が実用上の重要な貢献となっている。

AI SUMMARYYUKTI is a framework that converts natural-language decision scenarios into an uncertainty-typed proposition IR, then derives assumption-robust Pareto frontiers and regret certificates to produce verifiable, auditable decisions. It advances AI decision-making by explicitly handling ambiguous assumptions rather than ignoring them.

論文PaperPapers/Benchmarks·arXiv cs.SE

AfterVibe: 会話が終わった後に何が残るかAfterVibe: What Remains When the Conversation Ends

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本研究はAIとの対話終了後にユーザーの感情や認知にどのような影響が持続するかを分析し、チャットシステム設計における「余韻」の重要性を示した。

AI SUMMARYThis paper examines the emotional and cognitive residues that persist after human-AI conversations end, highlighting design implications for conversational systems that account for post-interaction effects.

論文PaperPapers/Benchmarks·arXiv cs.SE

AIエージェントが書いたコードはマージ後どうなるか?その追跡調査Do These Violent Delights Have Violent Ends? Measuring the Post-Merge Fate of Agentic Code

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本論文はAIエージェントが生成しマージされたコードのその後の運命を実証的に調査し、品質や保守性への長期的影響を定量化した研究である。エージェント生成コードの実用上のリスクを把握する上で重要な知見を提供する。

AI SUMMARYThis paper empirically tracks the post-merge lifecycle of code produced by AI coding agents, measuring its long-term quality, churn, and maintenance burden compared to human-written code. The findings inform real-world risk assessments of deploying agentic coding systems.

論文PaperPapers/Benchmarks·arXiv cs.SE

LLMを用いた静的解析アラートの判定とエラー低減技術Using LLMs to Adjudicate Static-Analysis Alerts with Error Reduction Techniques

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本論文は、静的解析ツールが生成する大量の誤検知アラートをLLMで自動判定し、エラー低減技術を組み合わせることで精度を高める手法を提案している。開発者の負担軽減とセキュリティ品質向上に貢献する研究成果である。

AI SUMMARYThis paper proposes using LLMs to automatically triage static-analysis alerts—distinguishing true bugs from false positives—while applying error reduction techniques to improve adjudication accuracy and reduce developer burden.

論文PaperPapers/Benchmarks·arXiv cs.CL

合意と反対意見:グループ推薦における主観的嗜好の動的LLMモデリングConsensus vs. Dissent: Dynamic LLM Modeling of Subjective Preferences in Group Recommenders

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本研究はグループ推薦システムにおいて、LLMを用いてメンバー間の合意と反対意見を動的にモデル化する手法を提案する。個人の主観的嗜好を柔軟に捉えることで、グループ全体の満足度向上を目指す。

AI SUMMARYThis paper proposes a dynamic LLM-based framework for group recommender systems that models both consensus and dissent among members' subjective preferences, improving collective satisfaction beyond simple preference aggregation.

Consensus vs. Dissent: Dynamic LLM Modeling of Subjective Preferences in Group Recommendersog
論文PaperPapers/Benchmarks·arXiv cs.CL

Index SLM テクニカルレポートIndex SLM Technical Report

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約Index SLMは、限られたパラメータ数で高い性能を実現する小型言語モデルシリーズであり、効率的なエッジ・オンデバイス推論への応用が期待される。

AI SUMMARYIndex SLM introduces a series of small language models designed to achieve competitive performance at reduced parameter counts, enabling practical deployment in edge and on-device scenarios.

論文PaperPapers/Benchmarks·arXiv cs.CL

RouteRec: 推薦エージェントの選択と集約に関する厳密な評価フレームワークRouteRec: Strict Evaluation of Recommender-Agent Selection and Aggregation

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約RouteRecは、複数の推薦エージェントをどう選択・集約するかを厳密に評価するベンチマークを提案し、エージェント間のルーティング戦略の有効性を体系的に測定できる点で重要。

AI SUMMARYRouteRec introduces a rigorous benchmark for evaluating how recommender agents are selected and aggregated, enabling systematic measurement of routing strategies across multiple agents.

論文PaperPapers/Benchmarks·arXiv cs.CL

言語モデルによるグローバルM&Aアービトラージ予測Global Merger-Arbitrage Forecasting with Language Models

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約大規模言語モデルをM&Aアービトラージの取引成否予測に応用し、財務テキストから有益なシグナルを抽出できることを示した研究。投資戦略への自然言語処理活用の可能性を広げる成果として注目される。

AI SUMMARYThis paper applies large language models to predict deal outcomes in merger-arbitrage investing, showing that NLP signals from financial text meaningfully improve forecasting accuracy across global markets.

論文PaperPapers/Benchmarks·arXiv cs.CL

忠実性を設計で担保:多様なステークホルダー向けLLM生成臨床試験サマリーの評価と改善Faithful by Design: Evaluating and Improving LLM-Generated Clinical Trial Summaries for Multi-Stakeholder Audiences

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約LLMが生成する臨床試験サマリーの事実忠実性を患者・医療者・研究者など複数の読者層に合わせて評価・改善する手法を提案。医療情報の誤りがもたらすリスクを低減する実用的な枠組みとして意義がある。

AI SUMMARYThis paper proposes methods to evaluate and improve the factual faithfulness of LLM-generated clinical trial summaries tailored to diverse audiences, including patients and clinicians. It addresses a critical safety concern by reducing hallucinations in high-stakes medical communication.

論文PaperPapers/Benchmarks·arXiv cs.CL

デバイス上でのリアルタイム字幕翻訳に向けたワークロード駆動最適化Workload-Driven Optimization for On-Device Real-Time Subtitle Translation

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本論文は、オンデバイス環境でリアルタイム字幕翻訳を実現するため、ワークロードの特性に基づいてモデルや処理を動的に最適化する手法を提案する。これにより、限られた計算資源でも低遅延かつ高品質な翻訳が可能になる。

AI SUMMARYThis paper proposes a workload-driven optimization framework for real-time subtitle translation running entirely on-device, dynamically adapting model execution to meet latency constraints without sacrificing translation quality.

論文PaperPapers/Benchmarks·arXiv cs.CL

量子化LLM推論におけるサイレント障害:「中空収束」と障害モードシフトの分類論的分析Silent Failures in Quantized LLM Reasoning: A Taxonomy-Based Analysis of Hollow Convergence and Failure Mode Shifts

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約量子化されたLLMが表面上は正しく見える回答を生成しながら内部推論が破綻する「中空収束」現象を分類・分析し、量子化が引き起こす障害モードの質的変化を明らかにした研究。

AI SUMMARYThis paper identifies and classifies "hollow convergence" in quantized LLMs—where models produce plausible-looking outputs while reasoning has silently broken down—revealing systematic failure mode shifts that standard benchmarks fail to detect.

論文PaperPapers/Benchmarks·arXiv cs.CL

シンガポールの言語環境に合わせた音声言語モデルの効率的な適応Efficiently Adapting Spoken Language Models for the Singaporean Context

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本研究は、シンガポール英語(Singlish)などの多言語混在環境に対応するため、既存の音声言語モデルを効率的にファインチューニングする手法を提案し、限られたリソースでの高精度な音声認識・理解を実現する。

AI SUMMARYThis paper proposes efficient adaptation methods for spoken language models targeting Singapore's multilingual context, achieving strong performance on Singlish and code-switching speech without requiring large-scale retraining.

論文PaperPapers/Benchmarks·arXiv cs.CL

非英語言語における推論コスト:日本語を事例とした研究Cost of Reasoning in non-English Languages: A Case Study on Japanese

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約推論型LLMが日本語などの非英語言語でタスクを処理する際、英語より大幅に多くのトークンを消費することを実証した研究。多言語展開におけるコストと効率の課題を明らかにしている。

AI SUMMARYThis paper demonstrates that reasoning LLMs consume significantly more tokens when processing non-English languages like Japanese compared to English, revealing hidden cost and efficiency disparities in multilingual deployments.

論文PaperPapers/Benchmarks·arXiv cs.CL

精度は同じ、証拠は不平等:ツール利用エージェントの意思決定面としての検索APIEqual Accuracy, Unequal Evidence: Search APIs as Decision Surfaces for Tool-Using Agents

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約検索APIが同程度の精度を示しても、返却される証拠の質や多様性に大きな差があり、ツール利用エージェントの意思決定に偏りをもたらすことを明らかにした研究。

AI SUMMARYThis paper shows that search APIs with similar accuracy can differ substantially in evidence quality and diversity, introducing hidden biases into tool-using agents' decisions.

Mon, Jul 1310 papers
論文PaperPapers/Benchmarks·arXiv cs.SE

プログラマーはLLMが生成したアサーションの評価が苦手で過信しがちProgrammers Are Poor and Overconfident Judges of LLM-Generated Assertions

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約LLMが生成したテストアサーションをプログラマーが評価する際、正確性を過大評価しミスを見落としやすいことが実験で示された。自動テスト生成の品質保証に人間のレビューだけでは不十分である可能性を示唆する。

AI SUMMARYA study found that programmers systematically overestimate the correctness of LLM-generated test assertions and miss significant errors, raising concerns about relying on human review as a quality gate for AI-generated tests.

Programmers Are Poor and Overconfident Judges of LLM-Generated Assertionsog
論文PaperPapers/Benchmarks·arXiv cs.SE

より良いハーネス、小さなモデル:自動ハーネス適応で90%コスト削減エージェントの構築Better Harnesses, Smaller Models: Building 90% Cheaper Agents via Automated Harness Adaptation

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約コーディングエージェントのテストハーネスを自動的に最適化することで、大型モデルに依存せず小型モデルでも高い性能を実現し、運用コストを約90%削減できることを示した研究。

AI SUMMARYThis paper shows that automatically adapting test harnesses for coding agents allows smaller, cheaper models to match large-model performance, cutting agent operational costs by roughly 90%.

論文PaperPapers/Benchmarks·arXiv cs.SE

LLM生成コードにおける「パッチワーク問題」The Patchwork Problem in LLM-Generated Code

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約LLMが生成するコードは断片的な知識を継ぎ接ぎした構造になりやすく、一貫性や保守性に欠けるという問題を論文が指摘している。この知見はAIコード生成ツールの評価・改善指針として重要な意味を持つ。

AI SUMMARYResearchers identify the "patchwork problem" in LLM-generated code, where outputs are stitched together from disparate training patterns, leading to inconsistency and poor maintainability. This has significant implications for how AI coding tools should be evaluated and improved.

論文PaperPapers/Benchmarks·arXiv cs.SE

SCATE: コスト効率の高いテスト生成のためにコーディングエージェントを監督する学習SCATE: Learning to Supervise Coding Agents for Cost-Effective Test Generation

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約SCATEはLLMベースのコーディングエージェントを監督者モデルで制御し、テスト生成コストを抑えながら品質を維持する手法を提案する。監督者がエージェントの行動を動的に評価することで、効率的なソフトウェアテスト自動化を実現する。

AI SUMMARYSCATE proposes training a supervisor model to guide LLM-based coding agents during automated test generation, reducing computational cost while maintaining coverage quality. This approach makes agent-driven software testing more practical for real-world use.

論文PaperPapers/Benchmarks·arXiv cs.SE

汎用から個別へ:ペルソナを考慮したコードレビュー説明の探求From Generic to Personalized: Exploring Persona-Aware Code Review Explanations

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本研究は、開発者の経験レベルや役割に応じてコードレビューのフィードバック説明を個別化するペルソナ対応アプローチを提案し、画一的な説明の限界を克服しようとしている。

AI SUMMARYThis paper proposes a persona-aware approach to generating code review explanations tailored to a developer's experience and role, showing that personalized feedback improves comprehension over generic explanations.

論文PaperPapers/Benchmarks·arXiv cs.SE

Bugs4Qを用いたQiskitプログラム修復におけるLLMのベンチマーク評価Benchmarking Large Language Models on Repairing Qiskit Programs using Bugs4Q

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約量子プログラムのバグ修復タスクにLLMを適用し、Bugs4Qベンチマークで性能を評価した研究。量子ソフトウェア開発における自動修復の可能性と限界を明らかにしている。

AI SUMMARYThis study evaluates large language models on automatically repairing buggy Qiskit quantum programs using the Bugs4Q benchmark, revealing both the promise and current limitations of LLM-based repair for quantum software.

論文PaperPapers/Benchmarks·arXiv cs.SE

Pythonの型アノテーションをジャストインタイムで自動更新する手法Automating Just-In-Time Python Type Annotation Updating

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約コード変更に伴い古くなったPythonの型アノテーションを自動的に検出・更新する手法を提案し、保守コストの削減と型安全性の維持を両立する。

AI SUMMARYThis paper proposes an automated approach to detect and update stale Python type annotations triggered by code changes, reducing maintenance burden while preserving type safety.

論文PaperPapers/Benchmarks·arXiv cs.SE

スキルマーケットの内側:ソフトウェアエンジニアリング活動から再利用可能なエージェントスキルへInside the Skill Market: From Software Engineering Activities to Reusable Agent Skills

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約ソフトウェアエンジニアリングのタスクをエージェントが再利用可能なスキルとして体系化する「スキルマーケット」フレームワークを提案し、エージェントの汎化性能と効率を向上させる研究。

AI SUMMARYThis paper proposes a "Skill Market" framework that distills software engineering activities into reusable agent skills, enabling more generalizable and efficient AI agents for SE tasks.

論文PaperPapers/Benchmarks·arXiv cs.SE

データ集約型コンピューティングにおけるプロパティテンプレートを用いたエージェント的証明とプロパティベーステストAgentic Proof and Property-Based Testing via Property-Templates in Data-Intensive Computing

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本論文は、データ集約型システムの検証にエージェントAIとプロパティテンプレートを組み合わせ、形式的証明とプロパティベーステストを自動化する手法を提案する。これにより複雑なデータ処理コードの信頼性検証コストを大幅に削減できる。

AI SUMMARYThis paper proposes using AI agents with reusable property-templates to automate formal proofs and property-based testing in data-intensive computing, reducing the manual effort required to verify correctness of complex data pipelines.

論文PaperPapers/Benchmarks·arXiv cs.SE

人間のテスト工程に着想を得たワークフローによるユニットテスト自動生成のためのマルチエージェントLLM協調Multi-Agent LLM Collaboration for Unit Test Generation via Human-Testing-Inspired Workflows

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約複数のLLMエージェントが人間のテスト設計プロセスを模倣して協調することで、ユニットテストの品質とカバレッジを向上させる手法を提案した研究。従来の単一モデルによる生成より効果的なテスト作成が可能になる。

AI SUMMARYThis paper proposes a multi-agent LLM framework that mimics human software testing workflows to collaboratively generate higher-quality unit tests with improved coverage, outperforming single-model approaches.

Sat, Jul 1112 papers
論文PaperPapers/Benchmarks·arXiv cs.AI

プロアクティブなエンタープライズエージェントのためのコンテキストグラフContext Graphs for Proactive Enterprise Agents

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本論文は、企業向けAIエージェントが自律的に行動するためのコンテキストグラフという新しい手法を提案し、エージェントが適切なタイミングで先回りして行動できる仕組みを示している。

AI SUMMARYThis paper proposes context graphs as a structured representation to enable proactive enterprise AI agents, allowing them to anticipate user needs and act autonomously at the right moment.

Context Graphs for Proactive Enterprise Agentsog
論文PaperPapers/Benchmarks·arXiv cs.AI

人間とLLMの混成集団に向けた対立的社会認識論Adversarial Social Epistemology for Assemblies of Humans and Large Language Models

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約人間とLLMが混在する集合的意思決定の場で、悪意ある操作や認識論的攻撃がどう機能するかを分析した研究。AIを含む社会的知識形成の堅牢性設計に重要な示唆を与える。

AI SUMMARYThis paper analyzes how adversarial actors can exploit mixed human-LLM assemblies to distort collective knowledge and decision-making, offering a framework for building more robust epistemic systems that include AI participants.

論文PaperPapers/Benchmarks·arXiv cs.AI

アライメント妥当性:ヘルスケアにおけるAI保証の新基準Alignment Plausibility: A New Standard for Assuring AI in Healthcare

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本論文は医療AIの安全性を評価する新概念「アライメント妥当性」を提案し、AIの挙動が臨床目標と一致しているかを体系的に検証する枠組みを示す。規制や倫理審査に応用可能な実用的基準として注目される。

AI SUMMARYThis paper proposes 'alignment plausibility' as a new standard for evaluating whether healthcare AI systems reliably act in accordance with clinical goals, offering a practical framework for regulatory and ethical review.

論文PaperPapers/Benchmarks·arXiv cs.AI

VectorizationLLM: ベクトル化に基づくスマートAIアシスタントVectorizationLLM: Smart Vectorization Based AI Assistant

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本論文はベクトル化技術を活用したLLMベースのAIアシスタント手法を提案し、効率的な情報検索と応答生成の改善を目指している。

AI SUMMARYThis paper proposes VectorizationLLM, an AI assistant leveraging smart vectorization to enhance retrieval and response quality in large language model systems.

論文PaperPapers/Benchmarks·arXiv cs.AI

ストレートスルー引受におけるエージェント型AIと検索拡張モデルAgentic AI and Retrieval-Augmented Models in Straight-Through Underwriting

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本論文は、保険引受の完全自動化(ストレートスルー処理)にエージェント型AIとRAGを組み合わせる手法を提案し、意思決定の精度と説明可能性の向上を示した。

AI SUMMARYThis paper proposes combining agentic AI with retrieval-augmented generation for fully automated insurance underwriting, demonstrating improved decision accuracy and explainability in straight-through processing.

論文PaperPapers/Benchmarks·arXiv cs.AI

フィードバック操作正則化:模倣学習のためのオフラインエージェントアライメントFeedback Manipulation Regularization: Enabling Offline Agent Alignment for Imitation Learning

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約オフライン模倣学習においてフィードバック操作を正則化する手法を提案し、エージェントのアライメントを改善する。オンライン環境なしに安全で整合性の高い行動方策を学習できる点が重要。

AI SUMMARYThis paper proposes Feedback Manipulation Regularization (FMR) to align agents with desired behavior in offline imitation learning settings, removing the need for online interaction while improving policy robustness.

論文PaperPapers/Benchmarks·arXiv cs.AI

Nigeria Machinery: ドメイン根拠推論層を備えた低リソース産業データセットNigeria Machinery: A Low-Resource Industrial Dataset with a Domain-Grounded Reasoning Layer

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約ナイジェリアの機械産業を対象とした低リソースNLPデータセットを構築し、ドメイン知識に基づく推論層を導入することで、資源の乏しい産業分野におけるAI応用の課題に取り組んでいる。

AI SUMMARYThis paper introduces a low-resource industrial dataset focused on Nigerian machinery, augmented with a domain-grounded reasoning layer to improve AI performance in underrepresented industrial settings.

論文PaperPapers/Benchmarks·arXiv cs.AI

Persona Cartography: 重み空間における言語モデルの性格特性のマッピングPersona Cartography: Charting Language Model Personality Traits in Weight Space

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約言語モデルの性格特性がモデルの重み空間においてどのように分布・構造化されているかを体系的に調査した研究で、AIの行動制御や安全性に新たな知見をもたらす。

AI SUMMARYThis paper investigates how personality traits of language models are encoded in weight space, offering new methods to map and understand model behavior for better alignment and control.

論文PaperPapers/Benchmarks·arXiv cs.AI

エージェント型ニューラルアーキテクチャ探索Agentic Neural Architecture Search

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約LLMベースのエージェントを用いてニューラルネットワークのアーキテクチャ探索を自律的に行う手法を提案。人手によるデザイン工数を削減しつつ高性能なモデル構造を発見できる点が注目される。

AI SUMMARYThis paper proposes using LLM-driven agents to autonomously conduct neural architecture search, reducing manual design effort while discovering high-performing network structures across tasks.

論文PaperPapers/Benchmarks·arXiv cs.AI

プロンプトから契約へ:監査可能なエンタープライズLLMエージェントのためのハーネスエンジニアリングFrom Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約企業向けLLMエージェントの動作を検証・監査可能にする「ハーネスエンジニアリング」手法を提案し、プロンプト設計を形式的な契約として扱うことで信頼性とガバナンスを高める。

AI SUMMARYThis paper proposes harness engineering, a framework that treats LLM agent prompts as formal contracts to enable auditability and governance in enterprise deployments, improving reliability and accountability.

論文PaperPapers/Benchmarks·arXiv cs.AI

LLMが一致するとき、それは正しいのか?自己一貫性とモデル間合意を信頼度シグナルとして検証When LLMs Agree, Are They Right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約複数のLLMが同じ答えを出す場合や単一モデルが一貫した回答を示す場合、それが正確さの指標になるかを実証的に検証した研究。合意が信頼度シグナルとして有効かを明らかにし、AI出力の信頼性評価に示唆を与える。

AI SUMMARYThis paper empirically audits whether self-consistency within a single LLM and agreement across multiple LLMs reliably signal factual correctness, finding nuanced limits to using consensus as a confidence proxy.

論文PaperPapers/Benchmarks·arXiv cs.AI

説得攻撃によりCoTモニタリングの有効性が低下する可能性Persuasion Attacks Can Decrease Effectiveness of CoT Monitoring

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約説得的なプロンプト操作がChain-of-Thoughtの監視機構を欺き、AIの安全性検査を回避できることを示した研究。CoTベースの監視手法の脆弱性として重要な警鐘となる。

AI SUMMARYThis research demonstrates that persuasion-based prompt attacks can undermine Chain-of-Thought monitoring, causing safety oversight mechanisms to miss harmful model behavior. The findings highlight a critical vulnerability in CoT-based AI supervision.

Fri, Jul 1015 papers
論文PaperPapers/Benchmarks·arXiv cs.AI

AgentLens: コーディングエージェント評価のための本番環境ベースのトラジェクトリレビューAgentLens: Production-Assessed Trajectory Reviews for Coding Agent Evaluation

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約AgentLensはコーディングエージェントの動作軌跡を本番環境の基準で評価する新しいフレームワークで、エージェントの実用的な性能をより正確に測定できる。

AI SUMMARYAgentLens introduces a trajectory review framework for coding agents that uses production-grounded assessment, enabling more realistic and reliable evaluation of agent behavior beyond traditional benchmarks.

論文PaperPapers/Benchmarks·arXiv cs.AI

文脈内探索はいつ有効か?リフレクション駆動推論のサンプリング複雑性理論When Does In-Context Search Help? A Sampling-Complexity Theory of Reflection-Driven Reasoning

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本論文は、LLMが推論中に自己反省・再探索を行う「リフレクション」の有効条件をサンプリング複雑性の観点から理論的に解析し、どのような問題設定で計算コストに見合う恩恵が得られるかを明らかにする。

AI SUMMARYThis paper develops a sampling-complexity theory to formally characterize when in-context search and reflection-driven reasoning improve LLM performance, offering principled guidance on the conditions under which iterative self-reflection is worth its computational cost.

論文PaperPapers/Benchmarks·arXiv cs.AI

エージェントベースモデリングにおけるLLMを活用した推論LLM-powered reasoning in agent-based modeling

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約LLMをエージェントベースモデルに組み込むことで、エージェントの意思決定に高度な推論能力を付与する手法を提案。社会シミュレーションや複雑系研究の表現力向上に貢献する。

AI SUMMARYThis paper proposes integrating large language models into agent-based modeling to enable human-like reasoning in simulated agents, improving the fidelity of social and complex-systems simulations.

論文PaperPapers/Benchmarks·arXiv cs.AI

計算・実験数学における SageMath 拡張 LLM エージェントの評価Evaluating SageMath-Augmented LLM Agents for Computational and Experimental Mathematics

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約SageMath を統合した LLM エージェントが計算数学・実験数学タスクをどの程度解けるかを体系的に評価した研究。数式処理システムとの連携がLLMの数学的推論能力を大幅に向上させることを示した。

AI SUMMARYThis paper systematically benchmarks LLM agents augmented with SageMath on computational and experimental mathematics tasks, showing that tool-integrated agents significantly outperform bare LLMs on complex mathematical problems.

論文PaperPapers/Benchmarks·arXiv cs.AI

「ハーネス効果」:オーケストレーション設計がエンタープライズ向けエージェントAIのトークン経済学を左右するThe Harness Effect: How Orchestration Design Sets the Token Economics of Enterprise Agentic AI

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約エージェントAIシステムにおけるオーケストレーション層の設計が、トークン消費量とコスト構造に直接影響することを実証した研究。企業導入における費用対効果の最適化に重要な示唆を与える。

AI SUMMARYThis paper demonstrates that the design of orchestration harnesses in agentic AI systems directly governs token consumption and cost structures, coining the term "Harness Effect." The findings offer actionable guidance for enterprises seeking to optimize the economics of large-scale AI agent deployments.

論文PaperPapers/Benchmarks·arXiv cs.CL

ソルバーから研究へ:LLM駆動の形式数学が研究フロンティアに到達From Solvers to Research: Large Language Model-Driven Formal Mathematics at the Research Frontier

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約大規模言語モデルが数学の問題を解くだけでなく、未解決問題の探索や定理証明など本格的な数学研究を支援できる段階に達しつつあることを論じたサーベイ論文。形式数学とLLMの融合が数学研究の在り方を変える可能性を示す。

AI SUMMARYThis survey paper examines how LLMs are advancing beyond competition-style problem solving into genuine mathematical research, including formal theorem proving and exploration of open problems, signaling a shift in how AI can contribute to frontier mathematics.

論文PaperPapers/Benchmarks·arXiv cs.CL

DeepSearch-World: 検証可能な環境における深層検索エージェントの自己蒸留DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約検証可能な環境でエージェントが自己蒸留により深層検索能力を向上させる新手法を提案。強化学習なしに検索精度を高められる点が注目される。

AI SUMMARYDeepSearch-World proposes a self-distillation framework that trains deep search agents in a verifiable environment, improving search accuracy without relying on reinforcement learning.

論文PaperPapers/Benchmarks·arXiv cs.CL

人間とLLMの協働による拡張可能な文化固有のステレオタイプデータセット構築Scalable and Culturally Specific Stereotype Dataset Construction via Human-LLM Collaboration

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約LLMと人間のアノテーターを組み合わせ、文化ごとに異なるステレオタイプを効率的に収集・構築する手法を提案。AIの公平性評価に役立つ多様なデータセット作成を可能にする。

AI SUMMARYThis paper proposes a scalable pipeline combining human annotators and LLMs to build culturally specific stereotype datasets, enabling more representative fairness evaluations for AI systems across diverse cultures.

論文PaperPapers/Benchmarks·arXiv cs.CL

非現実的なトークンが強化されるとき:LLM強化学習のためのテール考慮クレジット調整When Implausible Tokens Get Reinforced: Tail-Aware Credit Calibration for LLM Reinforcement Learning

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約強化学習でLLMを訓練する際、低確率トークンが誤って強化される問題を指摘し、テール分布を考慮したクレジット調整手法を提案。報酬の帰属精度を高めることでモデルの品質向上を図る。

AI SUMMARYThis paper identifies that low-probability tokens can be incorrectly reinforced during LLM RL training, and proposes a tail-aware credit calibration method to more accurately assign reward signals, improving overall model quality.

論文PaperPapers/Benchmarks·arXiv cs.CL

全二重音声エージェント向けLALMオーディオ審判の信頼性評価A Reliability Assessment of LALM Audio Judges for Full-Duplex Voice Agents

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本研究は、全二重音声エージェントの評価に用いられるLALMベースの自動審判モデルの信頼性を体系的に検証し、その一致度や偏りを明らかにした。自動評価の限界を把握することで、より堅牢な音声AIの品質測定が可能になる。

AI SUMMARYThis study systematically evaluates the reliability of large audio language model (LALM) judges used to assess full-duplex voice agents, revealing consistency and bias issues. The findings help establish more trustworthy automated evaluation pipelines for conversational speech AI.

論文PaperPapers/Benchmarks·arXiv cs.CL

Hallucination Self-Play: 進化した生成器による強化検出器のブートストラップHallucination Self-Play: Bootstrapping Reinforced Detector via Evolved Generator

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約LLMの幻覚検出を改善するため、生成器と検出器が自己対戦的に互いを強化し合うフレームワークを提案。ラベル付きデータなしで高精度な幻覚検出を実現できる点が重要。

AI SUMMARYThis paper proposes a self-play framework where a hallucination generator and detector iteratively improve each other, enabling robust hallucination detection without relying on expensive labeled data.

論文PaperPapers/Benchmarks·arXiv cs.CL

実行から教育へ:LLMにおける教育的制御を測定するBloom準拠フレームワークFrom Execution to Education: A Bloom-Aligned Framework for Measuring Educational Control in LLMs

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約ブルームの分類法に基づき、LLMが教育的文脈でどの程度学習者の認知レベルを制御できるかを定量評価するフレームワークを提案。AIチュータリングの品質保証に新たな指標をもたらす。

AI SUMMARYThis paper proposes a framework aligned with Bloom's Taxonomy to measure how well LLMs can control educational outputs across cognitive levels, enabling more rigorous evaluation of AI tutoring systems.

論文PaperPapers/Benchmarks·arXiv cs.CL

低遅延システムにおけるツール生成と自己進化型LLMエージェントTool-Making and Self-Evolving LLM Agents in Low-Latency Systems

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約LLMエージェントが自らツールを作成・改良しながら低遅延環境で動作する手法を提案し、エージェントの自律的な能力拡張と応答速度の両立を実現した研究。

AI SUMMARYThis paper proposes a framework where LLM agents autonomously create and refine tools while operating under low-latency constraints, enabling self-improvement without sacrificing response speed.

論文PaperPapers/Benchmarks·arXiv cs.CL

LLMの論理は信頼できるか?グラフベースのフレームワークで不確実性・一貫性・頑健性を定量化Can We Trust LLM's Logic? Quantifying Uncertainty, Coherence, and Robustness via a Graph-Based Framework

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本研究はグラフ構造を用いてLLMの推論における不確実性、論理的一貫性、入力変動への頑健性を定量的に評価する手法を提案する。LLMの信頼性を客観的に測る基盤として重要な貢献となる。

AI SUMMARYThis paper proposes a graph-based framework to quantify uncertainty, logical coherence, and robustness in LLM reasoning, enabling more objective evaluation of whether LLM outputs can be trusted for critical tasks.

論文PaperPapers/Benchmarks·arXiv cs.CL

PLURAL: 価値アラインメントのためのグローバルデータセットPLURAL: A Global Dataset for Value Alignment

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約PLURALは多文化・多言語にわたる価値観の違いを捉えた大規模データセットで、AIの価値アラインメント研究に多様な視点を提供する。単一文化に偏ったベンチマークの限界を補う点で重要。

AI SUMMARYPLURAL is a large-scale multilingual dataset capturing diverse human values across cultures, designed to improve value alignment in AI systems. It addresses the critical gap left by culturally homogeneous benchmarks.

Thu, Jul 93 papers
論文PaperPapers/Benchmarks·arXiv cs.AI

社会規範の学習が動的な人間とAIの協調における互換性を高めるLearning social norms enhances compatibility in dynamic human-AI coordination

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約AIエージェントが社会規範を学習することで、初対面の人間パートナーとの動的な協調タスクにおける適応性と互換性が向上することを示した研究。人間とAIの協働設計に新たな指針を提供する。

AI SUMMARYThis research shows that AI agents trained to learn social norms achieve better compatibility with unfamiliar human partners in dynamic coordination tasks, offering new design principles for human-AI teaming.

論文PaperPapers/Benchmarks·arXiv cs.AI

マルチエージェントLLM安全性における「操作的リフレーミング」と「承認フレーム委譲」Operational Reframing and Approval-Framed Delegation in Multi-Agent LLM Safety

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約複数のLLMエージェントが連携する環境で、タスクの表現を巧みに書き換えたり承認済みとして偽装することで安全制約を回避できる脆弱性を分析した研究。マルチエージェント設計における信頼伝播の危険性を示す点で重要。

AI SUMMARYThis paper analyzes how adversarial prompts can bypass safety constraints in multi-agent LLM systems by reframing tasks or falsely presenting delegated actions as pre-approved, exposing critical trust-propagation risks in agentic pipelines.

論文PaperPapers/Benchmarks·arXiv cs.AI

推論一貫性スキャン:AI安全性評価におけるChain-of-Thought妥当性監査フレームワークReasoning Consistency Scanning: A Framework for Auditing Chain-of-Thought Validity in AI Safety Evaluations

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約AIの思考連鎖(Chain-of-Thought)推論の一貫性を体系的に監査するフレームワークを提案し、安全性評価における推論の欠陥や矛盾を検出する手法を示した研究。信頼性の高いAI安全評価の実現に貢献する。

AI SUMMARYThis paper proposes a framework for systematically auditing chain-of-thought reasoning in AI safety evaluations, detecting logical inconsistencies and flawed reasoning steps. It matters because reliable safety assessments depend on valid reasoning chains.