HomeTags#paperPage 2

Tag timeline

#paperpage 2/3

同じキーワードで束ねられた更新の続きです。カテゴリをまたいだ関連ニュースや実装トピックの追跡に使えます。

Total80#paper の全掲載記事All listed entries tagged #paper
Showing30このページの表示件数Entries on this page
Page2/3静的ページ位置Static page position
Updated公開index snapshotPublished index snapshot

Entriespage 2/3 · 80 total

Tue, Jul 1410 entries
論文PaperPapers/Benchmarks·arXiv cs.CL

合意と反対意見:グループ推薦における主観的嗜好の動的LLMモデリングConsensus vs. Dissent: Dynamic LLM Modeling of Subjective Preferences in Group Recommenders

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本研究はグループ推薦システムにおいて、LLMを用いてメンバー間の合意と反対意見を動的にモデル化する手法を提案する。個人の主観的嗜好を柔軟に捉えることで、グループ全体の満足度向上を目指す。

AI SUMMARYThis paper proposes a dynamic LLM-based framework for group recommender systems that models both consensus and dissent among members' subjective preferences, improving collective satisfaction beyond simple preference aggregation.

Consensus vs. Dissent: Dynamic LLM Modeling of Subjective Preferences in Group Recommendersog
論文PaperPapers/Benchmarks·arXiv cs.CL

Index SLM テクニカルレポートIndex SLM Technical Report

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約Index SLMは、限られたパラメータ数で高い性能を実現する小型言語モデルシリーズであり、効率的なエッジ・オンデバイス推論への応用が期待される。

AI SUMMARYIndex SLM introduces a series of small language models designed to achieve competitive performance at reduced parameter counts, enabling practical deployment in edge and on-device scenarios.

論文PaperPapers/Benchmarks·arXiv cs.CL

RouteRec: 推薦エージェントの選択と集約に関する厳密な評価フレームワークRouteRec: Strict Evaluation of Recommender-Agent Selection and Aggregation

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約RouteRecは、複数の推薦エージェントをどう選択・集約するかを厳密に評価するベンチマークを提案し、エージェント間のルーティング戦略の有効性を体系的に測定できる点で重要。

AI SUMMARYRouteRec introduces a rigorous benchmark for evaluating how recommender agents are selected and aggregated, enabling systematic measurement of routing strategies across multiple agents.

論文PaperPapers/Benchmarks·arXiv cs.CL

言語モデルによるグローバルM&Aアービトラージ予測Global Merger-Arbitrage Forecasting with Language Models

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約大規模言語モデルをM&Aアービトラージの取引成否予測に応用し、財務テキストから有益なシグナルを抽出できることを示した研究。投資戦略への自然言語処理活用の可能性を広げる成果として注目される。

AI SUMMARYThis paper applies large language models to predict deal outcomes in merger-arbitrage investing, showing that NLP signals from financial text meaningfully improve forecasting accuracy across global markets.

論文PaperPapers/Benchmarks·arXiv cs.CL

忠実性を設計で担保:多様なステークホルダー向けLLM生成臨床試験サマリーの評価と改善Faithful by Design: Evaluating and Improving LLM-Generated Clinical Trial Summaries for Multi-Stakeholder Audiences

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約LLMが生成する臨床試験サマリーの事実忠実性を患者・医療者・研究者など複数の読者層に合わせて評価・改善する手法を提案。医療情報の誤りがもたらすリスクを低減する実用的な枠組みとして意義がある。

AI SUMMARYThis paper proposes methods to evaluate and improve the factual faithfulness of LLM-generated clinical trial summaries tailored to diverse audiences, including patients and clinicians. It addresses a critical safety concern by reducing hallucinations in high-stakes medical communication.

論文PaperPapers/Benchmarks·arXiv cs.CL

デバイス上でのリアルタイム字幕翻訳に向けたワークロード駆動最適化Workload-Driven Optimization for On-Device Real-Time Subtitle Translation

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本論文は、オンデバイス環境でリアルタイム字幕翻訳を実現するため、ワークロードの特性に基づいてモデルや処理を動的に最適化する手法を提案する。これにより、限られた計算資源でも低遅延かつ高品質な翻訳が可能になる。

AI SUMMARYThis paper proposes a workload-driven optimization framework for real-time subtitle translation running entirely on-device, dynamically adapting model execution to meet latency constraints without sacrificing translation quality.

論文PaperPapers/Benchmarks·arXiv cs.CL

量子化LLM推論におけるサイレント障害:「中空収束」と障害モードシフトの分類論的分析Silent Failures in Quantized LLM Reasoning: A Taxonomy-Based Analysis of Hollow Convergence and Failure Mode Shifts

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約量子化されたLLMが表面上は正しく見える回答を生成しながら内部推論が破綻する「中空収束」現象を分類・分析し、量子化が引き起こす障害モードの質的変化を明らかにした研究。

AI SUMMARYThis paper identifies and classifies "hollow convergence" in quantized LLMs—where models produce plausible-looking outputs while reasoning has silently broken down—revealing systematic failure mode shifts that standard benchmarks fail to detect.

論文PaperPapers/Benchmarks·arXiv cs.CL

シンガポールの言語環境に合わせた音声言語モデルの効率的な適応Efficiently Adapting Spoken Language Models for the Singaporean Context

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本研究は、シンガポール英語(Singlish)などの多言語混在環境に対応するため、既存の音声言語モデルを効率的にファインチューニングする手法を提案し、限られたリソースでの高精度な音声認識・理解を実現する。

AI SUMMARYThis paper proposes efficient adaptation methods for spoken language models targeting Singapore's multilingual context, achieving strong performance on Singlish and code-switching speech without requiring large-scale retraining.

論文PaperPapers/Benchmarks·arXiv cs.CL

非英語言語における推論コスト:日本語を事例とした研究Cost of Reasoning in non-English Languages: A Case Study on Japanese

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約推論型LLMが日本語などの非英語言語でタスクを処理する際、英語より大幅に多くのトークンを消費することを実証した研究。多言語展開におけるコストと効率の課題を明らかにしている。

AI SUMMARYThis paper demonstrates that reasoning LLMs consume significantly more tokens when processing non-English languages like Japanese compared to English, revealing hidden cost and efficiency disparities in multilingual deployments.

論文PaperPapers/Benchmarks·arXiv cs.CL

精度は同じ、証拠は不平等:ツール利用エージェントの意思決定面としての検索APIEqual Accuracy, Unequal Evidence: Search APIs as Decision Surfaces for Tool-Using Agents

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約検索APIが同程度の精度を示しても、返却される証拠の質や多様性に大きな差があり、ツール利用エージェントの意思決定に偏りをもたらすことを明らかにした研究。

AI SUMMARYThis paper shows that search APIs with similar accuracy can differ substantially in evidence quality and diversity, introducing hidden biases into tool-using agents' decisions.

Mon, Jul 1310 entries
論文PaperPapers/Benchmarks·arXiv cs.SE

プログラマーはLLMが生成したアサーションの評価が苦手で過信しがちProgrammers Are Poor and Overconfident Judges of LLM-Generated Assertions

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約LLMが生成したテストアサーションをプログラマーが評価する際、正確性を過大評価しミスを見落としやすいことが実験で示された。自動テスト生成の品質保証に人間のレビューだけでは不十分である可能性を示唆する。

AI SUMMARYA study found that programmers systematically overestimate the correctness of LLM-generated test assertions and miss significant errors, raising concerns about relying on human review as a quality gate for AI-generated tests.

Programmers Are Poor and Overconfident Judges of LLM-Generated Assertionsog
論文PaperPapers/Benchmarks·arXiv cs.SE

より良いハーネス、小さなモデル:自動ハーネス適応で90%コスト削減エージェントの構築Better Harnesses, Smaller Models: Building 90% Cheaper Agents via Automated Harness Adaptation

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約コーディングエージェントのテストハーネスを自動的に最適化することで、大型モデルに依存せず小型モデルでも高い性能を実現し、運用コストを約90%削減できることを示した研究。

AI SUMMARYThis paper shows that automatically adapting test harnesses for coding agents allows smaller, cheaper models to match large-model performance, cutting agent operational costs by roughly 90%.

論文PaperPapers/Benchmarks·arXiv cs.SE

LLM生成コードにおける「パッチワーク問題」The Patchwork Problem in LLM-Generated Code

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約LLMが生成するコードは断片的な知識を継ぎ接ぎした構造になりやすく、一貫性や保守性に欠けるという問題を論文が指摘している。この知見はAIコード生成ツールの評価・改善指針として重要な意味を持つ。

AI SUMMARYResearchers identify the "patchwork problem" in LLM-generated code, where outputs are stitched together from disparate training patterns, leading to inconsistency and poor maintainability. This has significant implications for how AI coding tools should be evaluated and improved.

論文PaperPapers/Benchmarks·arXiv cs.SE

SCATE: コスト効率の高いテスト生成のためにコーディングエージェントを監督する学習SCATE: Learning to Supervise Coding Agents for Cost-Effective Test Generation

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約SCATEはLLMベースのコーディングエージェントを監督者モデルで制御し、テスト生成コストを抑えながら品質を維持する手法を提案する。監督者がエージェントの行動を動的に評価することで、効率的なソフトウェアテスト自動化を実現する。

AI SUMMARYSCATE proposes training a supervisor model to guide LLM-based coding agents during automated test generation, reducing computational cost while maintaining coverage quality. This approach makes agent-driven software testing more practical for real-world use.

論文PaperPapers/Benchmarks·arXiv cs.SE

汎用から個別へ:ペルソナを考慮したコードレビュー説明の探求From Generic to Personalized: Exploring Persona-Aware Code Review Explanations

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本研究は、開発者の経験レベルや役割に応じてコードレビューのフィードバック説明を個別化するペルソナ対応アプローチを提案し、画一的な説明の限界を克服しようとしている。

AI SUMMARYThis paper proposes a persona-aware approach to generating code review explanations tailored to a developer's experience and role, showing that personalized feedback improves comprehension over generic explanations.

論文PaperPapers/Benchmarks·arXiv cs.SE

Bugs4Qを用いたQiskitプログラム修復におけるLLMのベンチマーク評価Benchmarking Large Language Models on Repairing Qiskit Programs using Bugs4Q

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約量子プログラムのバグ修復タスクにLLMを適用し、Bugs4Qベンチマークで性能を評価した研究。量子ソフトウェア開発における自動修復の可能性と限界を明らかにしている。

AI SUMMARYThis study evaluates large language models on automatically repairing buggy Qiskit quantum programs using the Bugs4Q benchmark, revealing both the promise and current limitations of LLM-based repair for quantum software.

論文PaperPapers/Benchmarks·arXiv cs.SE

Pythonの型アノテーションをジャストインタイムで自動更新する手法Automating Just-In-Time Python Type Annotation Updating

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約コード変更に伴い古くなったPythonの型アノテーションを自動的に検出・更新する手法を提案し、保守コストの削減と型安全性の維持を両立する。

AI SUMMARYThis paper proposes an automated approach to detect and update stale Python type annotations triggered by code changes, reducing maintenance burden while preserving type safety.

論文PaperPapers/Benchmarks·arXiv cs.SE

スキルマーケットの内側:ソフトウェアエンジニアリング活動から再利用可能なエージェントスキルへInside the Skill Market: From Software Engineering Activities to Reusable Agent Skills

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約ソフトウェアエンジニアリングのタスクをエージェントが再利用可能なスキルとして体系化する「スキルマーケット」フレームワークを提案し、エージェントの汎化性能と効率を向上させる研究。

AI SUMMARYThis paper proposes a "Skill Market" framework that distills software engineering activities into reusable agent skills, enabling more generalizable and efficient AI agents for SE tasks.

論文PaperPapers/Benchmarks·arXiv cs.SE

データ集約型コンピューティングにおけるプロパティテンプレートを用いたエージェント的証明とプロパティベーステストAgentic Proof and Property-Based Testing via Property-Templates in Data-Intensive Computing

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本論文は、データ集約型システムの検証にエージェントAIとプロパティテンプレートを組み合わせ、形式的証明とプロパティベーステストを自動化する手法を提案する。これにより複雑なデータ処理コードの信頼性検証コストを大幅に削減できる。

AI SUMMARYThis paper proposes using AI agents with reusable property-templates to automate formal proofs and property-based testing in data-intensive computing, reducing the manual effort required to verify correctness of complex data pipelines.

論文PaperPapers/Benchmarks·arXiv cs.SE

人間のテスト工程に着想を得たワークフローによるユニットテスト自動生成のためのマルチエージェントLLM協調Multi-Agent LLM Collaboration for Unit Test Generation via Human-Testing-Inspired Workflows

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約複数のLLMエージェントが人間のテスト設計プロセスを模倣して協調することで、ユニットテストの品質とカバレッジを向上させる手法を提案した研究。従来の単一モデルによる生成より効果的なテスト作成が可能になる。

AI SUMMARYThis paper proposes a multi-agent LLM framework that mimics human software testing workflows to collaboratively generate higher-quality unit tests with improved coverage, outperforming single-model approaches.

Sat, Jul 1110 entries
論文PaperPapers/Benchmarks·arXiv cs.AI

プロアクティブなエンタープライズエージェントのためのコンテキストグラフContext Graphs for Proactive Enterprise Agents

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本論文は、企業向けAIエージェントが自律的に行動するためのコンテキストグラフという新しい手法を提案し、エージェントが適切なタイミングで先回りして行動できる仕組みを示している。

AI SUMMARYThis paper proposes context graphs as a structured representation to enable proactive enterprise AI agents, allowing them to anticipate user needs and act autonomously at the right moment.

Context Graphs for Proactive Enterprise Agentsog
論文PaperPapers/Benchmarks·arXiv cs.AI

人間とLLMの混成集団に向けた対立的社会認識論Adversarial Social Epistemology for Assemblies of Humans and Large Language Models

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約人間とLLMが混在する集合的意思決定の場で、悪意ある操作や認識論的攻撃がどう機能するかを分析した研究。AIを含む社会的知識形成の堅牢性設計に重要な示唆を与える。

AI SUMMARYThis paper analyzes how adversarial actors can exploit mixed human-LLM assemblies to distort collective knowledge and decision-making, offering a framework for building more robust epistemic systems that include AI participants.

論文PaperPapers/Benchmarks·arXiv cs.AI

アライメント妥当性:ヘルスケアにおけるAI保証の新基準Alignment Plausibility: A New Standard for Assuring AI in Healthcare

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本論文は医療AIの安全性を評価する新概念「アライメント妥当性」を提案し、AIの挙動が臨床目標と一致しているかを体系的に検証する枠組みを示す。規制や倫理審査に応用可能な実用的基準として注目される。

AI SUMMARYThis paper proposes 'alignment plausibility' as a new standard for evaluating whether healthcare AI systems reliably act in accordance with clinical goals, offering a practical framework for regulatory and ethical review.

論文PaperPapers/Benchmarks·arXiv cs.AI

VectorizationLLM: ベクトル化に基づくスマートAIアシスタントVectorizationLLM: Smart Vectorization Based AI Assistant

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本論文はベクトル化技術を活用したLLMベースのAIアシスタント手法を提案し、効率的な情報検索と応答生成の改善を目指している。

AI SUMMARYThis paper proposes VectorizationLLM, an AI assistant leveraging smart vectorization to enhance retrieval and response quality in large language model systems.

論文PaperPapers/Benchmarks·arXiv cs.AI

ストレートスルー引受におけるエージェント型AIと検索拡張モデルAgentic AI and Retrieval-Augmented Models in Straight-Through Underwriting

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本論文は、保険引受の完全自動化(ストレートスルー処理)にエージェント型AIとRAGを組み合わせる手法を提案し、意思決定の精度と説明可能性の向上を示した。

AI SUMMARYThis paper proposes combining agentic AI with retrieval-augmented generation for fully automated insurance underwriting, demonstrating improved decision accuracy and explainability in straight-through processing.

論文PaperPapers/Benchmarks·arXiv cs.AI

フィードバック操作正則化:模倣学習のためのオフラインエージェントアライメントFeedback Manipulation Regularization: Enabling Offline Agent Alignment for Imitation Learning

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約オフライン模倣学習においてフィードバック操作を正則化する手法を提案し、エージェントのアライメントを改善する。オンライン環境なしに安全で整合性の高い行動方策を学習できる点が重要。

AI SUMMARYThis paper proposes Feedback Manipulation Regularization (FMR) to align agents with desired behavior in offline imitation learning settings, removing the need for online interaction while improving policy robustness.

論文PaperPapers/Benchmarks·arXiv cs.AI

Nigeria Machinery: ドメイン根拠推論層を備えた低リソース産業データセットNigeria Machinery: A Low-Resource Industrial Dataset with a Domain-Grounded Reasoning Layer

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約ナイジェリアの機械産業を対象とした低リソースNLPデータセットを構築し、ドメイン知識に基づく推論層を導入することで、資源の乏しい産業分野におけるAI応用の課題に取り組んでいる。

AI SUMMARYThis paper introduces a low-resource industrial dataset focused on Nigerian machinery, augmented with a domain-grounded reasoning layer to improve AI performance in underrepresented industrial settings.

論文PaperPapers/Benchmarks·arXiv cs.AI

Persona Cartography: 重み空間における言語モデルの性格特性のマッピングPersona Cartography: Charting Language Model Personality Traits in Weight Space

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約言語モデルの性格特性がモデルの重み空間においてどのように分布・構造化されているかを体系的に調査した研究で、AIの行動制御や安全性に新たな知見をもたらす。

AI SUMMARYThis paper investigates how personality traits of language models are encoded in weight space, offering new methods to map and understand model behavior for better alignment and control.

論文PaperPapers/Benchmarks·arXiv cs.AI

エージェント型ニューラルアーキテクチャ探索Agentic Neural Architecture Search

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約LLMベースのエージェントを用いてニューラルネットワークのアーキテクチャ探索を自律的に行う手法を提案。人手によるデザイン工数を削減しつつ高性能なモデル構造を発見できる点が注目される。

AI SUMMARYThis paper proposes using LLM-driven agents to autonomously conduct neural architecture search, reducing manual design effort while discovering high-performing network structures across tasks.

論文PaperPapers/Benchmarks·arXiv cs.AI

プロンプトから契約へ:監査可能なエンタープライズLLMエージェントのためのハーネスエンジニアリングFrom Prompts to Contracts: Harness Engineering for Auditable Enterprise LLM Agents

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約企業向けLLMエージェントの動作を検証・監査可能にする「ハーネスエンジニアリング」手法を提案し、プロンプト設計を形式的な契約として扱うことで信頼性とガバナンスを高める。

AI SUMMARYThis paper proposes harness engineering, a framework that treats LLM agent prompts as formal contracts to enable auditability and governance in enterprise deployments, improving reliability and accountability.