HomeTags#llmPage 7

Tag timeline

#llmpage 7/9

同じキーワードで束ねられた更新の続きです。カテゴリをまたいだ関連ニュースや実装トピックの追跡に使えます。

Total265#llm の全掲載記事All listed entries tagged #llm
Showing30このページの表示件数Entries on this page
Page7/9静的ページ位置Static page position
Updated公開index snapshotPublished index snapshot

Entriespage 7/9 · 265 total

Tue, Jul 1414 entries
論文PaperPapers/Benchmarks·arXiv cs.AI

忠実であって修正はしない:マルチホップエージェントリレーにおけるメッセージ形式の影響はティアに依存するFaithful, Not Corrective: Message-Format Effects in Multi-Hop Agent Relays Are Tier-Dependent

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約マルチホップエージェントリレーでは、メッセージの形式が下流エージェントの動作に与える影響がエージェントの階層によって異なり、上流エージェントは誤りを修正せず忠実に伝達することが示された。複数エージェント系の設計における信頼性評価に重要な知見を提供する。

AI SUMMARYThis study finds that message-format effects in multi-hop agent relay chains are tier-dependent: agents faithfully propagate upstream content rather than correcting errors, with implications for reliability in multi-agent system design.

論文PaperPapers/Benchmarks·arXiv cs.AI

潜在的CoT推論を動的システムとして解釈するInterpreting Latent CoT Reasoning as Dynamical Systems

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本研究はLLMの潜在空間におけるChain-of-Thought推論を動的システムの観点から分析し、推論過程の内部構造を数理的に解明する手法を提案する。推論メカニズムの解釈可能性向上に貢献する。

AI SUMMARYThis paper proposes a dynamical-systems framework for analyzing latent Chain-of-Thought reasoning in LLMs, offering a principled mathematical lens to understand how intermediate reasoning steps evolve in hidden states. It advances interpretability of complex multi-step reasoning.

論文PaperPapers/Benchmarks·arXiv cs.SE

AfterVibe: 会話が終わった後に何が残るかAfterVibe: What Remains When the Conversation Ends

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本研究はAIとの対話終了後にユーザーの感情や認知にどのような影響が持続するかを分析し、チャットシステム設計における「余韻」の重要性を示した。

AI SUMMARYThis paper examines the emotional and cognitive residues that persist after human-AI conversations end, highlighting design implications for conversational systems that account for post-interaction effects.

論文PaperPapers/Benchmarks·arXiv cs.SE

AIエージェントが書いたコードはマージ後どうなるか?その追跡調査Do These Violent Delights Have Violent Ends? Measuring the Post-Merge Fate of Agentic Code

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本論文はAIエージェントが生成しマージされたコードのその後の運命を実証的に調査し、品質や保守性への長期的影響を定量化した研究である。エージェント生成コードの実用上のリスクを把握する上で重要な知見を提供する。

AI SUMMARYThis paper empirically tracks the post-merge lifecycle of code produced by AI coding agents, measuring its long-term quality, churn, and maintenance burden compared to human-written code. The findings inform real-world risk assessments of deploying agentic coding systems.

論文PaperPapers/Benchmarks·arXiv cs.SE

LLMを用いた静的解析アラートの判定とエラー低減技術Using LLMs to Adjudicate Static-Analysis Alerts with Error Reduction Techniques

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本論文は、静的解析ツールが生成する大量の誤検知アラートをLLMで自動判定し、エラー低減技術を組み合わせることで精度を高める手法を提案している。開発者の負担軽減とセキュリティ品質向上に貢献する研究成果である。

AI SUMMARYThis paper proposes using LLMs to automatically triage static-analysis alerts—distinguishing true bugs from false positives—while applying error reduction techniques to improve adjudication accuracy and reduce developer burden.

論文PaperPapers/Benchmarks·arXiv cs.CL

合意と反対意見:グループ推薦における主観的嗜好の動的LLMモデリングConsensus vs. Dissent: Dynamic LLM Modeling of Subjective Preferences in Group Recommenders

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本研究はグループ推薦システムにおいて、LLMを用いてメンバー間の合意と反対意見を動的にモデル化する手法を提案する。個人の主観的嗜好を柔軟に捉えることで、グループ全体の満足度向上を目指す。

AI SUMMARYThis paper proposes a dynamic LLM-based framework for group recommender systems that models both consensus and dissent among members' subjective preferences, improving collective satisfaction beyond simple preference aggregation.

Consensus vs. Dissent: Dynamic LLM Modeling of Subjective Preferences in Group Recommendersog
論文PaperPapers/Benchmarks·arXiv cs.CL

言語モデルによるグローバルM&Aアービトラージ予測Global Merger-Arbitrage Forecasting with Language Models

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約大規模言語モデルをM&Aアービトラージの取引成否予測に応用し、財務テキストから有益なシグナルを抽出できることを示した研究。投資戦略への自然言語処理活用の可能性を広げる成果として注目される。

AI SUMMARYThis paper applies large language models to predict deal outcomes in merger-arbitrage investing, showing that NLP signals from financial text meaningfully improve forecasting accuracy across global markets.

論文PaperPapers/Benchmarks·arXiv cs.CL

忠実性を設計で担保:多様なステークホルダー向けLLM生成臨床試験サマリーの評価と改善Faithful by Design: Evaluating and Improving LLM-Generated Clinical Trial Summaries for Multi-Stakeholder Audiences

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約LLMが生成する臨床試験サマリーの事実忠実性を患者・医療者・研究者など複数の読者層に合わせて評価・改善する手法を提案。医療情報の誤りがもたらすリスクを低減する実用的な枠組みとして意義がある。

AI SUMMARYThis paper proposes methods to evaluate and improve the factual faithfulness of LLM-generated clinical trial summaries tailored to diverse audiences, including patients and clinicians. It addresses a critical safety concern by reducing hallucinations in high-stakes medical communication.

論文PaperPapers/Benchmarks·arXiv cs.CL

量子化LLM推論におけるサイレント障害:「中空収束」と障害モードシフトの分類論的分析Silent Failures in Quantized LLM Reasoning: A Taxonomy-Based Analysis of Hollow Convergence and Failure Mode Shifts

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約量子化されたLLMが表面上は正しく見える回答を生成しながら内部推論が破綻する「中空収束」現象を分類・分析し、量子化が引き起こす障害モードの質的変化を明らかにした研究。

AI SUMMARYThis paper identifies and classifies "hollow convergence" in quantized LLMs—where models produce plausible-looking outputs while reasoning has silently broken down—revealing systematic failure mode shifts that standard benchmarks fail to detect.

論文PaperPapers/Benchmarks·arXiv cs.CL

非英語言語における推論コスト:日本語を事例とした研究Cost of Reasoning in non-English Languages: A Case Study on Japanese

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約推論型LLMが日本語などの非英語言語でタスクを処理する際、英語より大幅に多くのトークンを消費することを実証した研究。多言語展開におけるコストと効率の課題を明らかにしている。

AI SUMMARYThis paper demonstrates that reasoning LLMs consume significantly more tokens when processing non-English languages like Japanese compared to English, revealing hidden cost and efficiency disparities in multilingual deployments.

コミュニティCommunityLocal Models·Qiita LLM

【AWS】Gemma 4をセルフホスティングしてみた〜クラッシュを回避するインスタンス選定とメモリのリアル〜A practical guide to self-hosting Gemma 4 on AWS, covering how to choose the…

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約AWSでGemma 4をセルフホスティングする際に発生するクラッシュを回避するため、適切なEC2インスタンス選定とメモリ管理の実践的な知見をまとめた記事。コスト効率と安定稼働を両立するための具体的な手順が参考になる。

AI SUMMARYA practical guide to self-hosting Gemma 4 on AWS, covering how to choose the right EC2 instance to avoid OOM crashes and manage memory effectively for stable inference.

【AWS】Gemma 4をセルフホスティングしてみた〜クラッシュを回避するインスタンス選定とメモリのリアル〜og
報道NewsNews/Policy·TechCrunch

HermesエージェントメーカーのNous Research、15億ドル評価額での新規資金調達を協議中Hermes agent maker Nous Research in talks for new funding at $1.5B valuation

重要度 MediumMedium priority技術記事 · Industry & Policytechnical post · Industry & Policy

AI要約オープンウェイトLLM「Hermes」シリーズで知られるNous Researchが、15億ドルの評価額で新たな資金調達を交渉中。AIエージェント分野への投資熱が続く中、独立系研究機関としての存在感が高まっている。

AI SUMMARYNous Research, known for its open-weight Hermes LLMs, is in talks to raise new funding at a $1.5B valuation, signaling strong investor appetite for agent-focused AI labs outside the major incumbents.

コミュニティCommunityLocal Models·Zenn LLM

AI導入で逆に非効率化した人へ:時代を超えて効く自動化5原則と実践コードThis article addresses developers who found AI adoption made them less…

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約AIツールを導入したにもかかわらず作業効率が下がった開発者向けに、ツールの種類を問わず有効な自動化の5原則と具体的な実装例を解説した記事。正しい原則を理解することで、AI活用の効果を最大化できる。

AI SUMMARYThis article addresses developers who found AI adoption made them less productive, offering five timeless automation principles and practical code examples that apply regardless of tooling. Understanding these fundamentals helps maximize the real-world value of AI integration.

報道NewsNews/Policy·Ars Technica

防御側もプロンプトインジェクションを活用し始めたNow, defenders are embracing the prompt injection, too

重要度 MediumMedium priority技術記事 · Industry & Policytechnical post · Industry & Policy

AI要約攻撃者の手法だったプロンプトインジェクションを、セキュリティ防御側が逆用する動きが広まっている。AIエージェントへの悪意ある指示を無力化するための新たな対策として注目される。

AI SUMMARYSecurity defenders are now repurposing prompt injection techniques to protect AI systems, turning an attacker's tool into a defensive mechanism that neutralizes malicious instructions targeting AI agents.

Mon, Jul 1311 entries
コミュニティCommunityLocal Models·Zenn LLM

9つの意図に絞ることで38MBのモデルで十分だった — 30Mパラメータモデルをゼロから学習した実測報告By limiting intent classification to just 9 categories, the author trained a…

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約意図分類を9種類に限定することで、30Mパラメータ・38MBという極小モデルをゼロから学習し実用精度を達成した実験報告。タスクを絞ることでローカルLLMの軽量化が現実的に可能であることを示している。

AI SUMMARYBy limiting intent classification to just 9 categories, the author trained a 30M-parameter (38 MB) model from scratch and achieved practical accuracy. This demonstrates that scoping tasks aggressively makes ultra-lightweight local LLMs viable.

コミュニティCommunityLocal Models·Zenn LLM

Apple IntelligenceのローカルLLMをPythonから呼び出す方法This article explains how to invoke Apple Intelligence's on-device LLM directly…

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約macOS上で動作するApple IntelligenceのローカルLLMをPython経由で直接呼び出す手法を解説した記事。オンデバイスAIをサードパーティアプリから活用できる点が注目される。

AI SUMMARYThis article explains how to invoke Apple Intelligence's on-device LLM directly from Python, enabling developers to leverage Apple's private local model in their own applications without relying on cloud APIs.

コミュニティCommunityLocal Models·Qiita LLM

Ollamaのモデル別同時実行制限だけでは防げなかった過負荷の話Even with per-model concurrency limits configured in Ollama, GPU resource…

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約Ollamaでモデルごとに同時実行数を制限しても、複数モデルの並列利用によりGPUリソースが枯渇し過負荷が発生するケースがあることを解説した記事。適切な運用には全体的なリソース管理が必要だと示している。

AI SUMMARYEven with per-model concurrency limits configured in Ollama, GPU resource exhaustion can still occur when multiple models run simultaneously, highlighting the need for holistic resource management beyond per-model settings.

Ollamaのモデル別同時実行制限だけでは防げなかった過負荷の話og
論文PaperPapers/Benchmarks·arXiv cs.SE

プログラマーはLLMが生成したアサーションの評価が苦手で過信しがちProgrammers Are Poor and Overconfident Judges of LLM-Generated Assertions

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約LLMが生成したテストアサーションをプログラマーが評価する際、正確性を過大評価しミスを見落としやすいことが実験で示された。自動テスト生成の品質保証に人間のレビューだけでは不十分である可能性を示唆する。

AI SUMMARYA study found that programmers systematically overestimate the correctness of LLM-generated test assertions and miss significant errors, raising concerns about relying on human review as a quality gate for AI-generated tests.

Programmers Are Poor and Overconfident Judges of LLM-Generated Assertionsog
論文PaperPapers/Benchmarks·arXiv cs.SE

LLM生成コードにおける「パッチワーク問題」The Patchwork Problem in LLM-Generated Code

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約LLMが生成するコードは断片的な知識を継ぎ接ぎした構造になりやすく、一貫性や保守性に欠けるという問題を論文が指摘している。この知見はAIコード生成ツールの評価・改善指針として重要な意味を持つ。

AI SUMMARYResearchers identify the "patchwork problem" in LLM-generated code, where outputs are stitched together from disparate training patterns, leading to inconsistency and poor maintainability. This has significant implications for how AI coding tools should be evaluated and improved.

論文PaperPapers/Benchmarks·arXiv cs.SE

SCATE: コスト効率の高いテスト生成のためにコーディングエージェントを監督する学習SCATE: Learning to Supervise Coding Agents for Cost-Effective Test Generation

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約SCATEはLLMベースのコーディングエージェントを監督者モデルで制御し、テスト生成コストを抑えながら品質を維持する手法を提案する。監督者がエージェントの行動を動的に評価することで、効率的なソフトウェアテスト自動化を実現する。

AI SUMMARYSCATE proposes training a supervisor model to guide LLM-based coding agents during automated test generation, reducing computational cost while maintaining coverage quality. This approach makes agent-driven software testing more practical for real-world use.

論文PaperPapers/Benchmarks·arXiv cs.SE

汎用から個別へ:ペルソナを考慮したコードレビュー説明の探求From Generic to Personalized: Exploring Persona-Aware Code Review Explanations

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本研究は、開発者の経験レベルや役割に応じてコードレビューのフィードバック説明を個別化するペルソナ対応アプローチを提案し、画一的な説明の限界を克服しようとしている。

AI SUMMARYThis paper proposes a persona-aware approach to generating code review explanations tailored to a developer's experience and role, showing that personalized feedback improves comprehension over generic explanations.

論文PaperPapers/Benchmarks·arXiv cs.SE

Bugs4Qを用いたQiskitプログラム修復におけるLLMのベンチマーク評価Benchmarking Large Language Models on Repairing Qiskit Programs using Bugs4Q

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約量子プログラムのバグ修復タスクにLLMを適用し、Bugs4Qベンチマークで性能を評価した研究。量子ソフトウェア開発における自動修復の可能性と限界を明らかにしている。

AI SUMMARYThis study evaluates large language models on automatically repairing buggy Qiskit quantum programs using the Bugs4Q benchmark, revealing both the promise and current limitations of LLM-based repair for quantum software.

論文PaperPapers/Benchmarks·arXiv cs.SE

スキルマーケットの内側:ソフトウェアエンジニアリング活動から再利用可能なエージェントスキルへInside the Skill Market: From Software Engineering Activities to Reusable Agent Skills

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約ソフトウェアエンジニアリングのタスクをエージェントが再利用可能なスキルとして体系化する「スキルマーケット」フレームワークを提案し、エージェントの汎化性能と効率を向上させる研究。

AI SUMMARYThis paper proposes a "Skill Market" framework that distills software engineering activities into reusable agent skills, enabling more generalizable and efficient AI agents for SE tasks.

論文PaperPapers/Benchmarks·arXiv cs.SE

人間のテスト工程に着想を得たワークフローによるユニットテスト自動生成のためのマルチエージェントLLM協調Multi-Agent LLM Collaboration for Unit Test Generation via Human-Testing-Inspired Workflows

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約複数のLLMエージェントが人間のテスト設計プロセスを模倣して協調することで、ユニットテストの品質とカバレッジを向上させる手法を提案した研究。従来の単一モデルによる生成より効果的なテスト作成が可能になる。

AI SUMMARYThis paper proposes a multi-agent LLM framework that mimics human software testing workflows to collaboratively generate higher-quality unit tests with improved coverage, outperforming single-model approaches.

コミュニティCommunityLocal Models·Zenn LLM

OpenFugu×ローカルLLM群でマルチAI駆動を検証③ 小型の群れは上位モデルを超えられるかThis third installment investigates whether a coordinated swarm of small local…

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約OpenFuguを用いて複数の小型ローカルLLMを協調させるマルチエージェント構成が、単体の大規模モデルの性能を上回れるかを実験的に検証した第3弾レポート。小型モデルの集合知が上位モデルに対抗できる可能性と限界を示している。

AI SUMMARYThis third installment investigates whether a coordinated swarm of small local LLMs running under OpenFugu can collectively outperform a single large model, revealing both the promise and practical limits of multi-agent ensemble approaches.

Sun, Jul 121 entries
コミュニティCommunityClaude Code·Qiita Claude

Claudeで日本株を分析するAIアナリストを無料APIとLLMで構築し書籍化A developer compiled a book on building a Japanese stock AI analyst using free…

重要度 MediumMedium priority技術記事 · Claude / Claude Codetechnical post · Claude / Claude Code

AI要約無料の財務データAPIとClaudeを組み合わせて日本株の決算書を自動解析するAIアナリストの構築手法を解説し、その内容を書籍としてまとめた。個人投資家がLLMを活用した財務分析を実践できる具体的な手順を提供している。

AI SUMMARYA developer compiled a book on building a Japanese stock AI analyst using free financial data APIs and Claude, enabling automated earnings report analysis. The project offers individual investors a practical guide to LLM-powered fundamental analysis.

Sat, Jul 114 entries
コミュニティCommunityLocal Models·Qiita LLM

LM StudioでローカルLLM環境を構築してみたA hands-on guide to setting up a local LLM environment using LM Studio,…

重要度 InfoInformational深掘り候補 · 技術記事 · Local LLM / Open ModelsDeep-dive candidate · technical post · Local LLM / Open Models

AI要約LM Studioを使ってローカル環境でLLMを動かす手順を解説した記事。クラウドに依存せずプライバシーを保ちながらAIを活用できる点が注目される。

AI SUMMARYA hands-on guide to setting up a local LLM environment using LM Studio, enabling private, offline AI inference without relying on cloud services.

LM StudioでローカルLLM環境を構築してみたog
論文PaperPapers/Benchmarks·arXiv cs.AI

人間とLLMの混成集団に向けた対立的社会認識論Adversarial Social Epistemology for Assemblies of Humans and Large Language Models

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約人間とLLMが混在する集合的意思決定の場で、悪意ある操作や認識論的攻撃がどう機能するかを分析した研究。AIを含む社会的知識形成の堅牢性設計に重要な示唆を与える。

AI SUMMARYThis paper analyzes how adversarial actors can exploit mixed human-LLM assemblies to distort collective knowledge and decision-making, offering a framework for building more robust epistemic systems that include AI participants.

論文PaperPapers/Benchmarks·arXiv cs.AI

VectorizationLLM: ベクトル化に基づくスマートAIアシスタントVectorizationLLM: Smart Vectorization Based AI Assistant

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本論文はベクトル化技術を活用したLLMベースのAIアシスタント手法を提案し、効率的な情報検索と応答生成の改善を目指している。

AI SUMMARYThis paper proposes VectorizationLLM, an AI assistant leveraging smart vectorization to enhance retrieval and response quality in large language model systems.

論文PaperPapers/Benchmarks·arXiv cs.AI

ストレートスルー引受におけるエージェント型AIと検索拡張モデルAgentic AI and Retrieval-Augmented Models in Straight-Through Underwriting

重要度 MediumMedium priority論文/研究 · Papers / Benchmarkspaper/research · Papers / Benchmarks

AI要約本論文は、保険引受の完全自動化(ストレートスルー処理)にエージェント型AIとRAGを組み合わせる手法を提案し、意思決定の精度と説明可能性の向上を示した。

AI SUMMARYThis paper proposes combining agentic AI with retrieval-augmented generation for fully automated insurance underwriting, demonstrating improved decision accuracy and explainability in straight-through processing.