HomeCategoriesPapers / Benchmarks

Category detail

Research15 total

Research カテゴリの更新を、新着順・30日トレンド・関連記事として確認できます。Browse Research updates by recency, 30-day trend, and related topics.

Total15現在のlive indexCurrent live index
Last 7d0直近7日の掲載数Entries in the latest 7 days
Vs prev 7d0%その前の7日間と比較Compared with the prior 7 days
Avg/day0直近7日 ÷ 7Latest 7 days divided by 7

Research navigation

Research は選定した論文・レポートを含みますResearch includes selected papers and reports

このカテゴリでは選定した論文・レポートを扱い、専用の arXiv レーンでは論文だけをまとめて閲覧できます。

Browse selected papers and reports here, or use the dedicated arXiv lane for a papers-only view.

arXiv Papers
TrendLast 30 days
This week 0— 0%Last week 0Daily avg 0
Jul 19Jul 26Aug 2Aug 9Aug 16Aug 17 ↑today
Papers/Benchmarks trend counts
DateCount
2026-07-190
2026-07-200
2026-07-210
2026-07-220
2026-07-230
2026-07-240
2026-07-250
2026-07-260
2026-07-270
2026-07-280
2026-07-290
2026-07-300
2026-07-310
2026-08-010
2026-08-020
2026-08-030
2026-08-040
2026-08-050
2026-08-060
2026-08-070
2026-08-080
2026-08-090
2026-08-100
2026-08-110
2026-08-120
2026-08-130
2026-08-140
2026-08-150
2026-08-160
2026-08-170
現在の Research 一覧(arXiv 専用レーンを除く)を元に、直近 7 日間は vivid、それ以前は薄色で表示Based on the current Research listing (excluding the dedicated arXiv lane); the last 7 days are vivid and earlier days are muted.

All articles15 total

新着順Newest first
Wed, Jul 11 entries
新規収集INDEXED公式OfficialPapers/Benchmarks·Hugging Face Blog

ScarfBench: エンタープライズJavaフレームワーク移行のためのAIエージェントベンチマークScarfBench: Benchmarking AI Agents for Enterprise Java Framework Migration

重要度 MediumMedium priority技術記事 · Papers / Benchmarkstechnical post · Papers / Benchmarks

AI要約IBM ResearchがScarfBenchを公開。AIエージェントがエンタープライズJavaのフレームワーク移行タスクをどれだけ自律的にこなせるかを評価するベンチマークで、実務での活用可能性を測る基準を提供する。

AI SUMMARYIBM Research introduced ScarfBench, a benchmark designed to evaluate AI agents on real-world enterprise Java framework migration tasks, providing a standardized way to measure how effectively models handle complex legacy modernization work.

Fri, Jun 191 entries
新規収集INDEXED公式OfficialPapers/Benchmarks·Hugging Face Blog

MosaicLeaks: リサーチエージェントは機密を守れるか?MosaicLeaks: Can your research agent keep a secret?

重要度 MediumMedium priority技術記事 · Papers / Benchmarkstechnical post · Papers / Benchmarks

AI要約AIリサーチエージェントが機密情報をモザイク的に漏洩させるリスク「MosaicLeaks」を検証した研究で、ローカルLLMを用いたエージェント設計における情報セキュリティの重要性を示している。

AI SUMMARYThis research introduces MosaicLeaks, a study examining how AI research agents can inadvertently leak sensitive information through aggregated outputs, highlighting critical security considerations for agent design using open models.

Thu, Jun 182 entries
新規収集INDEXED公式OfficialPapers/Benchmarks·Hugging Face Blog

LoRAを超えて:最も人気のあるファインチューニング手法に勝てるか?Beyond LoRA: Can you beat the most popular fine-tuning technique?

重要度 MediumMedium priority技術記事 · Papers / Benchmarkstechnical post · Papers / Benchmarks

AI要約HuggingFaceがLoRAと競合する各種PEFTアルゴリズムを比較検証し、タスクや制約に応じた最適な手法の選び方を解説している。LoRA一択ではなく用途次第でより優れた選択肢が存在することを示す点で重要。

AI SUMMARYHugging Face explores PEFT methods that rival or surpass LoRA, benchmarking alternatives across tasks to help practitioners choose the best fine-tuning approach for their specific constraints.

新規収集INDEXED公式OfficialPapers/Benchmarks·Hugging Face Blog

エージェント能力は十分か?独自ツールでオープンモデルをベンチマークするIs it agentic enough? Benchmarking open models on your own tooling

重要度 MediumMedium priority技術記事 · Papers / Benchmarkstechnical post · Papers / Benchmarks

AI要約オープンLLMのエージェント性能を自社ツール環境で評価するベンチマーク手法を解説し、モデル選定の実践的指針を提供する。

AI SUMMARYThis article presents a practical framework for benchmarking open LLMs on agentic tasks using custom tooling, helping developers choose the right model for real-world agent workflows.

Tue, Jun 21 entries
新規収集INDEXED公式OfficialPapers/Benchmarks·DORA Insights (Google)

「tokenmaxxing」時代におけるバランスの取り方——DORAが警鐘を鳴らすFinding balance in the era of tokenmaxxing

重要度 InfoInformational技術記事 · Papers / Benchmarkstechnical post · Papers / Benchmarks

AI要約AIトークン消費量を社内リーダーボードで競わせる「tokenmaxxing」がAI導入促進策として広がるなか、DORAの調査はこの数値偏重がグッドハートの法則に陥り生産性指標を歪めると警告し、質と量のバランスの重要性を訴えている。

AI SUMMARYDORA warns that "tokenmaxxing"—rewarding raw AI token consumption via internal leaderboards to spur adoption—can become a vanity metric that distorts real productivity, urging teams to balance quantity with quality.

Tue, Mar 101 entries
新規収集INDEXED公式OfficialPapers/Benchmarks·DORA Insights (Google)

DORA調査:AI導入から効果的なSDLC活用へ、緊張関係をどう調整するかBalancing AI tensions: Moving from AI adoption to effective SDLC use

重要度 InfoInformational技術記事 · Papers / Benchmarkstechnical post · Papers / Benchmarks

AI要約DORAは、AIがコード生成を加速し初期摩擦を減らす一方で、検証負担やスキル低下、統合の難しさという隠れたコストを生むと指摘。単なる導入から効果的活用へ移るには、生産性と品質、信頼と検証のバランスが鍵だと論じる。

AI SUMMARYDORA argues that while AI speeds initial code generation, it adds hidden taxes like verification overhead, skill degradation, and integration challenges, so leaders must balance productivity with quality to shift from adoption to effective SDLC use.

Tue, Feb 171 entries
新規収集INDEXED公式OfficialPapers/Benchmarks·DORA Insights (Google)

DORA調査:AI依存をどう管理するかManaging AI dependency

重要度 InfoInformational技術記事 · Papers / Benchmarkstechnical post · Papers / Benchmarks

AI要約DORAがUCバークレー校の学生と行った調査をもとに、AIツールへの過度な依存リスクを指摘し、ガードレール設定やスキル維持、批判的検証を通じて生産性と長期的な専門性を両立する方法を解説する。

AI SUMMARYDORA's research with UC Berkeley students examines the risk of over-relying on AI coding tools and shows how developers set guardrails, retain skills, and apply critical verification to balance productivity with lasting expertise.

Wed, Jan 71 entries
新規収集INDEXED公式OfficialPapers/Benchmarks·DORA Insights (Google)

DORA 2025年の振り返り:AI時代の開発生産性研究の進化DORA 2025: Year in review

重要度 InfoInformational技術記事 · Papers / Benchmarkstechnical post · Papers / Benchmarks

AI要約GoogleのDORAチームが2025年の活動を総括。AI支援開発に関する大規模調査「State of AI-assisted Software Development Report」の公開や、AI導入を成熟させるためのDORA AI Capabilities Modelの提示など、研究の重心がAI時代の開発生産性へとシフトしたことを報告している。

AI SUMMARYA look back at the highlights and community contributions from 2025.

Fri, Jan 21 entries
新規収集INDEXED公式OfficialPapers/Benchmarks·DORA Insights (Google)

DORAソフトウェアデリバリ指標の歴史と進化A history of DORA’s software delivery metrics

重要度 InfoInformational技術記事 · Papers / Benchmarkstechnical post · Papers / Benchmarks

AI要約DORAが10年以上にわたり研究してきたソフトウェアデリバリ指標の変遷を振り返る記事。デプロイ頻度やリードタイムなど4つの主要指標の成立経緯と、信頼性指標の追加など近年の改訂を解説している。

AI SUMMARYDORA’s software delivery performance metrics have evolved over time to reflect the changing technology landscape. Learn about the transition from the four keys to the current five-metric model.

Fri, Dec 121 entries
新規収集INDEXED公式OfficialPapers/Benchmarks·DORA Insights (Google)

AIを家庭教師として活用する開発組織の学習スタイルAI as a tutor

重要度 InfoInformational技術記事 · Papers / Benchmarkstechnical post · Papers / Benchmarks

AI要約DORAの調査によると、AIを単なるコード生成器ではなく家庭教師のように使い概念や設計の理解を深める開発者は、生産性や学習効果が高い傾向にある。組織はこの使い方を促す文化と仕組みを整えるべきだと指摘している。

AI SUMMARYDORA research finds developers who use AI as a tutor to deepen understanding of concepts and design outperform those treating it as a mere code generator, urging organizations to foster this learning practice.

Fri, Oct 171 entries
新規収集INDEXED公式OfficialPapers/Benchmarks·DORA Insights (Google)

AI時代における「ビルダー意図」の理解とDORAの新視点Understanding builder intent in the AI era

重要度 InfoInformational技術記事 · Papers / Benchmarkstechnical post · Papers / Benchmarks

AI要約AIがコード生成を担う時代に、DORAは人間の「ビルダーとしての意図」こそが成果を左右すると指摘。目的意識・判断・責任の質を捉える「ビルダーマインドセット」枠組みを提示し、4つの中核的意図を定義している。

AI SUMMARYAs AI decouples roles from tasks, DORA introduces a "Builder Mindset" framework defining four core intents, arguing that the quality of human intent now drives software delivery performance.

Tue, Sep 231 entries
新規収集INDEXED公式OfficialPapers/Benchmarks·DORA Insights (Google)

DORA調査: カスタマイズ可能なツールが開発者エンゲージメントを高めるHow customization supports developer engagement

重要度 InfoInformational技術記事 · Papers / Benchmarkstechnical post · Papers / Benchmarks

AI要約DORAの調査は、開発者がワークフローやAIツールを自分に合わせてカスタマイズできる環境がエンゲージメントを高め、生産性とウェルビーイングを向上させると示す。組織はツール選定や運用に柔軟性を持たせるべきだと提言する。

AI SUMMARYDORA research finds that letting developers customize their workflows and AI tools boosts engagement, productivity, and well-being, urging organizations to build flexibility into tool choice and operations.

Tue, Sep 161 entries
新規収集INDEXED公式OfficialPapers/Benchmarks·Hugging Face Blog

LeRobotDataset v3.0: lerobot に大規模データセットを導入(新しいタブで開きます)`LeRobotDataset:v3.0`: Bringing large-scale datasets to `lerobot`(opens in a new tab)

重要度 MediumMedium priority技術記事 · Papers / Benchmarkstechnical post · Papers / Benchmarks

AI要約LeRobotDataset v3.0 では大規模ロボティクスデータセットの効率的な管理・利用が可能になり、研究者が実機学習をスケールアップしやすくなった。

AI SUMMARYLeRobotDataset v3.0 introduces large-scale dataset support for the lerobot framework, making it significantly easier to manage and train on high-volume robot learning data.

Wed, Aug 61 entries
公式OfficialPapers/Benchmarks·DORA Insights (Google)

DORA、組織目標に合った測定フレームワークの選び方を解説Choosing measurement frameworks to fit your organizational goals

重要度 InfoInformational深掘り候補 · 技術記事 · Papers / BenchmarksDeep-dive candidate · technical post · Papers / Benchmarks

AI要約DORAメトリクス、SPACE、DevExなど複数の測定フレームワークを組織の目標に応じて使い分ける重要性を解説。単一指標に頼らず補完的に組み合わせる実践的な選択基準を示している。

AI SUMMARYDORA explains how to choose among measurement frameworks like DORA metrics, SPACE, and DevEx based on organizational goals, advocating complementary use over reliance on a single metric for effective software measurement.

Fri, Jul 41 entries
新規収集INDEXED公式OfficialPapers/Benchmarks·Hugging Face Blog

NeurIPS 2025 E2LMコンペティション発表:言語モデルの早期学習評価(新しいタブで開きます)Announcing NeurIPS 2025 E2LM Competition: Early Training Evaluation of Language Models(opens in a new tab)

重要度 MediumMedium priority技術記事 · Papers / Benchmarkstechnical post · Papers / Benchmarks

AI要約TII UAEがNeurIPS 2025向けにE2LMコンペを発表。学習初期段階のチェックポイントから最終性能を予測する手法を競い、LLM訓練コスト削減に貢献することを目指す。

AI SUMMARYTII UAE has announced the NeurIPS 2025 E2LM Competition, challenging participants to predict a language model's final performance from early training checkpoints, aiming to reduce the massive compute costs of full LLM training runs.