HomeTags#gpu

Tag timeline

#gpu21 total

同じキーワードで束ねられた更新を確認できます。カテゴリをまたいだ関連ニュースや実装トピックの追跡に使えます。

Total21#gpu の全掲載記事All listed entries tagged #gpu
Showing21このページの表示件数Entries on this page
Page1/1静的ページ位置Static page position
Updated公開index snapshotPublished index snapshot

Entriespage 1/1 · 21 total

Wed, Aug 121 entries
新規収集INDEXED公式OfficialNews/Policy·NVIDIA Blog

AIコンピュートのスケーリングに新しい電力アーキテクチャが必要な理由Why Scaling AI Compute Performance Requires a New Power Architecture

重要度 MediumMedium priority技術記事 · Industry & Policytechnical post · Industry & Policy

AI要約AIファクトリーの高密度化に伴い、従来の電力配電方式では非効率が生じているため、NVIDIAは800VDCアーキテクチャへの移行を推進している。これによりラック単位での電力損失を削減し、大規模GPUクラスタの性能向上を支える。

AI SUMMARYNVIDIA outlines why traditional AC power distribution struggles to keep pace with modern AI compute density, and makes the case for an 800V DC power architecture that reduces conversion losses and enables more efficient scaling of GPU-dense AI factories.

Sat, Aug 81 entries
コミュニティCommunityLocal Models·Qiita LLM

【止まっちゃう事件の記録 #2】GPUメモリ衝突を解消したのに、ホストが「無痕跡」で凍りつく問題 〜熱暴走という仮説にたどり着くまで〜After fixing GPU memory profiling conflicts between co-resident vLLM instances,…

重要度 InfoInformational深掘り候補 · 技術記事 · Local LLM / Open ModelsDeep-dive candidate · technical post · Local LLM / Open Models

AI要約vLLM複数常駐によるGPUメモリ衝突ハングを修正した後も原因不明のホスト凍結が続き、ログ・クラッシュダンプ一切なしの症状から熱暴走という新仮説に至るまでの調査過程を記録した記事。

AI SUMMARYAfter fixing GPU memory profiling conflicts between co-resident vLLM instances, the author found the host still froze silently with no logs or crash dumps, and documents the investigation that led to a thermal-runaway hypothesis.

【止まっちゃう事件の記録 #2】GPUメモリ衝突を潰したのに、まだホストが「無痕跡」で凍りつく話 〜熱暴走という仮説にたどり着くまで〜og
Fri, Aug 71 entries
🔥 HOT新規収集INDEXED公式OfficialNews/Policy·AWS News Blog

Amazon Bedrock AgentCoreに「ランタイムインスタンス」登場——本番AIエージェント向け永続コンピュートRuntime instances: persistent compute for production AI agents on Amazon Bedrock AgentCore

重要度 HighHigh priority技術記事 · Industry & Policytechnical post · Industry & Policy

AI要約Amazon Bedrock AgentCoreがランタイムインスタンスを発表。最大14日間のセッション継続、GPU対応、マルチエージェント協調を備えた管理型EC2基盤で、本番AIエージェントの安定稼働を実現する。

AI SUMMARYAmazon Bedrock AgentCore now offers runtime instances—managed, persistent EC2 infrastructure supporting sessions up to 14 days, GPU workloads, and multi-agent collaboration for production AI agents.

Runtime instances: persistent compute for production AI agents on Amazon Bedrock AgentCoreog
Thu, Aug 62 entries
公式OfficialGemini/Gemma·Google Cloud Blog

MirendilがAI HypercomputerのTPUとGPUを採用、モデルの事前・事後学習に活用Mirendil taps AI Hypercomputer TPUs and GPUs for pre- and post-training applications

重要度 MediumMedium priority技術記事 · Gemini / Gemmatechnical post · Gemini / Gemma

AI要約フロンティアAIラボのMirendilがGoogle CloudのAI Hypercomputerを採用し、TPUとNVIDIA GPUを組み合わせてモデルの事前学習・事後学習に活用することが発表された。主要AIラボのGoogle Cloud採用が相次ぐ中、新興スタートアップへの広がりを示す動きとして注目される。

AI SUMMARYFrontier AI startup Mirendil has selected Google Cloud's AI Hypercomputer—combining Google TPUs and NVIDIA GPU infrastructure—to power its model pre-training and post-training workloads, underscoring Google Cloud's growing dominance as the platform of choice for cutting-edge AI labs.

公式OfficialGemini/Gemma·Google Cloud Blog

エージェントAIのスケーリング:UiPathがAI HypercomputerでGPUプラットフォームを構築した方法Scaling agentic AI: How UiPath built its high-performance GPU platform on AI Hypercomputer

重要度 MediumMedium priority技術記事 · Gemini / Gemmatechnical post · Gemini / Gemma

AI要約UiPathはGoogle CloudのAI Hypercomputer上に高性能GPUプラットフォームを構築し、数百のGPUを連携させることでエンタープライズ向けエージェントAIの大規模展開を実現した。

AI SUMMARYUiPath leveraged Google Cloud's AI Hypercomputer to build a high-performance GPU platform capable of orchestrating hundreds of GPUs, enabling reliable large-scale agentic AI for enterprise automation workloads.

Scaling agentic AI: How UiPath built its high-performance GPU platform on AI Hypercomputermedia
Tue, Aug 41 entries
新規収集INDEXED公式OfficialNews/Policy·Meta Engineering

GEMトレーニング:MetaがLLMスケールの広告基盤モデルの効率を2倍にした方法GEM Training: How Meta Doubled the Efficiency of Its LLM-Scale Ads Foundation Model

重要度 MediumMedium priority技術記事 · Industry & Policytechnical post · Industry & Policy

AI要約MetaはInstagram・Facebook向け広告推薦基盤モデルGEMのトレーニング効率を2倍(MFU 20〜25%)に引き上げつつ、計算量を4倍にスケールさせることに成功した。

AI SUMMARYMeta achieved a 2x improvement in end-to-end training efficiency for GEM, its ads recommendation foundation model, reaching 20–25% MFU while scaling training FLOPs 4x on thousands of latest-generation GPUs.

Mon, Jul 271 entries
コミュニティCommunityLocal Models·Zenn LLM

ローカルLLM向けハードウェアを「容量・帯域・MoE・TTFT」で選ぶThis article explains how to choose hardware for running local LLMs by…

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約ローカルLLMを快適に動かすには、メモリ容量・メモリ帯域・MoEアーキテクチャへの対応・初回トークン生成速度(TTFT)という4軸でハードウェアを評価することが重要だと解説した記事。GPU・CPUオフロード・専用アクセラレータの選び方に実践的な指針を提供している。

AI SUMMARYThis article explains how to choose hardware for running local LLMs by evaluating four key axes: memory capacity, memory bandwidth, MoE architecture support, and time-to-first-token (TTFT), offering practical guidance for selecting GPUs, CPU offload setups, and dedicated accelerators.

ローカルLLM向けハードウェアを「容量・帯域・MoE・TTFT」で選ぶog
Fri, Jul 241 entries
公式OfficialGemini/Gemma·Google Cloud Blog

アイドルアクセラレータを最小化:llm-dの協調タイムスライシングによるネイティブRLジョブインターリービングMinimize idle accelerators: Native RL job interleaving with co-operative time-slicing in llm-d

重要度 MediumMedium priority技術記事 · Gemini / Gemmatechnical post · Gemini / Gemma

AI要約llm-dに協調タイムスライシング機能が導入され、強化学習の推論・訓練ジョブをGPU上でインターリーブすることでアクセラレータの遊休時間を大幅に削減できるようになった。

AI SUMMARYllm-d now supports cooperative time-slicing for RL workloads, allowing inference and training jobs to interleave on the same GPUs and significantly reducing accelerator idle time.

Minimize idle accelerators: Native RL job interleaving with co-operative time-slicing in llm-dmedia
Wed, Jul 222 entries
公式OfficialNews/Policy·NVIDIA Blog

NVIDIAが初のGPUアクセラレーション医療物理シミュレーションフレームワークをオープンソース化NVIDIA Open Sources First GPU-Accelerated Medical Physics Simulation Framework

重要度 MediumMedium priority技術記事 · Industry & Policytechnical post · Industry & Policy

AI要約NVIDIAは医療物理シミュレーション向け初のGPUアクセラレーション対応オープンソースフレームワークを公開した。放射線治療などの計算を大幅に高速化し、研究・臨床応用の発展に貢献することが期待される。

AI SUMMARYNVIDIA has open-sourced its first GPU-accelerated medical physics simulation framework, enabling dramatically faster radiation therapy and related calculations that could accelerate both research and clinical workflows.

🔥 HOT公式OfficialNews/Policy·NVIDIA Blog

NVIDIA Vera Rubin、ワット当たり性能とトークンコストで世界のパートナーに優位性をもたらすNVIDIA Vera Rubin Driving Performance Per Watt, Lowest Token Cost for Partners Worldwide

重要度 HighHigh priority技術記事 · Industry & Policytechnical post · Industry & Policy

AI要約NVIDIAの新アーキテクチャ「Vera Rubin」は、ワット当たり性能を大幅に向上させ、AI推論のトークンコストを削減することで、データセンター運用コストの最適化に貢献する。

AI SUMMARYNVIDIA's Vera Rubin platform delivers significant gains in performance per watt and reduces per-token inference costs, giving cloud and enterprise partners a more efficient and economical path to large-scale AI deployment.

Tue, Jul 211 entries
公式OfficialLocal Models·Ollama Releases

v0.32.2-rc0: CUDA v12 の Linux 向けに CC 10.0 サポートを追加v0.32.2-rc0: cuda: add CC 10.0 for linux in CUDA v12 (#17025)

重要度 MediumMedium priority公式リリース · Local LLM / Open Modelsofficial release · Local LLM / Open Models

AI要約Ollama v0.32.2-rc0 では、Linux 環境の CUDA v12 に Compute Capability 10.0 対応が追加され、最新世代の NVIDIA GPU でのローカル LLM 実行が可能になります。

AI SUMMARYOllama v0.32.2-rc0 adds Compute Capability 10.0 support for Linux under CUDA v12, enabling local LLM inference on the latest generation of NVIDIA GPUs.

v0.32.2-rc0: cuda: add CC 10.0 for linux in CUDA v12 (#17025)media
Mon, Jul 131 entries
コミュニティCommunityLocal Models·Qiita LLM

Ollamaのモデル別同時実行制限だけでは防げなかった過負荷の話Even with per-model concurrency limits configured in Ollama, GPU resource…

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約Ollamaでモデルごとに同時実行数を制限しても、複数モデルの並列利用によりGPUリソースが枯渇し過負荷が発生するケースがあることを解説した記事。適切な運用には全体的なリソース管理が必要だと示している。

AI SUMMARYEven with per-model concurrency limits configured in Ollama, GPU resource exhaustion can still occur when multiple models run simultaneously, highlighting the need for holistic resource management beyond per-model settings.

Ollamaのモデル別同時実行制限だけでは防げなかった過負荷の話og
Mon, Jul 61 entries
新規収集INDEXED公式OfficialLocal Models·Hugging Face Blog

🤗 Kernels: 主要アップデート🤗 Kernels: Major Updates

重要度 MediumMedium priority技術記事 · Local LLM / Open Modelstechnical post · Local LLM / Open Models

AI要約Hugging FaceがKernelsプラットフォームを大幅に刷新した。最適化されたGPUカーネルの統合・共有の仕組みが改善され、ローカル環境でのLLM推論の高速化と開発体験の向上が期待できる。

AI SUMMARYHugging Face has revamped its Kernels platform, making it significantly easier to integrate and share optimized GPU kernels within inference pipelines, delivering better performance for local LLM deployments.

Tue, Jun 301 entries
公式OfficialNews/Policy·NVIDIA Blog

ClaudeがBlackwell Ultraに対応:AnthropicのモデルがAzureでNVIDIA GB300上での稼働を開始Claude Meets Blackwell Ultra: Anthropic’s Models Now Run on NVIDIA GB300 in Azure

重要度 InfoInformational深掘り候補 · 技術記事 · Industry & PolicyDeep-dive candidate · technical post · Industry & Policy

AI要約AnthropicのClaudeモデルが、Microsoft AzureにてNVIDIAの最新GPU「GB300 Blackwell Ultra」上での稼働に正式対応した。これにより、クラウド上でのAI推論の速度と処理効率が大幅に向上することが期待される。

AI SUMMARYAnthropic's Claude models are now available on NVIDIA's GB300 Blackwell Ultra GPUs within Microsoft Azure, enabling significantly faster and more efficient AI inference for large-scale cloud deployments.

Sun, Jun 281 entries
コミュニティCommunityLocal Models·Qiita LLM

WhichLLM入門 — 自分のGPUで最速のローカルLLMをCLIで選ぶハンズオンWhichLLM is an open-source CLI that recommends the best-performing local LLM…

重要度 InfoInformational深掘り候補 · 技術記事 · Local LLM / Open ModelsDeep-dive candidate · technical post · Local LLM / Open Models

AI要約WhichLLMは自分のハードウェアで最も高性能なローカルLLMをコマンド1発で推薦するOSSのCLIツールで、パラメータ数ではなくベンチマーク品質・VRAM適合・推定速度を統合したスコアで選定する。

AI SUMMARYWhichLLM is an open-source CLI that recommends the best-performing local LLM for your own hardware, scoring candidates by benchmark quality, VRAM fit, and estimated speed rather than raw parameter count.

WhichLLM入門 — 自分のGPUで最速のローカルLLMをCLIで選ぶハンズオンog
Wed, Jun 241 entries
公式OfficialNews/Policy·NVIDIA Blog

NVIDIAとAWSが連携、大規模AIの本番運用を加速NVIDIA and AWS Collaborate to Bring AI to Production at Scale

重要度 InfoInformational深掘り候補 · 技術記事 · Industry & PolicyDeep-dive candidate · technical post · Industry & Policy

AI要約NVIDIAとAWSが提携し、低レイテンシ推論や高速ベクトル検索、優れたGPUの価格性能を組み合わせ、運用の複雑さを増やさずにスケール可能なAI本番インフラの構築を支援すると発表した。

AI SUMMARYNVIDIA and AWS are collaborating to bring AI to production at scale, combining low-latency inference, fast vector search and strong GPU price-performance without added operational complexity.

NVIDIA and AWS Collaborate to Bring AI to Production at Scaleog
Mon, Jun 221 entries
公式OfficialNews/Policy·NVIDIA Blog

ジャグジーより熱い45°C冷却液──NVIDIAが大規模AIマシン向け液体冷却技術を刷新Hotter Than a Hot Tub: The 45°C Breakthrough to Cool AI’s Biggest Machines

重要度 InfoInformational深掘り候補 · 技術記事 · Industry & PolicyDeep-dive candidate · technical post · Industry & Policy

AI要約NVIDIAの最新AIサーバーはジャグジー(38〜40°C)を超える最大45°Cの冷却液で動作可能になり、電力を消費する冷却装置(チラー)への依存を減らせる。これによりAIデータセンターの冷却効率と省エネ性能が大幅に向上する。

AI SUMMARYNVIDIA's newest AI servers can run cooling liquid up to 45°C — hotter than a hot tub — letting AI data centers cut their reliance on power-hungry chillers and significantly improve energy efficiency.

Fri, Jun 191 entries
新規収集INDEXED公式OfficialNews/Policy·AWS News Blog

Amazon EC2 G7インスタンス正式提供開始 — NVIDIA RTX PRO 4500 Blackwell Server Edition GPU搭載Announcing Amazon EC2 G7 instances accelerated by NVIDIA RTX PRO 4500 Blackwell Server Edition GPUs

重要度 MediumMedium priority技術記事 · Industry & Policytechnical post · Industry & Policy

AI要約NVIDIA RTX PRO 4500 Blackwell Server Edition GPUを搭載したAmazon EC2 G7インスタンスが一般提供開始。AIインファレンス、グラフィックス、データ分析ワークロード向けの高性能GPU高速化を実現します。

AI SUMMARYAnnouncing the general availability of Amazon Elastic Compute Cloud (Amazon EC2) G7 instances, delivering high performance GPU acceleration for AI inference, graphics, and data analytics workloads.

Wed, Jun 101 entries
公式OfficialNews/Policy·Meta Newsroom

インフラ解説:コンピュートパワーInfrastructure Explained: Compute Power

重要度 InfoInformational深掘り候補 · 技術記事 · Industry & PolicyDeep-dive candidate · technical post · Industry & Policy

AI要約Metaが、AI開発を支えるコンピュートパワーの基礎を解説。GPU・CPU・独自のカスタムシリコン「MTIA」がどのように連携し、増大するAI処理需要やイノベーションを支えているかを紹介する解説記事。

AI SUMMARYMeta explains the fundamentals of compute power and how GPUs, CPUs, and its custom MTIA silicon work together to support AI innovation and meet rapidly growing computing demands.

Infrastructure Explained: Compute Powermedia
Thu, Jun 41 entries
新規収集INDEXED公式OfficialCopilot·Microsoft Foundry Blog

Foundry Managed Compute 発表:Microsoft Foundry でオープンモデルを実行Announcing Foundry Managed Compute: Run open models in Microsoft Foundry

重要度 MediumMedium priority技術記事 · GitHub Copilottechnical post · GitHub Copilot

AI要約Microsoft Foundry Managed Computeが発表され、オープンソースやカスタムAIモデルをフロンティアモデルと同じエンドポイント・SDK・請求体系でホストできるGPU PaaSが提供される。

AI SUMMARYMicrosoft announced Foundry Managed Compute, a new GPU platform-as-a-service that lets developers host open-source and custom AI models behind the same endpoints, SDKs, and billing as frontier models.

Tue, May 191 entries
公式OfficialGemini/Gemma·Google Developers Blog

Google Cloud × NVIDIA 開発者コミュニティが設立1周年、会員数10万人を突破(新しいタブで開きます)One Year of Innovation: Celebrating 100k Members in the Google Cloud x NVIDIA Developer Community(opens in a new tab)

重要度 InfoInformational深掘り候補 · 技術記事 · Gemini / GemmaDeep-dive candidate · technical post · Gemini / Gemma

AI要約Google Cloud と NVIDIA の共同開発者コミュニティが設立1年で会員10万人を達成し、AI・GPU インフラを活用する開発者向けの支援やリソース提供を強化する方針を示した。

AI SUMMARYThe joint Google Cloud and NVIDIA developer community reached 100,000 members in its first year, reaffirming its focus on giving builders advanced AI and GPU infrastructure, tools, and resources.

One Year of Innovation: Celebrating 100k Members in the Google Cloud x NVIDIA Developer Communityog