Foundry Managed Compute 発表:Microsoft Foundry でオープンモデルを実行Announcing Foundry Managed Compute: Run open models in Microsoft Foundry
匿名の公開いいねです。記事の保存・お気に入りではなく、Featured、Top 3、重要度、掲載順位には影響しません。仕組みとプライバシーAnonymous public likes are reactions, not saved articles or bookmarks. They do not affect Featured, Top 3, importance, or listing order.How it works and privacy
Microsoft Foundry Managed Computeが発表され、オープンソースやカスタムAIモデルをフロンティアモデルと同じエンドポイント・SDK・請求体系でホストできるGPU PaaSが提供される。
Microsoft announced Foundry Managed Compute, a new GPU platform-as-a-service that lets developers host open-source and custom AI models behind the same endpoints, SDKs, and billing as frontier models.
要約と収集メタデータをもとに生成した AI 解説本文です。元記事全文の転載・翻訳ではありません。This AI explainer is generated from the summaries and collected metadata, not from a reproduction or translation of the full source article.
マイクロソフトは、開発者向けAIプラットフォーム「Microsoft Foundry」に、オープンソースやカスタムのAIモデルをGPU上で運用できる新サービス「Foundry Managed Compute」を発表した。フロンティアモデルと同じエンドポイント、SDK、課金体系のまま自前のモデルをホストできる点が特徴で、商用の大規模モデルと独自モデルを一貫した仕組みで扱えるようになる。
Foundry Managed Computeは、いわゆるGPUのPaaS(サービスとしてのプラットフォーム)に位置づけられる。利用者は推論基盤となるサーバーやドライバ、スケーリングの管理を直接行う必要がなく、モデルをデプロイすればマネージドなGPU環境上で推論エンドポイントが提供される。これにより、インフラ運用の負担を抑えつつ、オープンウェイトのモデルや社内で微調整したモデルを本番環境に乗せやすくなると見られる。
背景には、Meta(Llama系)やMistral、DeepSeek、Qwenといったオープンウェイトモデルの急速な普及がある。これらは重みが公開され、用途に応じた微調整や自社データへの適応がしやすい一方、実運用には相応のGPUリソースと運用知識が必要だった。Managed Computeはその実行環境を取り込むことで、フロンティアモデルとオープンモデルを同じワークフローで併用する選択肢を広げる狙いがあるとみられる。
Microsoft Foundryは、従来Azure AI Foundry(旧Azure AI Studio)として展開されてきたAI開発基盤を指すブランドで、モデルカタログやエージェント開発、評価などの機能を一括して提供している。今回の追加により、既存のフロンティアモデル向けAPIと同じ呼び出し方やコスト管理の中で、自前モデルの推論まで統合できることになる。
同種のマネージド推論や独自モデルのホスティングは、Amazon BedrockやGoogle Vertex AI、Hugging Faceなど各社も力を入れている領域であり、競争は激しい。エンドポイントやSDK、請求を共通化する今回のアプローチは、複数モデルを使い分ける企業にとって移行や運用のコストを下げうる。一方で、実際の性能や対応モデルの範囲、料金の詳細については、公式情報の確認が重要となる。
Microsoft has announced Foundry Managed Compute, a new GPU platform-as-a-service designed to let developers host open-source and custom AI models inside Microsoft Foundry using the same endpoints, software development kits, and billing arrangements that already serve the platform's frontier models. The move matters because it narrows the operational gap between proprietary hosted models, such as those from OpenAI, and the growing universe of open-weight and bespoke models that teams increasingly want to run for cost, customization, or data-control reasons.
The central idea is unification. Until now, organizations that wanted to deploy an open model such as a Llama, Mistral, or Phi variant, or a model they fine-tuned themselves, often had to manage separate infrastructure, distinct authentication paths, and different cost-tracking systems compared with calls to a commercial frontier model. Foundry Managed Compute aims to collapse that divide by placing both categories behind a consistent inference interface. According to the announcement, an application that already talks to a frontier model through the Foundry SDK should be able to point at a managed-compute deployment with minimal changes, because the request format, endpoint structure, and billing flow are intended to be the same.
Technically, the offering is positioned as a managed GPU service, meaning Microsoft provisions and operates the underlying accelerator hardware, scaling, and serving stack while the customer focuses on selecting a model and configuring a deployment. This is distinct from raw infrastructure-as-a-service, where a team would rent virtual machines with GPUs and assemble the serving software themselves. It also differs from fully serverless, pay-per-token model-as-a-service endpoints, in that managed compute typically implies dedicated capacity allocated to a deployment. That distinction tends to give more predictable latency and throughput for production workloads, at the cost of paying for reserved GPU time rather than only per request, though the precise pricing and capacity details would need to be confirmed against Microsoft's documentation.
The launch fits into a broader repositioning. Microsoft renamed Azure AI Foundry to Microsoft Foundry as it consolidated its enterprise AI tooling into a single platform spanning a model catalog, agent services, evaluation tooling, and deployment options. Foundry's model catalog already lists a large number of models from multiple providers, and Managed Compute appears intended to be the execution layer that lets customers actually run the open and custom entries in that catalog, not just the commercially hosted ones. By keeping the developer surface consistent, Microsoft is likely trying to make the choice between an open model and a frontier model a configuration decision rather than an architectural one.
There are practical reasons enterprises pursue this flexibility. Open-weight models can be fine-tuned on proprietary data and deployed within a controlled environment, which helps with domain accuracy and with governance requirements in regulated sectors. Running a smaller open model can also be cheaper at scale than repeatedly calling a large commercial model, particularly for narrow, high-volume tasks. At the same time, self-hosting introduces responsibilities around GPU availability, optimization, and reliability that a managed service is meant to absorb.
The competitive context is significant. Amazon Web Services offers comparable capabilities through Amazon Bedrock and SageMaker, and Google Cloud provides model hosting through Vertex AI, while specialized providers such as Hugging Face, Together AI, Fireworks, and Baseten focus on open-model inference. Microsoft's differentiation appears to rest on integration: tying managed GPU hosting to the same identity, security, monitoring, and billing infrastructure that enterprises already use across Azure and Foundry. For organizations standardizing on the Microsoft stack, that continuity can reduce engineering overhead and procurement friction.
Several details will determine how compelling the service proves in practice, including which GPU types are available, how quickly deployments scale, what regional coverage looks like, and how the economics compare with both serverless endpoints and self-managed alternatives. Organizations evaluating it will also want to understand how it interacts with Foundry's agent and evaluation features, and whether custom models require specific packaging formats. As with any newly announced platform service, capabilities and limits may evolve, so teams planning production deployments should validate the current specifications and pricing directly through Microsoft's official Foundry documentation before committing.
本ページの本文と要約は AI による自動生成です。日本語版と英語版は言語ごとに独立して生成されるため、表現や詳しさが異なる場合があります。正確性は元記事 (devblogs.microsoft.com) をご確認ください。The body and summaries are AI-generated independently for each language, so wording and detail may differ. Verify accuracy at the original source (devblogs.microsoft.com).





