HomeIndustry & PolicyGoogle DeepMindの新AIモデル「Gemini Robotics 2」がロボットの全身制御に対応

Google DeepMindの新AIモデル「Gemini Robotics 2」がロボットの全身制御に対応Google DeepMind’s new AI model can control a robot’s entire body

AI2 点サマリSummary highlight
  • Google DeepMindはGemini Robotics 2を発表し、従来の上半身制御から足先から指先までの全身動作制御へと機能を拡張した。
  • ヒューマノイドロボットの実用化に向けた重要な進展となる。

Google DeepMind upgraded its Gemini Robotics model to version 2, expanding control from the upper body to full whole-body humanoid motion from feet to fingertips, marking a significant step toward capable humanoid robots.

要約と収集メタデータをもとに生成した AI 解説本文です。元記事全文の転載・翻訳ではありません。This AI explainer is generated from the summaries and collected metadata, not from a reproduction or translation of the full source article.

Google DeepMindは、ロボット制御向けAIモデルの最新版「Gemini Robotics 2」を発表した。従来モデルが人型ロボットの上半身操作に主眼を置いていたのに対し、新版は足先から指先までを含む「全身動作」の制御に対応するとしており、実用的なヒューマノイド(人型ロボット)の実現に向けた重要な一歩と位置づけられている。

Gemini Roboticsは、Googleが手がけるマルチモーダル基盤モデル「Gemini」をロボット分野に応用した系譜に連なる。こうしたモデルは一般に、カメラなどが捉えた視覚情報と自然言語による指示を受け取り、それを具体的な関節や手先の動作へと変換する役割を担う。今回のアップデートは、これまで腕や手といった上半身に限られていた制御範囲を、脚部を含む姿勢の保持や移動に近い動きへ広げる点が核心と見られる。

全身の統合制御が難しいとされてきたのは、二足歩行に伴うバランス維持や、地面との接触を考慮した重心移動が、上半身の把持動作に比べて格段に複雑になるためだ。足先から指先までを一つのモデルで扱えれば、物を運びながら移動する、しゃがんで作業するといった連続的な動作の実現に近づく可能性がある。

Google DeepMindはGemini Robotics 2を発表し、従来の上半身制御から足先から指先までの全身動作制御へと機能を拡張した。
📰 Industry & Policy · 本記事のポイント

ヒューマノイド開発をめぐっては、Figureやテスラ(Optimus)、Boston DynamicsのAtlas、さらにNVIDIAが提供する開発基盤など、ハードウェアとソフトウェアの双方で競争が活発化している。Google DeepMindはAIモデル側の強化によって存在感を高めようとしていると見られ、汎用的な動作を学習させたモデルを複数のロボットへ展開できるかが今後の焦点となりそうだ。

一方で、対応するロボットの機種や安全性の検証、実環境での信頼性など、実用化に向けて残る課題は少なくない。発表時点で公開された情報は限られており、具体的な提供形態や適用範囲については、今後の追加情報を待つ必要がある。

Google DeepMind has released Gemini Robotics 2, an updated version of its robotics-focused AI model that the company says can "control entire humanoid robots." The upgrade matters because whole-body coordination has long been one of the hardest problems in humanoid robotics, and it signals that large AI developers are pushing their models beyond narrow tabletop tasks toward machines that can move and act more like people.

The headline change is the expansion of what the model can command. According to Google DeepMind, the previous version of Gemini Robotics focused primarily on controlling a humanoid robot's upper body, such as its arms and hands for manipulation tasks. Gemini Robotics 2 now supports what the company describes as "whole-body motions," spanning from a robot's feet to its fingertips. In practice, that appears to mean the model can coordinate locomotion, balance, and posture alongside the fine motor control needed to grasp and manipulate objects, rather than treating the upper and lower halves of a robot as separate problems.

That distinction is more significant than it may sound. Controlling a humanoid's legs and maintaining balance while the arms reach, lift, or carry involves managing shifting weight, contact forces, and stability in real time. Systems that handle only the upper body often assume the robot is stationary or externally supported. By extending control to the full body, Gemini Robotics 2 is positioned to tackle tasks that require a robot to walk, crouch, reach across a space, or stabilize itself while working, though the exact range of demonstrated behaviors and their reliability outside controlled settings remains to be seen.

Gemini Robotics belongs to a category increasingly known as vision-language-action, or VLA, models. These systems build on the multimodal foundations of large language models like Gemini, but instead of producing only text, they translate visual input and natural-language instructions into physical actions a robot can execute. The approach is meant to give robots more general-purpose flexibility, allowing them to interpret a spoken or written command and adapt to objects and environments they were not explicitly programmed for, rather than relying solely on hand-coded routines for each task.

The release fits into a broader and highly competitive push across the industry toward capable humanoid robots. Google DeepMind has previously worked with humanoid hardware makers, and companies such as Figure, Tesla with its Optimus program, Boston Dynamics, Apptronik, and others are all developing humanoid platforms or the software to run them. Chipmaker Nvidia has also invested heavily in robotics through simulation and foundation-model tools aimed at training robots more efficiently. The common thread is a bet that general-purpose AI models, trained on large and diverse datasets, can make robots easier to instruct and more adaptable than earlier generations built around task-specific programming.

For context, whole-body control has traditionally been the domain of specialized control engineering and reinforcement learning, often tuned for a single robot design. Folding that capability into a broader AI model that also handles perception and language is the more ambitious goal, because it could allow the same underlying system to be adapted across different robot bodies and tasks. Whether Gemini Robotics 2 generalizes cleanly across hardware, and how much task-specific tuning it still requires, are open questions that will likely become clearer as researchers and partners test it in real deployments.

As with earlier robotics announcements, it is worth treating demonstration results with some caution. Robotics systems that perform impressively in curated videos or lab conditions can behave differently in unstructured, real-world environments, where lighting, clutter, and unexpected obstacles complicate both perception and movement. Google DeepMind has not framed the model as a finished consumer product, and the practical timeline for such technology reaching factories, warehouses, or homes remains uncertain.

Even so, the move from upper-body manipulation to full-body control represents a meaningful step in how AI models are being applied to physical machines. It reflects a wider industry direction in which foundation models increasingly serve as the "brain" for robots, and it adds to the growing evidence that major AI labs see embodied systems as a key frontier alongside their work on text, image, and video generation.

  • 出典SourceThe Verge報道News
  • 直近30件の平均重要度Avg importance, last 301=Info · 2=Medium · 3=High
  • 配信形式FormatブログBlog
  • 重要度Importance重要度 MediumMedium priority(Industry & Policy 427件中、同等以上 318件)(318 of 427 Industry & Policy entries are equal or higher)
  • 情報の寿命Half-life⏱️ 短命 (ニュース)Short-lived (news)
  • 原文言語Source languageEN
  • 収集日時Collected2026/07/31 06:56

本ページの本文と要約は AI による自動生成です。日本語版と英語版は言語ごとに独立して生成されるため、表現や詳しさが異なる場合があります。正確性は元記事 (theverge.com) をご確認ください。The body and summaries are AI-generated independently for each language, so wording and detail may differ. Verify accuracy at the original source (theverge.com).

📰Industry & Policy の他の記事More from Industry & Policyもっと見る →View more →