RETROSPECTIVE RECORD · PREPARED 16 SEPTEMBER 2026The field guide · 120 retrospective records ↗
Turntaking Review

The field guide / Agents & tools

Agents & tools / From the field guide · 13 March 2024 event · prepared 16 September 2026

DeepMind trained one agent to follow instructions across nine games

SIMA's own report says human experts scored only 60% on its hardest test, showing the evaluation's own difficulty.

deepmind.googleprimary record

SIMA: A generalist AI agent for 3D virtual environments

Document
13 March 2024
Event
13 March 2024
Retrieved
16 September 2026
No visual was published with this record, so its primary document stands in its place.

The conversation

On 13 March 2024, Google DeepMind published a blog post and an accompanying technical report introducing SIMA, the Scalable Instructable Multiworld Agent, described as trained to 'follow natural-language instructions to carry out tasks in a variety of video game settings.' DeepMind states it partnered with eight external game studios, including Hello Games and Tuxedo Labs, to train and test SIMA across nine commercial video games plus four research environments, including a custom Unity environment the report calls the Construction Lab.

What the documents show

SIMA takes only two inputs, on-screen images and a text instruction, and outputs keyboard-and-mouse actions, without access to a game's source code or a bespoke API, which the post frames as letting one agent interface with any game a human could play. DeepMind reports evaluating SIMA across 600 basic skills spanning navigation, object interaction and menu use, and separately on nearly 1,500 unique in-game tasks judged in part by human raters. The report states an agent trained across all nine games 'significantly outperformed' agents trained on only one game each, and that an agent trained on all but one held-out game performed 'nearly as well' on the unseen game as a specialist trained on it directly, DeepMind's own comparison rather than a third-party score.

The system boundary

The report is explicit that this is not about game scores: DeepMind states the project 'isn't about achieving high game scores,' framing SIMA instead as testing whether language can be grounded in real-time action across many environments. A control test removed language entirely, and the agent 'behaves in an appropriate but aimless manner,' for instance gathering resources rather than moving to an instructed location, showing the agent's behavior without an instruction is not the same as failure to act, but a documented loss of direction. DeepMind describes the current version as evaluated only on short tasks, completable in about ten seconds, not the multi-step planning it names as future work.

Where it fails

The technical report's own human-comparison experiment is a documented limitation of the evaluation itself: expert human players, using the same judges and criteria as the agents, achieved only a 60% success rate on a focused set of No Man's Sky tasks, which the report says demonstrates 'the difficulty of the tasks' and 'the stringency of' the grading criteria, including human failures caused by performing an unnecessary action before the instructed one. The report's own closing section states 'many skills and tasks remain out of reach' and lists more environments, more robust agents, and 'more comprehensive and carefully controlled evaluations' as future work, not present capability.

  • Is a reported SIMA result measured against other agents, a fixed skill list, or human performance on the same criteria?
  • What does the agent do when given no instruction at all, and does that reveal a goal the agent has learned versus a coincidence?
  • Does the cited task involve the short, single-step actions SIMA was evaluated on, or longer multi-step planning it was not?

DeepMind's own materials describe SIMA as early-stage research demonstrating cross-game generalization on short tasks, explicitly not a claim of human-level or general game-playing competence.

Sources & reading trail

SIMA: A generalist AI agent for 3D virtual environments ↗

DeepMind's own framing of SIMA's goal, its game/studio partnerships, the 600-skill and human-comparison evaluations.

Source published: 13 March 2024 · Retrieved: 16 September 2026

Scaling Instructable Agents Across Many Simulated Worlds (v1) ↗

The technical report's generalization results, the 60% human success-rate comparison, and its own future-work limitations.

Source published: 13 March 2024 · Retrieved: 16 September 2026

Documentation, rulings and incident records establish the entry; the boundary reading is Chatbot Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.