SIMA: A generalist AI agent for 3D virtual environments
- Document
- 13 March 2024
- Event
- 13 March 2024
- Retrieved
- 16 September 2026
The conversation
On 13 March 2024, Google DeepMind published a blog post and an accompanying technical report introducing SIMA, the Scalable Instructable Multiworld Agent, described as trained to 'follow natural-language instructions to carry out tasks in a variety of video game settings.' DeepMind states it partnered with eight external game studios, including Hello Games and Tuxedo Labs, to train and test SIMA across nine commercial video games plus four research environments, including a custom Unity environment the report calls the Construction Lab.
What the documents show
SIMA takes only two inputs, on-screen images and a text instruction, and outputs keyboard-and-mouse actions, without access to a game's source code or a bespoke API, which the post frames as letting one agent interface with any game a human could play. DeepMind reports evaluating SIMA across 600 basic skills spanning navigation, object interaction and menu use, and separately on nearly 1,500 unique in-game tasks judged in part by human raters. The report states an agent trained across all nine games 'significantly outperformed' agents trained on only one game each, and that an agent trained on all but one held-out game performed 'nearly as well' on the unseen game as a specialist trained on it directly, DeepMind's own comparison rather than a third-party score.
The system boundary
The report is explicit that this is not about game scores: DeepMind states the project 'isn't about achieving high game scores,' framing SIMA instead as testing whether language can be grounded in real-time action across many environments. A control test removed language entirely, and the agent 'behaves in an appropriate but aimless manner,' for instance gathering resources rather than moving to an instructed location, showing the agent's behavior without an instruction is not the same as failure to act, but a documented loss of direction. DeepMind describes the current version as evaluated only on short tasks, completable in about ten seconds, not the multi-step planning it names as future work.
Where it fails
The technical report's own human-comparison experiment is a documented limitation of the evaluation itself: expert human players, using the same judges and criteria as the agents, achieved only a 60% success rate on a focused set of No Man's Sky tasks, which the report says demonstrates 'the difficulty of the tasks' and 'the stringency of' the grading criteria, including human failures caused by performing an unnecessary action before the instructed one. The report's own closing section states 'many skills and tasks remain out of reach' and lists more environments, more robust agents, and 'more comprehensive and carefully controlled evaluations' as future work, not present capability.
- Is a reported SIMA result measured against other agents, a fixed skill list, or human performance on the same criteria?
- What does the agent do when given no instruction at all, and does that reveal a goal the agent has learned versus a coincidence?
- Does the cited task involve the short, single-step actions SIMA was evaluated on, or longer multi-step planning it was not?
DeepMind's own materials describe SIMA as early-stage research demonstrating cross-game generalization on short tasks, explicitly not a claim of human-level or general game-playing competence.
Sources & reading trail
DeepMind's own framing of SIMA's goal, its game/studio partnerships, the 600-skill and human-comparison evaluations.
Source published: 13 March 2024 · Retrieved: 16 September 2026
The technical report's generalization results, the 60% human success-rate comparison, and its own future-work limitations.
Source published: 13 March 2024 · Retrieved: 16 September 2026
Documentation, rulings and incident records establish the entry; the boundary reading is Chatbot Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.