SIMA (Scalable Instructable Multiworld Agent), which DeepMind published in 2024, is a generalist agent built for 3D virtual environments. Give it a screen and a simple natural-language instruction, and it plays a 3D game almost the way a human would.
Not One Game, But Many
SIMA isn’t a bot tuned for one specific game, and that’s what makes it interesting. Working with eight game studios, DeepMind trained it across nine titles, including No Man’s Sky (exploring alien planets), Satisfactory (building automated factories on an alien world), and Valheim (a Norse-mythology survival crafting game). It can follow human instructions and play games it has never seen before.
Playing Like a Human
SIMA’s most distinctive trait is that its inputs and outputs match a human’s exactly. It takes in the game screen (raw pixels) plus a spoken or typed instruction, and it outputs keyboard and mouse actions. Unlike a typical game bot that calls into the engine’s internal API, SIMA plays through the same interface a person uses: watching a screen and operating a keyboard and mouse. That’s why the researchers call it a versatile agent, one that adapts quickly to environments it hasn’t seen.
The training instructions described short tasks that could be completed in under 10 seconds: things like “open the map” or “climb the ladder.”
Training with Behavior Cloning
SIMA learns through behavior cloning: collect data on expert behavior, then train a policy to reproduce it directly.

The training data starts with recordings of expert players, and natural-language commands get attached to those trajectories after the fact. The agent receives both a gameplay video and a language instruction; the video is encoded with a pretrained image/video encoder, and the instruction with a text model.
The model itself is a newly trained cross-attended transformer capable of handling text, image, and video together, built on Transformer-XL for its long-context memory. It outputs a set of 8 actions (keyboard/mouse operations) to take next.
Training minimizes the cross-entropy loss between the model’s predicted actions and what the expert did. On top of that, SIMA borrows classifier-free guidance (CFG), a technique common in image generation, to give natural-language instructions more weight:
𝜋_CFG = 𝜋(image, language) + 𝜆 · (𝜋(image, language) − 𝜋(image))
The larger 𝜆 gets, the more the language instruction influences the predicted action.
Evaluation
When evaluated on following natural-language instructions, SIMA achieved overall success rates in the 50-65% range.

The gap between command types is fairly stark. SIMA handles movement-related commands like stop, move, and drive well, but success drops noticeably on commands that lean heavily on game systems, like cook, build, and collect.

Compared directly against expert human players on No Man’s Sky, SIMA hit a 34% success rate versus the experts’ 60%. It’s not at human level yet, but SIMA performed reasonably well even in zero-shot conditions, playing that game for the first time. That suggests SIMA learned skills that transfer between games, rather than only memorizing game-specific rules.

Finally, dropping natural-language instructions or CFG causes performance to fall off sharply, which shows SIMA follows the language instruction rather than reacting to screen state alone.
