Xiaoyu Zhan*,
Xinyu Wang*,
Xiaohong Zhang*,
Huanjie Zhu*,
Tengjiao Sun,
Pengcheng Fang,
Jiaxing Yu,
Yanwen Guo📧,
and Dongjie Fu📧
Modern game development relies heavily on conventional graphics pipelines. High-quality visual content requires modeling, material authoring, animation, lighting, effects, and runtime optimization, making asset production expensive and extending the development cycle of game prototypes. Recently, video foundation models are beginning to change film and video production, but games differ from linear media, they require not only continuous and realistic imagery, but also stable and reproducible gameplay rules, object states, and interaction outcomes. We present Magpie, a real-time generative world-rendering system for interactive games. Magpie separates gameplay execution from visual generation. Designers define scenes and rules in a game engine. At runtime, the Game Engine resolves player actions and maintains world state, while an independent Render Server generates visual output from white-box frames produced by the engine. The Render Server receives a text prompt and a first-frame image specifying the visual style only during initialization. During subsequent interaction, white-box frames serve as the continuing denoising condition, and camera poses retrieve historical frames relevant to the current viewpoint. Player actions, state variables, object properties, and event signals remain in the Game Engine and will not be passed directly to the Render Server. The generative model is therefore responsible for visual presentation, while the Game Engine continues to execute gameplay rules and control state. To train Magpie, we manually collect approximately 300 hours of interactive video in Unreal Engine scenes. The data covers basic locomotion, viewpoint changes, driving, sitting, collision interactions, and idle states, with synchronized high-fidelity renderings, white-box renderings, camera poses, and structured interaction records. Magpie provides a system-level implementation path for applying generative models to real-time game rendering. It preserves gameplay designability and reproducibility, and reduces the dependence of early game prototypes on complete visual assets.