Generative world simulation

DAO (道) computes the shared world’s state and dynamics, while JING (镜) generates each observer’s first-person experience. Actions change the world; those changes shape what humans and agents experience and how they act next. This continuous loop connects perception, decision, action, and world evolution within a world that persists beyond any individual observer.

01 / INTERACTIVE EXPERIENCE MODEL

JING.

Generates what one observer sees, hears, and experiences next from a first-person point of view.

EXPLORE THE EXPERIENCE MODEL ↓
02 / GENERATIVE WORLD ENGINE

DAO.

Maintains shared world state and rules, shapes local observations, and supports autonomous agents in a persistent society.

EXPLORE THE WORLD ENGINE ↓

JING.

An Egocentric Interactive Experience Model

XGEN-JING generates first-person experiences conditioned on your actions and observation history. Define the scenes you explore, the objects you manipulate, and the characters you communicate with—each interaction shaping what you experience next.

What JING renders is one view
into a larger world.

Beyond one observer’s experience, the world continues.
DAO maintains and updates this shared world, and provides JING with only what the current observer can perceive. It also enables agents to make decisions and act independently within the evolving shared world.

DAO.

A Computable World Engine for Shared Reality

DAO keeps a shared world running. Actions change its state, observers experience it from different viewpoints, and their experiences become part of what happens next. Agents also pursue their own goals in the world simulated by DAO.

GLOBAL STATERULE VALIDATIONSTATE UPDATEOBSERVER FILTERJING PROMPT

Discussion & Limitations

Together, DAO and JING move beyond generating the next frame toward simulating the world behind it. We hope this paradigm can help pave the way towards OASIS: a persistent shared world that people can inhabit, rather than a sequence of scenes generated on demand. In this vision, encounters carry history, choices have lasting consequences, and social life continues beyond any one participant’s presence. The aim is not just a world that responds to you, but a world—and a society—that evolves with and without you.

Realizing this vision reliably remains an open challenge. Generative World Simulation is a research preview, and practical economical real-time deployment requires lower latency, better visual fidelity, and more reliable interaction. Fast motion and rapid viewpoint changes can reduce camera and action precision, while extended sessions may introduce drift in scene details, spatial relationships, and interaction context. Improving motion robustness, persistent memory, and responsive control remains central to our future work.