Dual Exam: Controlled Evaluation of Multimodal Agents and World Models
How do agents and world models work together to complete tasks? In such a system, an agent selects an action based on the current visual observation. A world model generates a view of the environment after that action, which the agent then uses to decide what to do next. This process requires the two components to work together: the agent must make accurate judgments and take effective actions, while the world model must correctly depict the results and preserve any state that the task requires to remain unchanged.
Consider packing a backpack before a trip. The agent needs to place a passport, a medicine case, and a charger into the same blue backpack, then carry it to the entryway and inspect its contents. During execution, the visual observations do not confirm whether some items have been picked up, yet the agent continues with its packing plan. By the later packing and carrying steps, it is using a different, smaller blue bag. This violates the requirement to use the same backpack throughout, and the task ultimately remains incomplete.

The execution record shows that none of the four discrepancies in item or backpack identity received a targeted response. The two actions counted as “corrective responses” were simply retries at navigating to the target location; they did not resolve these discrepancies. This case shows that continuing with a plan does not mean the task requirements are being met, and taking corrective action does not mean the problem has been resolved. Further testing is needed to determine whether the failure stems from the agent’s decisions, the environment states generated by the world model, or their interaction. To address this, Dual Exam introduces two controlled evaluation tracks that assess agents and world models separately.
Coupled Evaluation: How Agents and World Models Perform Together
In a coupled evaluation, we paired GPT-5.6 Sol with Seedance 2.5 and ran 150 cases across ten scenarios. Each case was run once, with at most 20 interaction steps and no episode-level retries; initialization failures and technical failures were excluded from these 150 cases. Under this model pairing and evaluation setup, strict task success was 0%, World State Consistency was 26.6%, and Agent Corrective Response was 9.2%.

The three metrics measure, respectively, whether the task is fully completed, whether task-relevant environment states remain consistent, and whether the agent takes a targeted action within two steps of a visible discrepancy that meets the scoring criteria. The Overall column reports the unweighted mean of the ten scenario scores, rounded to one decimal place.
Two Controlled Evaluation Tracks, Assessing Agents and World Models Separately
The coupled evaluation above examines the interaction between the paired models. The two separate evaluations assess the agent and the world model individually: the environment’s response rules are held fixed when evaluating the agent, while a validated action sequence is provided in advance when evaluating the world model.
| Component Evaluated | Controlled Condition | Evaluation Focus |
|---|---|---|
| Agent | Validated environment response mechanism | Interpreting observations, selecting actions, verifying progress, and completing the task |
| World Model | Validated action sequence | Correctly realizing action effects while maintaining consistency in object identity, relevant state, and rules |
In the agent evaluation, the agent selects its own actions, and the environment responds according to predefined rules. The rules remain fixed, while the visual observations update as actions are taken. In the world model evaluation, the evaluators provide the action sequence in advance and focus on whether the model correctly generates the resulting environment states. Both tracks score performance against task requirements and execution evidence; a model’s claim that it is “done” is not enough to establish success.

The separate evaluations cover four domains: physical, digital, social, and scientific. The agent evaluation comprises 1,088 evaluation items across 15 task categories and compares 10 model configurations. The world model evaluation comprises 400 cases across eight scenarios, with 50 cases per scenario, and reports results for 16 model configurations.
Of the world model configurations, 15 cover all scenarios. Face-generation restrictions prevented Seedance 2.5 from being evaluated on Daily Life and Human Interaction, so its average over the remaining six scenarios is reported separately. The two separate evaluations use different tasks and statistical definitions from the 150-case coupled evaluation described above, so their results cannot be pooled for comparison.
The separate agent and world model evaluations also do not require cases to be paired one to one, and their aggregate scores are not directly comparable. Task categories are used to summarize evaluation results and may overlap; a category score therefore should not be interpreted simply as a measure of one independent ability. Below, we use category-level results and individual cases to examine four issues worth considering.
Finding 1: Performing Individual Actions Does Not Mean Completing the Whole Task
The Agent Index is weighted by the number of evaluation items in each task category, rather than averaged equally across the four domains. GPT 6 Astra ranks first with 31.3, followed by GPT 5.6 Sol (20.7), Gemini 3.6 Flash (20.1), Qwen 3.8 Max (20.0), and Kimi K3 (19.8).

A closer look at the category-level results shows that GPT 6 Astra scores 63.3 on Planning and 49.7 on Navigation, but 7.1 on Manipulation and 11.2 on Action Sequencing. Even the highest scores across all model configurations in the latter two categories are only 9.5 and 11.2, respectively.
The digital domain also needs to be examined by task category. The highest scores for Web Interaction, Desktop Application Interaction, Mobile Application Interaction, and Game Environment Interaction are 29.6, 5.6, 42.9, and 47.1, respectively. Because task difficulty and scoring rules differ, these scores are useful for comparing models within the same category; score differences across categories cannot be interpreted directly as capability gaps.
Across all 15 task categories, GPT 6 Astra has the highest score in nine; DeepSeek V4.1 Flash, GPT 5.6 Sol, and Seed 2.1 Pro lead in the remaining three, two, and one categories, respectively. These differences show why model selection requires examining category-level results in relation to the intended task.
Beyond category scores, execution recordings help clarify the distinction between “performing an action” and “completing a task.” In the tower-building video below, the robotic arm picks up and moves parts on the table but ultimately fails to build the required tower. The evaluation record also marks the task as incomplete. This recording comes from a historical configuration and cannot be used to judge GPT 6 Astra’s performance on this task.
By contrast, the puzzle case below shows a task being completed. The agent drags the pieces into place step by step until they form a complete image, with the final visual state and the evaluation record corroborating each other.
The outcomes differ, but the principle for checking them is the same: assess whether the task requirements have been met. Picking up parts and dragging puzzle pieces are intermediate steps. Whether the target structure or image is complete determines the task outcome.
Finding 2: A Model’s Claim of Completion Does Not Mean the Task Is Complete
In actual evaluations, the agent’s completion report, the outcome in the environment, and the evaluator’s verdict may not agree. Determining whether a task succeeded requires checking all three against one another.
In a phone-brightness task, the model claims after four actions that it has set the brightness to maximum. The evaluation record, however, shows an actual brightness of 183/255 and a task score of 0. The model has reported completion, but the target state has not been reached.
This case suggests adding a check of the target state before ending a task to verify that the actual outcome meets the requirements. This is a potential improvement that remains to be validated; experiments are needed to determine whether it increases success rates.
Low scores, too, need to be interpreted alongside the evaluation records. In an earlier shopping-cart task, a score of 0 was recorded because the scoring system did not return a result. In this situation, the score alone cannot establish whether the model completed the task. Distinguishing “the goal was not achieved” from “scoring evidence is missing” helps avoid misattributing failure.
Finding 3: Making the Requested Change Does Not Mean Preserving Everything Else
For a world model, generating action effects is only part of the task. After an action, it must also preserve the identities, counts, and states of objects that the task requires to remain unchanged. The table below reports results across eight scenarios, grouped by image and video outputs. Image and video models differ in how they are interacted with and in the information they provide about the unfolding process. Score differences between the two groups can therefore be presented, but should not be used directly to judge which type of model is stronger.

Individual cases make these changes easier to see. Qwen-Image-3.0 restores the viewing angle as requested in a single-pendulum case, yet generates additional cups in a cup-stacking case. MiniMax H3 preserves the cup count in the stacking case, yet adds a red tray that was not originally present in a separate tray case. Action effects and state preservation need to be checked separately.
A separate single-pendulum video shows a change in object count: there is initially one orange pendulum bob, but another appears while it is swinging. The task requires preserving the original single-bob apparatus without adding objects, yet the generated visuals change the structure of the apparatus. The motion still looks smooth, but the experimental conditions have changed.
Similar problems appear in everyday scenarios. An action that should simply move an object may produce an additional copy of it. Some tasks require a state to remain unchanged throughout, yet the generated action may violate that requirement. The following two videos illustrate these situations.
The task requires moving the original pair of scissors back to the left. Instead, the scissors on the right remain in place while another pair appears on the left, increasing the number of pairs.
The task requires the refrigerator to remain closed throughout, but the server in the generated video pulls the door open, violating this state constraint.
In an interactive system, each new visual observation informs the agent’s next decision. World model evaluation therefore needs to check both whether the requested action effects are correctly realized and whether the state that must remain unchanged is preserved. Once object counts, identities, or attributes deviate from the requirements, the conditions on which subsequent plans depend may change as well.
Finding 4: Leading the Individual Rankings Does Not Mean Performing Best Together
Choosing a world model for an agent starts with the scenarios the intended task involves. In this evaluation, all world model configurations score at most 16.5 on Web Interaction and 13.3 on Game Interaction. MiniMax H3, for example, scores 53.8 on Embodiment but 3.7 on Web Interaction. Performance varies substantially across scenarios, but the scores alone cannot disentangle the effects of model capability, task difficulty, and scoring rules.
Overall rankings cannot replace scenario-level analysis either. GPT-Image-2 averages 21.7 across the eight scenarios, above MiniMax H3’s 17.3. On Embodiment, however, their scores are 20.9 and 53.8, respectively, reversing their order. This descriptive comparison shows that averages can mask differences across scenarios. Selecting components requires examining results relevant to the target task.
Even when both components perform well in the target scenario, their ability to work together still needs to be tested. Can the agent recognize an error generated by the world model? If it does, can it take effective corrective action? Two separate leaderboards cannot answer these questions. Different model pairings need to be tested under the same task conditions and interaction budgets.
Next Steps: From Identifying Problems to Testing Improvements
The results above offer clues for improving interactive systems, but whether proposed improvements work still needs to be tested. Experiments could be designed around the following three questions.
- Can checks during execution improve task completion rates? Under the same interaction budget, compare when checks are performed and what they examine. Record discrepancy detection, corrective responses, successful repairs, and final task completion separately to determine where checks are effective.
- Can more reasoning improve decisions in social interaction? Vary the reasoning budget while holding the model configuration and task conditions constant, and observe how decisions and outcomes change. Current comparisons between models also involve differences in scale, training methods, and reasoning settings, so performance differences cannot yet be attributed to “overthinking.”
- Can recording key information, such as object identities and counts, help models complete multi-step tasks? Check during execution whether this information has changed in ways it should not, then assess action effects, state preservation, and final completion rates separately to test whether this approach works.
Dual Exam offers a more concrete way to analyze failure: break “the task failed” down into actions, states, and outcomes that can be checked. Returning to the task of packing for a trip: did the passport, medicine case, and charger go into the original backpack? Was the backpack carried to the entryway? Did its contents meet the task requirements? Only by checking each of these conditions can we determine whether the task was truly completed and identify a basis for further improvement.