Command Palette
Search for a command to run...
MULTI-AGENT EGOCENTRIC WORLD MODEL WITH FINE-GRAINED EMBODIED INTERACTION
MULTI-AGENT EGOCENTRIC WORLD MODEL WITH FINE-GRAINED EMBODIED INTERACTION
Dahyun Chung Siyoon Jin Hyunwook Choi Honggyu An Junyoung Seo Hyunsung Kim Seung Wook Kim Seungryong Kim
Abstract
Egocentric world models predict first-person observations conditioned on an agent’s actions, but most focus on a single agent. Real embodied settings often involve multiple agents that act and interact within a shared environment. Existing multi-agent world models rely on coarse actions like locomotion, camera control, or discrete commands, leaving fine-grained embodied interactions underexplored. We formulate multi-agent egocentric world modeling as synchronized ego-stream generation for multiple agents interacting through fine-grained actions in a shared world. This requires cross-view action consistency, sharedenvironment consistency, and consistent propagation of interaction-induced state updates. We propose Multi-agent Egocentric World Model (ME-World), which jointly denoises multiple ego streams in a shared token sequence, conditions each stream on all agents’ target-view poses, and grounds generation with shared environment memory. We train and evaluate on real and synthetic multi-agent data and introduce shared-world consistency metrics for environment, update, and identity consistency. Experiments show ME-World improves shared-world consistency, action control, identity preservation, and video quality over existing methods.
One-sentence Summary
Researchers from KAIST AI propose the Multi-agent Egocentric World Model (ME-World), which jointly denoises multiple egocentric streams in a shared token sequence, conditions each stream on all agents’ target-view poses, and grounds generation with shared environment memory, thereby improving shared-world consistency, action control, identity preservation, and video quality over existing multi-agent world models.
Key Contributions
- The paper introduces embodied multi-agent world modeling as synchronized egocentric stream generation for multiple agents interacting through fine-grained actions in a shared world.
- The paper proposes ME-World, which jointly denoises multiple ego streams in a shared token sequence, conditions each stream on all agents' target-view poses, and grounds generation with shared environment memory.
- The paper constructs real and synthetic multi-agent benchmarks with shared-world consistency metrics for environment, update, and identity; experiments show ME-World improves shared-world consistency, action control, identity preservation, and video quality over existing methods.
Introduction
Video world models enable agents to predict future observations conditioned on their actions, with recent egocentric approaches extending this idea to first-person embodied streams. In multi-agent settings, however, prior work typically assumes a single embodied ego or abstracts actions into low-dimensional controls, which misses the dual role of fine-grained body, hand, and head motion as both first-person control and cross-view observable behavior. This limits the ability to keep cross-view actions synchronized, shared environments consistent, and interaction-induced object or state changes aligned across all views. The authors introduce ME-World, a multi-agent egocentric world model that jointly denoises all ego streams in a shared token sequence, conditions on shared actions, and grounds generation in shared environment memory and common anchor references. They also combine real and synthetic multi-ego data and propose shared-world consistency metrics to evaluate whether generated streams represent the same evolving shared world.
Method
The authors address the challenge of embodied multi-agent world modeling, where multiple agents act and interact within a shared, evolving environment while observing it from their own egocentric viewpoints. Given the future embodied action sequences, text prompts, initial ego frames, and observation histories for N agents, the goal is to generate future ego streams for all agents. Each generated stream must follow its respective action sequence while remaining globally consistent with the shared environment and world state. For the n-th agent, the future ego stream and action sequence are denoted as X(n)={xt(n)}t=1T and A(n)={at(n)}t=1T, respectively. The observation histories of all agents are pooled into a shared observation-history pool E. The full conditioning signal is defined as C={E,A(1:N),P(1:N),x0(1:N)}. To ensure that the generated streams agree on the shared environment and interaction-induced state changes, the framework couples the ego streams through joint multi-agent generation, guides them with shared action conditioning, and grounds them with shared environment memory.
As shown in the figure below:
Joint Multi-Agent Generation Because agents interact within the same environment, their ego streams are inherently dependent. Rather than generating each stream independently, the model jointly denoises their latent sequences within a single pretrained video diffusion transformer. This approach allows cross-stream information exchange through self-attention, ensuring that all streams evolve as mutually informed observations of the same shared world. Formally, the model learns the joint conditional distribution pθ(X(1),…,X(N)∣C).
Using a frozen video VAE, each ego stream X(n) is encoded into a latent sequence Z(n). During denoising, the noisy latents of all agents at timestep τ are concatenated along the token dimension:
Zτgen=[Zτ(1);…;Zτ(N)].The transformer processes this joint sequence using self-attention. To maintain spatial and temporal alignment, the authors apply identical rotary positional embeddings to all streams, aligning their space-time positions in a common positional frame. They distinguish the streams by adding a learnable agent embedding e(n) to the tokens of the n-th agent. For text conditioning, each stream utilizes a separate cross-attention path, allowing each agent to follow its own prompt P(n) while sharing visual information through joint self-attention.
Shared Action Conditioning Joint generation enables information exchange, but the model requires explicit guidance on how each agent should move and interact. A naive approach of conditioning each stream solely on its own private action leads to inconsistent rendering of embodied interactions across viewpoints. To resolve this, the authors implement shared action conditioning, guiding each ego stream with the motion of all agents.
For the n-th agent, the action at frame t is decomposed into body motion at,body(n) and head motion at,head(n). The body motion is encoded as a self-action skeleton map capturing ego-visible hand and arm motion. The head motion is represented with Plucker ray maps, providing a geometric encoding of the induced camera rays.
To construct the shared action condition for the n-th agent, a pose condition sequence Xpose(n) is defined. At each frame, this is rendered on the target image plane using the agent's self-action skeleton map, the 3D body poses of other agents, and relative camera geometry. This unified pose condition represents both the agent's own motion and the visible motions of other agents in its ego view. For head motion, Plucker ray maps Xray(n) are collected and expressed in a shared canonical frame, making camera trajectories geometrically comparable across streams. The encoded pose and ray features are concatenated with the noisy latent Zτ(n) along the channel dimension for joint denoising.
Shared Environment Memory While shared action conditioning makes agent movements explicit, it does not define the surrounding world layout or appearance. The authors introduce shared environment memory from the pooled observation history E to provide a common reference. This memory is integrated via two complementary paths.
The stream-aligned geometric memory path transforms relevant observations into target-view conditions. For each target stream, the history image Xhist that provides the largest valid coverage of the target view is selected and warped according to the target ray condition Xray(n). Dynamic agent regions are masked during projection to avoid transferring outdated poses. This warping yields a warped RGB condition Xwarp(n) and a visibility mask Mvis(n), which are encoded and concatenated with the noisy latent to directly condition denoising on aligned scene evidence.
The cross-stream anchor memory path complements this by preserving selected history images in their original, unwarped form as shared scene references. The authors greedily select K history observations to maximize target-view coverage. These selected observations are encoded into reference features Zref={zrefk}k=1K and appended to the joint denoising sequence along the token dimension. The extended input joint sequence to the transformer becomes:
Zτjoint=[Zτgen;Zref],allowing all ego streams to attend to the same anchors and grounding them in the same environment.
Training Objective The authors fine-tune the pretrained model using rectified flow. Let Z0gen=[Z0(1);…;Z0(N)] denote the clean latents of all agents, and ϵgen=[ϵ(1);…;ϵ(N)] where ϵ(n)∼N(0,I). For τ∼U(0,1), the noisy latent is constructed as Zτgen=(1−τ)Z0gen+τϵgen. The model fθ is trained by minimizing the flow matching loss:
LFM=EZ0,ϵ,τ,C[fθ(Zτjoint,τ,C)−(ϵgen−Z0gen)22].Experiment
The paper evaluates ME-World on real and synthetic multi-agent egocentric benchmarks using shared-world consistency metrics for unseen environment regions, state updates, and counterpart identity, along with camera control, action control, and video quality measures. Comparisons against multi-view, single-ego, general world, and adapted multi-agent baselines show that existing methods struggle with cross-view environment and interaction consistency, while ME-World achieves the strongest consistency, action control, and visual quality. Ablations confirm that joint multi-agent generation couples ego streams, shared action conditioning aligns embodied behavior, and shared environment memory anchors backgrounds to a common evolving world.
On the real benchmark, existing methods show trade-offs across shared-world consistency, camera control, action control, and video quality: multi-view baselines can still produce inconsistent scenes and action artifacts, single-ego models lack other-agent conditioning, and general world models often fail to keep ego streams aligned. On the synthetic benchmark, baselines tend to suffer from identity drift and mismatched motions across ego views. ME-World yields stronger shared-world consistency and improves camera control, self- and other-agent action control, and video quality. Existing real-benchmark methods trail on shared-world consistency and action alignment, and multi-view baselines still show blurry or ghosted actions despite using a reference ego stream. Single-ego world models exhibit weak other-agent action control, while general world models often fail to maintain consistent environment and interaction outcomes across ego streams. On synthetic data, baselines commonly lose counterpart identity and produce inconsistent interactions, whereas ME-World preserves identity and achieves stronger camera control, action control, and video quality.
In a controlled real-benchmark comparison of multi-agent world models, ME-World outperforms all baselines across shared-world consistency, camera control, self and other-agent action control, and video quality. MetaWorld is the strongest baseline, particularly for shared-world consistency and video quality, while Solaris and γ-World show larger drops in consistency and fine-grained action control. MultiWorld maintains relatively strong action control but still trails ME-World overall. ME-World achieves the best shared-world consistency, camera control, self- and other-agent action control, and video quality among all compared methods. Among the adapted baselines, MetaWorld shows the strongest shared-world consistency and video quality, whereas Solaris and γ-World exhibit marked degradation in shared-world consistency and fine-grained action control.
The adapted architectures differ mainly in how they handle joint multi-agent generation and shared environment memory. ME-World is the only method that includes explicit counterparts for joint multi-agent generation, shared action conditioning, and shared environment memory. This full combination aligns with its reported advantages in shared-world consistency, action control, and video quality. ME-World is the only adapted architecture with explicit support for joint multi-agent generation, shared action conditioning, and shared environment memory. MultiWorld and MetaWorld replace shared environment memory with method-specific components, while Solaris and γ-World lack an explicit shared environment memory module. The architectural differences align with ME-World's stronger shared-world consistency, action control, and video quality relative to adapted multi-agent baselines.
Fine-tuning the base model alone still produces implausible multi-agent ego futures, with weak shared-world consistency and action control. Adding self-hand pose and history warping brings the largest gains in camera control, action metrics, and video quality, while joint generation and shared action conditioning further improve multi-view consistency and action alignment. The shared-action variant achieves the strongest shared-world consistency and full-body action control among the evaluated rows. Off-the-shelf and fine-tuned Cosmos-Predict2.5 baselines show weak shared-world consistency and poor multi-agent action and video quality. Self-hand pose with history warping yields the largest improvement in camera control, hand and full-body action metrics, and video quality. Joint generation further improves shared-world consistency and visual quality, and shared action conditioning gives the best full-body action alignment and shared-world consistency. Shared environment memory is reported to anchor streams to a common evolving world, addressing background inconsistencies.
The experiments evaluate multi-agent world models on real and synthetic benchmarks, compare adapted architectures, and ablate components of ME-World. On both benchmarks, existing baselines show trade-offs such as cross-view inconsistency, weak other-agent conditioning, identity drift, and mismatched interactions, while ME-World consistently achieves stronger shared-world consistency, camera and action control, and video quality. Architecture comparisons indicate ME-World is the only method with explicit joint multi-agent generation, shared action conditioning, and shared environment memory, and ablations show that hand-pose and history warping provide the largest control gains, with joint generation and shared action conditioning further improving multi-view consistency and action alignment.