Consistent Modeling of 4D Scenes for Perception and Generation

August 2026

Consistent Modeling of 4D Scenes for Perception and Generation

Authors:

Jinhyung Park

Abstract:

Learning scene representations that capture 3D structure through time is important for both perception and interactive generation. Perception estimates the state of the world from observations; generation creates new, plausible states from scratch. For both tasks, we want consistency across space, over time, and across modalities, so that a single scene state does not drift across observations, viewpoints, or rollouts. This thesis explores a single question: what representation allows us to consistently model 4D driving scenes, for both perception and generation?

Our works trace a path through three families of representations: dense grids, boxes, and sparse queries. We begin with dense temporal voxel grids, leveraging consistency across space: enforcing agreement of image features across many views of the same 3D location delivers strong multi-camera 3D detection through long-horizon temporal stereo, but also reveals recurring difficulties of grid-centric modeling: a grid has no notion of individual entities, covers the full scene only at considerable cost, and its features smear when propagated through time. Studying depth completion, we turn to consistency across modalities: learning pixel-point affinities lets a single model explain an RGB image together with depth measurements of any density and pattern, though the resulting groupings remain purely geometric and single-frame. With box-centric representations, in semi-supervised 2D-3D and temporal object detection, we use consistency between modalities and over time as supervision directly: enforcing agreement between a scene's 2D and 3D detections, and across neighboring frames, produces better pseudo-labels and stronger detectors. Boxes, however, are coarse and cover only the foreground. These threads come together in S2GO, which summarizes the scene into a compact set of streaming, grounded queries that decode to semantic Gaussians: an entity-centric, full-scene representation that stays consistent across space, time, and modalities, and achieves state-of-the-art semantic occupancy perception at real-time speeds.

We then explore whether the same representation can support generation, where the scene state is created rather than estimated. We factorize 4D scene generation around these sparse grounded 3D latents: a layout diffusion stage generates latent positions, classes, and orientations; a feature diffusion stage generates per-latent features encoding local geometry; and a motion diffusion stage evolves the persistent latents through time, with ego motion applied as an explicit rigid transform of the latent set. Each latent decodes to semantic Gaussians that splat to complete semantic occupancy, so editing the world reduces to editing latents: individual actors can be translated or rotated directly, and the merging, flickering, and splitting artifacts common to dense voxel generators are substantially reduced. The resulting framework, LatentWorld, achieves state-of-the-art 3D and 4D generation quality on the CarlaSC and Waymo datasets. With this, every part of our question is addressed by one family of representation: a sparse set of grounded, entity-centric latents. More broadly, we hope that a single scene state serving both perception and generation opens the door to unified, entity-centric world models: a state estimated from sensors that can be rolled forward, edited, and imagined, supporting prediction, planning, and simulation for autonomous systems.
@phdthesis{Park-2026-88363,
author = {Jinhyung Park},
title = {Consistent Modeling of 4D Scenes for Perception and Generation},
year = {2026},
month = {August},
school = {Carnegie Mellon University},
address = {Pittsburgh, PA},
number = {CMU-RI-TR-26-104},
keywords = {4D Scene Generation, 4D Scene Understanding, Autonomous Driving, Gaussian Splatting, World Models},
}
Copyright notice: This material is presented to ensure timely dissemination of scholarly and technical work. Copyright and all rights therein are retained by authors or by other copyright holders. All persons copying this information are expected to adhere to the terms and constraints invoked by each author's copyright. These works may not be reposted without the explicit permission of the copyright holder.