Google DeepMind introduced D4RT, a unified AI model for 4D scene reconstruction and tracking across space and time. The model addresses the challenge of recovering volumetric 3D worlds in motion from flat 2D video projections. D4RT tracks every pixel of every object as it moves through three spatial dimensions and time, while disentangling object motion from camera motion. Traditionally, this required computationally intensive processes or multiple specialized AI models for depth, movement, and camera angles.
No score is assigned. Sources and their independence are shown in the citation chain below.