UrbanGround: From Local Perception to Spatial Agency in a Real-Scale City
A real-scale urban sandbox that turns territory-wide 3D geospatial data of Hong Kong into a continuously rendered and physically grounded environment for multimodal agents.






01
Want to explore UrbanGround yourself?
Click the scene to load the complete browser edition here—no download or separate page required.
Control mode
First visit: during periods of high visitor traffic or on slower connections, loading the complete Unity scene may take 3–5 minutes. It is cached locally for faster return visits. For faster access, download the project from GitHub and serve the tile datasets over a local network.
02
Explore, navigate, and interact at city scale.
03
Hong Kong at real scale.
UrbanGround is a real-scale urban sandbox built from territory-wide 3D geospatial data. It supports direct first-person play and programmatic control by MLLM agents through the same interface. We release the sandbox on the web and as native builds for macOS, Windows, and Linux. It also includes diverse tasks for studying how multimodal agents perceive and act in a real city.
04
Diverse urban environments.
Figure 2 groups the dynamic simulation by time of day, weather, and pedestrian activity. Select a condition to inspect the original panel from the paper.
The same Hong Kong viewpoint under clear daytime illumination.
05
From seeing a street to adapting within a city.
The experimental data follows a five-level progression from local understanding to explicit navigation, implicit goal inference, multi-task planning, and interaction with environmental change.
The evaluation exposes only task instructions, first-person observations, physical controls, and map interaction. It does not provide hidden coordinates, remaining distance, or privileged simulator state.
06
Coverage and composition.
The experimental data spans geographically diverse regions of Hong Kong while distributing thirteen task types across five capability levels. These views summarize where tasks occur and how the evaluation set is composed.