ACE-Data-0: Human-Centric Ambient Capture as Embodied Data Engine
A human-centric capture system and dataset with 150 hours, 17M video frames, 200 household task categories, 50 participants, two environments, and 75,000 interaction episodes.
Research Engineer · ACE Robotics
Computer vision · generative modeling · embodied intelligence
I am a Research Engineer at ACE Robotics, previously at Shanghai AI Laboratory, working at the intersection of computer vision, generative modeling, and embodied intelligence. My research focuses on video generation, 3D/4D visual computing, world models, and multimodal data engines for embodied agents, with an emphasis on understanding and generating dynamic visual worlds.
My recent work spans efficient video synthesis, 4D content generation and evaluation, digital-human datasets, and human–scene interaction. I work closely with Dr. Liang Pan and Prof. Ziwei Liu.
Research interests. Video generation · 3D/4D visual computing · world models · embodied AI.
Recent updates
Released HSImul3R: Physics-in-the-Loop Reconstruction of Simulation-Ready Human-Scene Interactions. Project
WorldLens was accepted as an oral presentation at CVPR 2026. Project · Paper
Hi3DEval was accepted to the NeurIPS 2025 Datasets and Benchmarks Track. Project
One paper was accepted to IEEE TDSC 2025.
Released Vchitect-2.0, a parallel transformer for scaling up video diffusion models. Code · Paper
Fast-Vid2Vid++ was accepted to IEEE TPAMI 2024. Paper
RenderMe-360 was accepted to the NeurIPS 2023 Datasets and Benchmarks Track. Project
One paper was accepted to IEEE TIFS 2022.
Research output
Selected papers, datasets, and benchmarks in chronological order. See the full list on Google Scholar →
A human-centric capture system and dataset with 150 hours, 17M video frames, 200 household task categories, 50 participants, two environments, and 75,000 interaction episodes.