Wenbo Ji   jī / yí文博

I am an M.Sc. student at the Technical University of Munich and a research intern at Agile Robots SE. I build generative and geometric models for interactive humans, dynamic 3D scenes, and robot learning.

My current work studies camera-controlled human motion video diffusion with Prof. Matthias Nießner and cross-embodiment video generation for dexterous manipulation.

I am seeking Ph.D. opportunities for Fall 2027 in 3D/4D scene representation, video generation, and robot world models.

Previously, I collaborated with Prof. Daniel Cremers, Prof. Nassir Navab, and Prof. Benjamin Busam on 3D reconstruction and tracking. I earned an M.Sc. in Computer Science from Tongji University and a B.Sc. in Information and Computing Science from Nanjing Tech University.

News

I joined Agile Robots SE as a research intern working on world models for robot dexterous manipulation.

Our paper CSG-Fusion received the Best Paper Award at the ICCV 2025 Workshop E2E3D (workshop track).

I graduated from Tongji University with a Master's degree in Computer Science.

More news

Our paper LiteTracker is accepted by MICCAI 2025.

Our paper RE0 is accepted by ICRA 2025.

I joined ImFusion as a research intern on dense video point tracking.

Research Interests

  1. Past

    Built foundations in long-term tracking, 3D/4D reconstruction, and scene decomposition.

  2. Now

    Working on camera-controlled human motion video generation and video world models for robot dexterous manipulation.

  3. Next

    Unifying these threads into perception-action models of interactive humans and dynamic scenes.

3D/4D Scene Representation Human-Centric Video Generation Dynamic Visual Perception Embodied World Models

Selected Publications

* equal contribution, † corresponding author

ViDS teaser showing identity-preserving portrait animation driven by 3D face normal maps
Human-Centric Video Generation

ViDS: Video Diffusion Shader using 3D Face Tracking

3D face tracking-conditioned video diffusion for expressive, identity-preserving portrait animation from a single image, with autoregressive sampling for longer videos. On VFHQ, ViDS ranked first on 8 of 13 reported metrics.

CSG-Fusion teaser image
3D/4D Scene Representation

CSG-Fusion: Consistent Sparse-View Gaussian Splatting via Matching-based Fusion

Best Paper Award

Matching-based fusion of sparse-view pointmaps into compact, cross-view-consistent 3D Gaussians. At 90% ScanNet++ overlap, it improved PSNR by 2.8 dB over Splatt3R while using approximately 124K fewer Gaussians.

LiteTracker teaser image
Dynamic Visual Perception

LiteTracker: Leveraging Temporal Causality for Accurate Low-latency Tissue Tracking

Causal temporal feature reuse with prior-motion initialization for accurate, low-latency online tissue tracking. It ran approximately 7× faster than its predecessor and 2× faster than prior state of the art, reaching 29.67 ms P95 for 1,024 points.

RE0 teaser image
Dynamic Visual Perception

RE0: Recognize Everything with 3D Zero-shot Instance Segmentation

Training-free 3D zero-shot instance segmentation from multi-view masks and CLIP semantics.

Experiences

Agile Robots SE logo
Agile WRD logo
Embodied World Models

Video World Model for Robot Dexterous Manipulation

Research Internship
  • Developing a cross-embodiment video generation method that translates egocentric human demonstrations into robot-domain videos for downstream policy learning.
Mentors
Mahdi Mustapha Hamad Agile Robots SE / WRD Group
TUM Visual Computing Lab logo
Human-Centric Video Generation

Human Motion Video Diffusion

Master's Thesis
  • Developing a camera-controlled video diffusion model for controllable synthesis of human motion and scene interactions across changing viewpoints.
Mentors
TUM Visual Computing Lab logo
Human-Centric Video Generation

Human Head Avatar Animation

Research Internship
  • Led the development of ViDS, an identity-preserving video diffusion method for long-form portrait animation.
  • ViDS ranked first on 8 of 13 VFHQ metrics, improving reenactment quality and identity preservation.
Mentors
Jiapeng Tang, Prof. Matthias Nießner TUM Visual Computing
ImFusion logo
TUM CAMP logo
Dynamic Visual Perception

Dense Point Tracking

Research Internship
  • Implemented LiteTracker’s online inference and EMA-flow initialization, and conducted low-latency tracking experiments.
  • LiteTracker remained competitive on STIR and SuPer while running ~7× faster than its predecessor and 2× faster than prior state of the art.
Mentors
University of Oxford logo
Technical University of Munich logo
TUM DI Lab logo
3D/4D Scene Representation

3D Scene Decomposition

Guided Research
  • Designed and evaluated CSG-Fusion; received Best Paper at ICCV Workshop E2E3D.
  • Improved ScanNet++ PSNR by 2.8 dB over Splatt3R at 90% overlap with ~124K fewer Gaussians, and demonstrated zero-shot generalization on DTU.
Zhejiang University logo
3D/4D Scene Representation

Large Scale 3D Scene Reconstruction

Research Assistant
  • Contributed to the development and experimental evaluation of a large-scale 3D scene reconstruction pipeline.
Mentors
Prof. Yiyi Liao Zhejiang University

Thesis

Thesis reconstruction pipeline diagram
3D/4D Scene Representation

Endoscopic Scene Reconstruction with 4D Half Gaussian Splatting

Master's Thesis

Developed a 4D Half-Gaussian splatting pipeline for deformable stereo endoscopic reconstruction with depth-prior initialization, HexPlane spatiotemporal deformation, and edge-aware depth regularization. Achieved 38.1 PSNR on EndoNeRF versus prior endoscopic GS/NeRF baselines; also evaluated on SCARED.

Technical Report

Cover of the Object-Centric 3D Reconstruction and Decomposition technical report
3D/4D Scene Representation

Object-Centric 3D Reconstruction and Decomposition

TUM DI Lab Report

A TUM DI Lab report on object-centric 3D reconstruction and decomposition with 3D Gaussian Splatting.

Education

Technical University of Munich logo

M.Sc. Electrical Engineering and Information Technology

Technical University of Munich

M.Sc. Electrical Engineering and Information Technology

Double-degree program with Tongji University (Tongji M.Sc. awarded 2025).
Thesis: camera-controlled human motion video diffusion at the Visual Computing Group.

Tongji University logo

M.Sc. Computer Science

Tongji University

M.Sc. Computer Science

Thesis: endoscopic scene reconstruction with 4D half-Gaussian splatting.

Nanjing Tech University logo

B.Sc. Information and Computing Science (Embedded Software)

Nanjing Tech University

B.Sc. Information and Computing Science (Embedded Software)

A computing major within the Department of Mathematics.

Selected Awards

Munich, Germany

Deutscher Akademischer Austauschdienst (DAAD) Scholarship

Recognition
Nanjing, China

National Encouragement Scholarship

Recognition

Projects

InfraLens slogan banner

InfraLens

AI Infrastructure Handbook

A static handbook for understanding how modern AI systems train, serve, generate, route, compress, and fail.

OpenUserStudyKit slogan banner

OpenUserStudyKit

Reusable User Study Infrastructure

An open-source toolkit for building reusable user study questionnaires and experiment workflows.

LiteAvatar WASM slogan banner

LiteAvatar - WASM Version

2D Audio-driven Human Avatar Animation

A lightweight audio-driven 2D avatar solution that runs entirely in the browser using WASM based on Lite-avatar. No backend server required.

Blog

Thoughts on research, 3D, video generation, and the occasional in-between.

View all posts
Latest posts are available on the blog page.