I am an M.Sc. student at the Technical University of Munich and a research intern at Agile Robots SE. I build generative and geometric models for interactive humans, dynamic 3D scenes, and robot learning.
My current work studies camera-controlled human motion video diffusion with Prof. Matthias Nießner and cross-embodiment video generation for dexterous manipulation.
I am seeking Ph.D. opportunities for Fall 2027 in 3D/4D scene representation, video generation, and robot world models.
3D face tracking-conditioned video diffusion for expressive, identity-preserving portrait animation from a single image, with autoregressive sampling for longer videos. On VFHQ, ViDS ranked first on 8 of 13 reported metrics.
3D/4D Scene Representation
CSG-Fusion: Consistent Sparse-View Gaussian Splatting via Matching-based Fusion
Yan Xia*†, Wenbo Ji*, Weirong Chen, Daniel Cremers
Overview
Matching-based fusion of sparse-view pointmaps into compact, cross-view-consistent 3D Gaussians. At 90% ScanNet++ overlap, it improved PSNR by 2.8 dB over Splatt3R while using approximately 124K fewer Gaussians.
Dynamic Visual Perception
LiteTracker: Leveraging Temporal Causality for Accurate Low-latency Tissue Tracking
MICCAI, 2025
Authors
Mert Asim Karaoglu, Wenbo Ji, Ahmed Abbas, Nassir Navab, Benjamin Busam, Alexander Ladikos†
Overview
Causal temporal feature reuse with prior-motion initialization for accurate, low-latency online tissue tracking. It ran approximately 7× faster than its predecessor and 2× faster than prior state of the art, reaching 29.67 ms P95 for 1,024 points.
Dynamic Visual Perception
RE0: Recognize Everything with 3D Zero-shot Instance Segmentation
ICRA, 2025
Authors
Xiaohan Yan*, Zijian Jiang*, Yinghao Shuai*, Nan Wang, Xiaowei Song, Wenbo Ji, Ge Wu, Jinyu He, Gang Wei, Zhicheng Wang†
Overview
Training-free 3D zero-shot instance segmentation from multi-view masks and CLIP semantics.
Experiences
Embodied World Models
Video World Model for Robot Dexterous ManipulationApr 2026 - Now
Research Internship
Contributions
Developing a cross-embodiment video generation method that translates egocentric human demonstrations into robot-domain videos for downstream policy learning.
Mentors
Mahdi Mustapha HamadAgile Robots SE / WRD Group
Human-Centric Video Generation
Human Motion Video DiffusionMarch 2026 - Now
Master's Thesis
Contributions
Developing a camera-controlled video diffusion model for controllable synthesis of human motion and scene interactions across changing viewpoints.
Endoscopic Scene Reconstruction with 4D Half Gaussian Splatting
2025
Master's Thesis
Overview
Developed a 4D Half-Gaussian splatting pipeline for deformable stereo endoscopic reconstruction with depth-prior initialization, HexPlane spatiotemporal deformation, and edge-aware depth regularization. Achieved 38.1 PSNR on EndoNeRF versus prior endoscopic GS/NeRF baselines; also evaluated on SCARED.
Technical Report
3D/4D Scene Representation
Object-Centric 3D Reconstruction and Decomposition
2025
TUM DI Lab Report
Authors
Wenbo Ji, Michael Neumayr, Nina Kirakosyan, Filip Skubacz
Overview
A TUM DI Lab report on object-centric 3D reconstruction and decomposition with 3D Gaussian Splatting.
Education
M.Sc. Electrical Engineering and Information Technology
Technical University of Munich2023 – Now
M.Sc. Electrical Engineering and Information Technology
Double-degree program with Tongji University (Tongji M.Sc. awarded 2025). Thesis: camera-controlled human motion video diffusion at the Visual Computing Group.
M.Sc. Computer Science
Tongji University2021 – 2025
M.Sc. Computer Science
Thesis: endoscopic scene reconstruction with 4D half-Gaussian splatting.
B.Sc. Information and Computing Science (Embedded Software)
Nanjing Tech University2017 – 2021
B.Sc. Information and Computing Science (Embedded Software)
A computing major within the Department of Mathematics.
Selected Awards
2023 - 2024Munich, Germany
Deutscher Akademischer Austauschdienst (DAAD) Scholarship
Recognition
2019Nanjing, China
National Encouragement Scholarship
Recognition
Projects
InfraLens
May 2026
Focus
AI Infrastructure Handbook
Overview
A static handbook for understanding how modern AI systems train, serve, generate, route, compress, and fail.
OpenUserStudyKit
March 2026
Focus
Reusable User Study Infrastructure
Overview
An open-source toolkit for building reusable user study questionnaires and experiment workflows.
LiteAvatar - WASM Version
Jan 2026
Focus
2D Audio-driven Human Avatar Animation
Overview
A lightweight audio-driven 2D avatar solution that runs entirely in the browser using WASM based on Lite-avatar. No backend server required.
Blog
Thoughts on research, 3D, video generation, and the occasional in-between.