About Wenbo Ji
I am an M.Sc. student at the Technical University of Munich and a research intern at Agile Robots SE. I build generative and geometric models for interactive humans, dynamic 3D scenes, and robot learning.
My current work spans two threads: camera-controlled human motion video diffusion at the TUM Visual Computing Group with Yu Chi, Jiapeng Tang, and Prof. Matthias Nießner; and cross-embodiment video generation for dexterous manipulation at Agile Robots with Mahdi Mustapha Hamad.
I am seeking Ph.D. opportunities for Fall 2027 in 3D/4D scene representation, video generation, and robot world models.
Research direction
- Past. My earlier work built foundations in long-term tracking, 3D/4D reconstruction, and scene decomposition.
- Now. I work on camera-controlled human motion video generation and video world models for robot dexterous manipulation.
- Next. I want to unify these threads into perception-action models of interactive humans and dynamic scenes.
The broader goal is to connect geometry, generation, tracking, and interaction rather than treat them as separate endpoints: models that can reconstruct what is present, generate how it may change, and reason about how human or robot actions alter a world over time.
Selected publications
- Ji, W., Davoli, D., Chen, Z., Schoneveld, L., Nießner, M., & Tang, J. (2026). ViDS: Video Diffusion Shader using 3D Face Tracking. arXiv preprint arXiv:2607.24124.
- Xia, Y., Ji, W., Chen, W., & Cremers, D. (2025). CSG-Fusion: Consistent Sparse-View Gaussian Splatting via Matching-based Fusion. In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops (pp. 2653–2662).
- Karaoglu, M. A., Ji, W., Abbas, A., Navab, N., Busam, B., & Ladikos, A. (2025). LiteTracker: Leveraging Temporal Causality for Accurate Low-latency Tissue Tracking. In Medical Image Computing and Computer Assisted Intervention – MICCAI 2025 (Lecture Notes in Computer Science, Vol. 15969, pp. 308–317). Springer Nature Switzerland. https://doi.org/10.1007/978-3-032-05127-1_30
- Yan, X., Jiang, Z., Shuai, Y., Wang, N., Song, X., Ji, W., Wu, G., He, J., Wei, G., & Wang, Z. (2025). RE0: Recognize Everything with 3D Zero-Shot Instance Segmentation. In 2025 IEEE International Conference on Robotics and Automation (ICRA) (pp. 12655–12662). IEEE. https://doi.org/10.1109/ICRA55743.2025.11127468
Background
Before my current work, I worked on human-centric video generation at TUM Visual Computing. I also worked on dense point tracking at ImFusion and TUM CAMP, and on 3D scene reconstruction and decomposition with TUM-CVG and Oxford VGG. Earlier, I completed a research internship on large-scale 3D reconstruction at Zhejiang University under the supervision of Prof. Yiyi Liao.
Before joining TUM, I earned a master's degree in Computer Science from Tongji University and a bachelor's degree in Information and Computing Science from Nanjing Tech University.
This blog
This blog is my working notebook for ideas that are not finished yet. It is where I make research questions concrete, document the engineering decisions behind visual computing, video and world models, robotics, and AI systems, and occasionally write about work and life. My personal homepage holds the more formal record of publications and projects.