About Wenbo Ji

I am an M.Sc. student at the Technical University of Munich and a research intern at Agile Robots SE. I build generative and geometric models for interactive humans, dynamic 3D scenes, and robot learning.

My current work spans two threads: camera-controlled human motion video diffusion at the TUM Visual Computing Group with Yu Chi, Jiapeng Tang, and Prof. Matthias Nießner; and cross-embodiment video generation for dexterous manipulation at Agile Robots with Mahdi Mustapha Hamad.

I am seeking Ph.D. opportunities for Fall 2027 in 3D/4D scene representation, video generation, and robot world models.

Research direction

The broader goal is to connect geometry, generation, tracking, and interaction rather than treat them as separate endpoints: models that can reconstruct what is present, generate how it may change, and reason about how human or robot actions alter a world over time.

Selected publications

Background

Before my current work, I worked on human-centric video generation at TUM Visual Computing. I also worked on dense point tracking at ImFusion and TUM CAMP, and on 3D scene reconstruction and decomposition with TUM-CVG and Oxford VGG. Earlier, I completed a research internship on large-scale 3D reconstruction at Zhejiang University under the supervision of Prof. Yiyi Liao.

Before joining TUM, I earned a master's degree in Computer Science from Tongji University and a bachelor's degree in Information and Computing Science from Nanjing Tech University.

This blog

This blog is my working notebook for ideas that are not finished yet. It is where I make research questions concrete, document the engineering decisions behind visual computing, video and world models, robotics, and AI systems, and occasionally write about work and life. My personal homepage holds the more formal record of publications and projects.