PhD student in Computer Science and Technology at Beijing Jiaotong University, advised by Prof. Yunchao Wei. I am also a research intern at MT Lab, Meitu.
My research focuses on visual generative AI, with an emphasis on:
- Image generation and editing
- Diffusion models and Diffusion Transformers (DiT)
- Unified multimodal generation
- Data curation, model post-training, and evaluation
- Intelligent visual creation systems
- StructGen: Disambiguating Multi-Reference Image Generation via Structured Context Modeling — ACM MM 2026. Structured context modeling for precise attribute-subject grounding in multi-reference generation. [page] [paper] [code]
- CharaConsist: Fine-Grained Consistent Character Generation — ICCV 2025. A training-free method for fine-grained character consistency in DiT-based generation. [paper] [code]
- DCEdit: Dual-Level Controlled Image Editing via Precisely Localized Semantics — ACM MM Workshop 2025. Precise local image editing through dual-level semantic control. [paper] [code]
- On Exact Editing of Flow-Based Diffusion Models — Under Review, AAAI 2027. Exact image editing through flow-based diffusion modeling. [paper]
Python · PyTorch · Diffusers · Transformers · C++ · LaTeX