I obtained my Ph.D. degree in the Department of Computer Science and Engineering at the Chinese University of Hong Kong, supervised by Prof. Jiaya Jia and Prof. Bei Yu. Before that, I obtained my B.E. degree in Information Engineering at the College of Information Science and Electronic Engineering, Zhejiang University.
During my Ph.D. life, I have spent wonderful times collaborating with, among others:
In the industrial community, I conduct research and development on AGI.
Zhejiang Lab
I primarily led the construction of a 10,000-card GPU cluster and spearheaded the training of domain-specific large science models, in close collaboration with Academician Jian Wang, Prof. Hujun Bao, and Prof. Zhe Liu.
Huawei
My research and development focused on training unified inference models for generation and understanding, as well as advancing AIGC models including text-to-image and text-to-video. I successfully drove the deployment of these models into core Huawei products, and established strong academic collaborations with Prof. Hanwang Zhang from NTU Singapore.
MiroMind
I dedicated myself to the post-training of agentic foundation models and the development of self-evolving harnesses, working in close partnership with Prof. Shuicheng Yan from NUS Singapore.
In parallel, I have been a visiting researcher at a number of academic institutions. The full record of these positions is in Experiences.
Several specific topics of our current research interests and focus:
01Multi-modality data (image, video, 3D, etc.) generation & manipulation via AIGC
02Multi-modality large model and agent → AGI
03Generative Computational Photography: large model and efficiency optimization
04Security and alignment for large models / AGI
05Embodied AI (VLA, VA, etc.) and World Model
I carry out development and research in the industrial and academic communities at the same time. If you are interested in collaborating with me, please feel free to contact me through the email.
Selected Collaborations
News
[07/2026]Two papers are accepted by ACM MM 2026
[06/2026]We created a personal research project dedicated to forecasting the results of the "2026 World Cup" and the "2026 Gaokao (高考) scorelines". Explore the predictions on WOWCAI Website and let us know your thoughts! Report 今日头条
[06/2026]One paper is accepted by ECCV 2026
[06/2026]The Apodex (agentic model family) has been released across a diverse range of parameter sizes, Technical Report, Huggingface
[05/2026]Two papers are accepted by ICML 2026
[04/2026]We release the most comprehensive survey about the Latent Space of Large Models Reported by 机器之心
[04/2026]We release the Hierarchical Autonomy Evolution framework for Agent Security Reported by 新智元
[04/2026]One paper is accepted by TCSVT
[04/2026]Three papers are accepted by ACL 2026 (one main, two findings)
[03/2026]We release the most comprehensive survey about Intelligent Remote Sensing Agents Reported by 机器之心
[03/2026]We release the first Fairness benchmark for UMLLMs at ICLR 2026 Reported by 新智元
[03/2026]We release the agentic large model Mirothinker 1.7&H (Paper, Code)
[02/2026]One paper is accepted by TIP
[02/2026]Three papers are accepted by CVPR 2026 (two main, one finding, one workshop)
[01/2026]We release the first multi-agent RL training framework with technical report, MATPO-PR
[01/2026]One paper is accepted by Neurocomputing
[01/2026]Two papers are accepted by ICLR 2026
[01/2026]One paper is accepted by ICASSP 2026
[01/2026]One paper is accepted by TIP
News in 2025
[12/2025]We release the most robust MLLM Robust-R1 at AAAI 2026 Reported by 新智元 and CVer
[12/2025]I am selected as an area chair (AC) for ARR January/ACL 2026
[11/2025]I am selected as an area chair (AC) for ICML 2026
[11/2025]We release the agentic large model Mirothinker (Paper, Demo)
[11/2025]Two papers are accepted by AAAI 2026
[10/2025]One paper is accepted by IJCV
[09/2025]I was selected for the list of the world's top 2% scientists.
[09/2025]Two papers are accepted by NeurIPS 2025
[08/2025]One paper is accepted by SIGGRAPH ASIA 2025
[07/2025]One paper is accepted by ACM MM 2025
[07/2025]One paper is accepted by TIP
[07/2025]Two papers are accepted by ECAI 2025
[05/2025]Three papers are accepted by ICCV 2025
[06/2025]One paper is accepted by ACM Computing Surveys
[05/2025]One paper is accepted by ICCP
[05/2025]One paper is accepted by TMM
[05/2025]One paper is accepted by TPAMI
[05/2025]One paper is accepted by ICML 2025
[05/2025]One paper is accepted by IJCAI 2025
[05/2025]One paper is accepted by SIGGRAPH 2025
[05/2025]We release the new MLLM framework of selftok, Project page. I mainly lead the post-training stage, especially Reinforcement Learning. Its preceding work has received the Best Student Paper Honorable Mention in CVPR 2025.
[04/2025]One paper is accepted by ICMR 2025
[01/2025]One paper is accepted by ICLR 2025
[12/2024]Two papers are accepted by AAAI 2025
News in 2024
[11/2024]One paper is accepted by 3DV 2025
[09/2024]Two papers are accepted by NeurIPS 2024
[09/2024]One paper is accepted by EMNLP 2024
[09/2024]One paper is accepted by TVCG
[07/2024]Three papers are accepted by ECCV 2024
[06/2024]We release the demo and code of Sagiri, which is a representative model to incorporate restoration and AIGC, especially for HDR, Project page.
[06/2024]We release the demo and code of DepthAnything V2, which is a stronger open-world depth estimation model, Project page.
[05/2024]We release several works about large models (Model Safety, Embodied AI, Federated LLMs, VLM)