I am pursuing my Ph.D. in computer science and technology at the PCA Lab, Nanjing University of Science and Technology, under the supervision of Professor Jin Xie, and the lab is headed by Professor Jian Yang, and I am expected to graduate in March 2027. Previously, I completed my bachelor's degrees in electrical engineering & automation and master's degrees in control science & engineering at the B-DAT Lab, Nanjing University of Information Science and Technology, under the supervision of Professor Kaihua Zhang, with the lab directed by Professor Qingshan Liu.

My research interests lie in multimodal open-world computer vision, with a particular focus on embodied AI, generative 3D scene reconstruction and simulation, as well as world-model-driven foreseeing and decision-making. Previously, I also worked on 2D vision, focusing on open-world image segmentation and depth estimation.

I’m currently exploring opportunities in both academia (postdoc) and industry (research scientist / engineer roles). Feel free to reach out if you’re interested.

📢 News

  • 2026.02: Two paper was accepted by CVPR 2026.
  • 2026.01: One paper was accepted by ICRA 2026.
  • 2025.09: One paper was accepted by IEEE Transactions on Multimedia.
  • 2025.02: One paper was accepted by CVPR 2025.
  • 2024.07: One paper was accepted by ECCV 2024.
  • 2024.04: One paper was accepted by Frontiers of Computer Science.
  • 2024.03: One paper was accepted by ICME 2024.
  • 2023.10: One paper was accepted by IEEE Geoscience and Remote Sensing Letters.
  • 2023.07: Our paper "Object-Aware Calibrated Depth-Guided Transformer for RGB-D Co-Salient Object Detection" was selected as the Best Student Paper 🏆 of ICME 2023 (1/1413).
  • 2023.03: One paper was accepted by CVPR 2023.
  • 2023.03: One paper was accepted by ICME 2023.
  • 2023.02: One paper was accepted by ICASSP 2023.
  • 2022.12: One paper was accepted by Chinese Journal of Computers.
  • 2022.08: One paper was accepted by IEEE Transactions on Multimedia.

📝 Selected Publications

CoPhy arXiv 2026
Distill to Think, Foresee to Act: Cognitive-Physical Reinforcement Learning for Autonomous Driving, arXiv2026
Yang Wu, Qiang Meng, Zhaojiang Liu, Youquan Liu, Jian Yang, Jin Xie
  • We ground cognitive priors into spatial perception at zero inference cost via cognitive prior distillation.
  • Our method can integrate an explicit BEV world model to perform cognitive-physical policy optimization, breaking through the ceiling of imitation learning.
  • Our method can function as a controllable yet safe system capable of interpreting diverse human instructions.
GEM CVPR 2026
GEM: Generating LiDAR World Model via Deformable Mamba, CVPR2026 (Code, Project)
Yang Wu, Zhaojiang Liu, Qiang Meng, Youquan Liu, Renliang Weng, Jianjun Qian, Jian Yang, Jin Xie
  • We introduce a LiDAR scene tokenizer, designed for latent representation of LiDAR data.
  • We present a LiDAR world model with explicit dynamic-static disentanglement in driving scenes.
  • Our framework makes a pioneering contribution to LiDAR world models by enabling "what-if" reasoning.
T2LDM CVPR 2026
A Self-Conditioned Representation Guided Diffusion Model for Realistic Text-to-LiDAR Scene Generation, CVPR2026 (Code)
Wentao Qu, Guofeng Mei, Yang Wu, Yongshun Gong, Xiaoshui Huang, Liang Xiao
  • We propose a Text-to-LiDAR Diffusion Model, T2LDM, with a self-conditioned representation guidance.
  • We construct a high-quality content-composable Text-LiDAR benchmark, T2nuScenes.
  • By leveraging a directional position prior, T2LDM alleviates road distortion, further improving scene fidelity.
SMAFormer IEEE TMM 2026
Learning Semantic-level Multi-modal Alignment Transformer for RGB-D Co-salient Object Detection, IEEE TMM2026
Yang Wu, Shenglong Hu, Kaihua Zhang, Lingyan Liang, Yaqian Zhao, Gang Dong
  • Our SMAFormer utilizes semantic information to calibrate the depth maps through learnable semantic-level clustering.
  • We design a SAM guided cluster center extractor to dynamically generate adaptive cluster centers from RGB images.
  • We further design a Recurrent Deep Clustering Module to capture patch-level associations iteratively.
WeatherGen CVPR 2025
WeatherGen: A Unified Diverse Weather Generator for LiDAR Point Clouds via Spider Mamba Diffusion, CVPR2025 (Code)
Yang Wu, Yun Zhu, Kaihua Zhang, Jianjun Qian, Jin Xie, Jian Yang
  • We propose WeatherGen, the first unified framework for generating diverse weather LiDAR data.
  • Spider Mamba is devised, which can model the LiDAR feature interactions effectively in a way that can best maintain the physical structure of the LiDAR data.
Text2LiDAR ECCV 2024
Text2LiDAR: Text-guided LiDAR Point Cloud Generation via Equirectangular Transformer, ECCV2024 (Code)
Yang Wu, Kaihua Zhang, Jianjun Qian, Jin Xie, Jian Yang
  • We propose the first effective text-controllable LiDAR point cloud generation framework, Text2LiDAR.
  • To advance the field of LiDAR point cloud generation, nuLiDARtext is constructed, comprising 34,149 pairs of text-LiDAR data.
CoGEM CVPR 2023
Co-Salient Object Detection with Uncertainty-aware Group Exchange-Masking, CVPR2023 (Code)
Yang Wu, Huihui Song, Bo Liu, Kaihua Zhang, Dong Liu
  • A robust CoSOD model learning mechanism, called group exchange-masking is proposed.
  • Our framework is different from the traditional CoSOD model learning frameworks that use groups of relevant images as training data.
  • We propose a dual-path feature extraction module to model the uncertainty and consensus feature.
OCDFormer ICME 2023
Object-Aware Calibrated Depth-Guided Transformer for RGB-D Co-Salient Object Detection, ICME2023 🏆 Best Student Paper Award
Yang Wu, Lingyan Liang, Yaqian Zhao, Kaihua Zhang
  • We design a DCM to split different depth layers and subdivide co-salient regions to produce calibrate depth maps, and set them as intermediate self-supervised signals to improve the generalization of the model.
  • A novel CMT is designed for feature extraction and fusion that can fully integrate the complementary knowledge among the RGB and depth features.
HrSSNM IEEE TMM 2022
Deep Object Co-segmentation and Co-saliency Detection via High-order Spatial-Semantic Network Modulation, IEEE TMM2022
Kaihua Zhang*, Yang Wu*(Equal Contribution), Mingliang Dong, Bo Liu, Dong Liu, Qingshan Liu
  • We further analyze and elucidate the similarities and differences between CSG and CSD tasks, and propose a joint processing framework.
  • We design an adaptive spatial modulation branch to learn a spatial representation for each cluster.
  • We design a high-order semantic modulation branch to transform the convolutional features with high-order statistics for object classification.

💼 Internships

  • 2025.08–Present: Long-term Intern, Momenta Momenta, Suzhou
  • 2025.04–2025.08: Star Program Intern, DiDi DiDi | Kargobot Kargobot, Shanghai
  • 2021.03–2023.06: HPC Cluster Intern, Sitonholy Sitonholy, Nanjing

🏆 Selected Honors and Awards

  • 2026.06: Outstanding Doctoral Fellowship (南理工优秀博士培养对象资助) (Top 1%), NJUST
  • 2023.07: Best Student Paper Award, IEEE International Conference on Multimedia and Expo (ICME)
  • 2023.06: Jiangsu Provincial Merit Student (省三好) (Top 1‰), JSDE | CYLC-JS
  • 2022.11: National Graduate Scholarship (国家奖学金) (Top 2%), NUIST
  • 2021.12: Second Prize in National Post-Graduate Mathematical Contest in Modeling, JSDE
  • 2019.05: Jiangsu Provincial Outstanding Student Leader (省优干) (Top 1‰), JSDE | CYL-JS
  • 2018.11: Second Prize in National College Student Robot Competition (Jiangsu), JSDE
  • 2018.07: First Prize in National Undergraduate Electronics Design Contest (Jiangsu), Texas Instruments

🤝 Academic Services

  • Conference Reviewer: CVPR, ECCV, ICCV, NeurIPS, ICLR, ICML
  • Journal Reviewer: TPAMI, IJCV, TIP, TNNLS, TMM, TCSVT, PR