Bo Gao

MSc. in AI Engineering @ Carnegie Mellon University

Multi-modality · Image & Video Generation · World Model

I received my Bachelor degree in the School of Intelligence Engineering at Sun Yat-sen University (SYSU), during which I also studied at The Chinese University of Hong Kong (CUHK) through a 2+2 program. At SYSU, I was advised by Prof. Xiaodan Liang. I am now pursuing my Master degree at Carnegie Mellon University (CMU), where I am advised by Prof. Andrea Zanette.
My current research interests mainly lie in Vision Language Model, image & video generation, and World Model.
📚 Passionate about Research 📷 Passionate about Photography 🎮 Passionate about Counter-Strike and Valorant 🎵 Passionate about Rap Music 🏀 Passionate about Basketball

📄 CV 💻 Github 🧠 Google Scholar
Email1: golbert2002 [at] 163 [dot] com
Email2: bogao [at] andrew [dot] edu

profile photo
📢 News
  • 📄 Submitted several papers about LLM to ICLR 2027 (2026)
  • 🏆 SLVR accepted as Poster at NeruIPS 2026 (2026)
  • 🏆 One paper accepted as Findings at EMNLP 2026 (2026)
  • 🏆 Bookagent accepted as Findings at ACL 2026 (2026)
  • 🏆 Free-Mask accepted as Oral at ACM Multimedia 2025 (2025)
  • 🎓 Received CMU Master offer and Dartmouth CS PhD offer for 2025 Fall (2025)
🏅 Honors & Awards
  • National Scholarship for 2021-2022
  • First Prize Scholarship for 2021-2022
  • Academic Innovation Scholarship for 2021-2022
  • Special Research Scholarship for 2022-2023
  • First Prize of National College Students Mathematical Competition 2022
  • Meritorious Winner Prize of Mathematical Contest in Modeling (MCM) 2023
  • Top 3% (24/929) in Kaggle: Image Matching Challenge 2024 – Hexathlon (CVPR'24 workshop) [link]

🧩 Academic Service
  • Reviewer: npj Digital Medicine (Nature Partner Journals), AAAI, NeruIPS, ICLR and CVPR.
📚 Publications
Conference
SLVR: Structured Latent Visual Reasoning via Human-like Reasoning Flows
Albert Gao* , Bing Xue † , Andrea Zanette †
The Fortieth Annual Conference on Neural Information Processing Systems (NeruIPS), 2026 Poster Accepted
paper / project

We propose the Linear Fitness Subspace (LFS) hypothesis: mutation-induced fitness variation encoded by PLMs is concentrated in a low-dimensional linear subspace of residue-level representation changes.

Linear Fitness Subspace in Protein Language Models Enables Sample-Efficient Directed Evolution
SiYuan Ma, Canran Xiao, Zikai Xiao, Albert Gao , Liang He, Xuan-Yu Wang, Shuying Cao, Xiaojun Jia †
Empirical Methods in Natural Language Processing (EMNLP), 2026 Poster Accepted
paper / project

We propose the Linear Fitness Subspace (LFS) hypothesis: mutation-induced fitness variation encoded by PLMs is concentrated in a low-dimensional linear subspace of residue-level representation changes.

BOOKAGENT: Orchestrating Safety-Aware Visual Narratives via Multi-Agent Cognitive Calibration
Bo Gao*, Chang Liu, Yuyang Miao, SiYuan Ma, Ser-Nam Lim †
The 64th Annual Meeting of the Association for Computational Linguistics (ACL), 2025 Poster Accepted
paper / project

We explore collaborative AI agents with instruction tuning to generate long books with improved realism and long-context consistency.

Free-Mask: A Novel Paradigm of Integration Between the Segmentation Diffusion Model and Image Editing to Improve Segmentation Ability
Bo Gao*, Jianhui Wang, Xinyuan Song, Yangfan He, Fangxu Xing, Tianyu Shi †
ACM Multimedia (ACM MM), 2025 Oral Accepted
paper / code

We propose Free-Mask, combining diffusion-based segmentation with image editing to create realistic datasets that better match open-world settings while generating accurate masks.

Exploring Warping-Guided Features via Adaptive Latent Diffusion Model for Virtual Try-on
Bo Gao*, Junchi Ren, Fei Shen †, Mengwan Wei, Zijun Huang
IEEE International Conference on Multimedia and Expo (ICME), 2024 Oral Accepted
paper / code

We present ALDM, a warping-guided adaptive latent diffusion model for more photorealistic virtual try-on results.

Journal
Research on Two-Way Detection of YOLO V5s+Deep Sort Road Vehicles Based on the Attention Mechanism
Bo Gao*, Ronghui Zhang †
Journal of Physics: Conference Series, 2022
paper

We propose a "double vertical line" algorithm by introducing attention mechanisms into YOLOv5 and combining them with Deep SORT-based tracking for robust vehicle detection.

🎓 Education
B.Eng. in Intelligence Engineering @ Sun Yat-sen University
Sep. 2021 - Jun. 2025
GPA: 3.9 / 4.0, Ranking: 3 / 110
MSc. in AI Engineering @ Carnegie Mellon University
Aug. 2025 - Now
🧪 Experience
Chinese Academy of Sciences
Research Intern
Beijing, China
May. 2024 - July. 2024
Mentor: Dr. Lianlei Shan
Image and Video Generation
Tsinghua University
Research Intern
Beijing, China
July. 2024 - Oct. 2024
Mentor: Dr. Jun Zhu
Image Editing
Everlyn
Research Intern
Hong Kong, China
Mar. 2025 - Now
Mentor: Dr. Ser-Nam Lim
Controllable Generation