Yuwen Tan

I'm a second-year PhD student in the Computer Science department at Boston University, where I am advised by Prof. Boqing Gong.
My current research interests focus on visual understanding in multimodal models and video generation.
Previously, I received my master's degree from Huazhong University of Science and Technology, where my research focused on continual learning and machine unlearning under the supervision of Prof. Xiang Xiang.

Email  /  CV  /  Scholar  /  Github

profile photo

Research

I'm interested in computer vision, deep learning, and generative AI. Most of my research is about vision language models and continual learning. Some papers are highlighted.

Examples of camera movements in cinematographic practice Natural Language Camera Movement Understanding
Yuwen Tan, Joey Huang, Jin Huang, Haoxiang Li, Boqing Gong

ECCV, 2026
project page / arXiv / code / dataset

We establish natural language camera movement understanding as a standalone task and introduce ACaM, an extensive benchmark and instruction-tuning framework spanning real and synthetic videos.

Vision LLMs Are Bad at Hierarchical Visual Understanding, and LLMs Are the Bottleneck
Yuwen Tan, Yuan Qing, Boqing Gong

CVPR, 2026
project page / arXiv

This paper reveals that many state-of-the-art large language models (LLMs) lack hierarchical knowledge about the visual world, failing to recognize even well-established biological taxonomies.

Lifting Data-Tracing Machine Unlearning to Knowledge-Tracing for Foundation Models
Yuwen Tan, Boqing Gong

TMLR, 2026
project page / arXiv

In this position paper, we propose to lift data-tracing machine unlearning to knowledge-tracing for foundation models (FMs). We support this position based on practical needs and insights from cognitive studies.

Semantically-Shifted Incremental Adapter-Tuning is A Continual ViTransformer
Yuwen Tan, Qinghao Zhou, Xiang Xiang, Ke Wang, Yuchuan Wu, Yongbin Li


CVPR, 2024
video / code

The proposed method eliminates the need for constructing an adapter pool and avoids retaining any image samples. Experimental results on five benchmarks demonstrate the effectiveness of our method which achieves the SOTA performance.

Coarse-to-fine few-shot class-incremental learning setting Coarse-To-Fine Incremental Few-Shot Learning
Xiang Xiang, Yuwen Tan, Qian Wan, Jing Ma, Alan L. Yuille, Gregory D. Hager

ECCV, 2022
paper / arXiv / code

We formulate coarse-to-fine few-shot class-incremental learning (C2FSCIL) and propose Knowe, a simple and effective strategy for learning fine-grained classes without forgetting coarse-grained knowledge.

Work Experience

Google (YouTube)
Student Researcher, May 2026 – Present
Supervisor: Dr. Yilin Wang
Pixocial, Human-First AI
Applied Research Intern, May 2025 – Aug. 2025
Supervisor: Dr. Haoxiang Li

Miscellanea

Academic Service

Conference Reviewer: CVPR 2026, ECCV 2026, ICLR 2026, NeurIPS 2025 & 2026
Journal Reviewer: TPAMI, IJCV, ACM Computing Surveys (CSUR)