Dahua Lin is an Associate Professor in the Department of Information Engineering at The Chinese University of Hong Kong and Director of the Institute of Interdisciplinary Artificial Intelligence at CUHK. He is also Co-founder and Chief Scientist of SenseTime, and concurrently serves as a Lead Scientist at the Shanghai Artificial Intelligence Laboratory and Chair of the IEEE Working Group on Evaluation Standards for Large-Scale Deep Learning Models. Professor Lin received his Ph.D. in Computer Science from the Massachusetts Institute of Technology in 2012. He has long been engaged in research on computer vision, deep learning, multimodal foundation models, and generative artificial intelligence. As of April 2026, he has published more than 300 papers in top-tier AI conferences and journals, with over 90,000 total citations. Among them, the most-cited paper for which he serves as the corresponding author has received approximately 7,000 citations, and his h-index has reached 130. From 2021 to 2025, he was selected for five consecutive years as one of Elsevier’s World’s Top 2% Scientists.
In terms of academic contributions, Professor Dahua Lin has continuously led frontier developments in visual intelligence and multimodal intelligence, driving the field from video understanding toward a new stage of large-scale multimodal models. In his early work, he proposed representative methods such as Temporal Segment Networks (TSN) and Spatio-Temporal Graph Convolutional Networks (ST-GCN), which broke through key bottlenecks in temporal modeling and action understanding. He also introduced Non-Parametric Instance Discrimination (NPID), establishing an important paradigm for modern contrastive learning and laying the foundation for subsequent methods such as MoCo and SimCLR. In recent years, he has led the open-source “Intern” model ecosystem, addressing core challenges in multimodal learning, including fine-grained alignment, high-resolution understanding, and the scarcity of post-training data. Through a series of innovative methods for data construction and model training, his work has significantly enhanced the ability to understand and reason over complex images, long videos, and three-dimensional spaces, thereby accelerating the broad adoption of key technologies in multimodal large models. In his latest research, he has further developed the SenseNova-SI framework and achieved major breakthroughs in spatial intelligence training, pushing multimodal intelligence from the digital world into the physical world.