Song-Chun Zhu (Chinese: 朱松纯; born June 1968 in Ezhou, Hubei, China) is a Chinese computer scientist and applied mathematician known for his work in computer vision, cognitive artificial intelligence, and robotics. He received a Bachelor of Science from the University of Science and Technology of China and completed his Master of Science and PhD in 1996 at Harvard University under advisor David Mumford, with a thesis titled "Statistical and Computational Theories for Image Segmentation, Texture Modeling and Object Recognition".
Zhu currently works at Peking University and was previously a professor in the Departments of Statistics and Computer Science at the University of California, Los Angeles (UCLA), where he also served as Director of the UCLA Center for Vision, Cognition, Learning and Autonomy (VCLA).[1,2] In 2005, Zhu founded the Lotus Hill Institute, an independent non-profit organization to promote international collaboration within the fields of computer vision and pattern recognition.[3] He has published extensively, lectured globally on artificial intelligence, and received accolades including the David Marr Prize, the Helmholtz Test-of-Time Award, and being named an IEEE Fellow in 2011 for "contributions to statistical modeling, learning and inference in computer vision."[12,4]
Zhu has two daughters, Stephanie and Yi, the latter of whom is a competitive figure skater.[5,6]
Early life and education
Born and raised in Ezhou, China, Zhu found inspiration when he was young in the development of computers playing chess, sparking his interest in artificial intelligence. In 1991, Zhu earned his B.S. in Computer Science from the University of Science and Technology of China at Hefei. During his undergraduate years, Zhu, finding the computational theory of vision by the late MIT neuroscientist David Marr deeply influential, aspired to pursue a general unified theory of vision and AI.[7]
In 1992, Zhu continued his study of computer vision at the Harvard Graduate School of Arts and Sciences. At Harvard, Zhu studied under the supervision of American mathematician David Mumford and gained an introduction to "probably approximately correct" (PAC) learning under the instruction of Leslie Valiant. Zhu concluded his studies at Harvard in 1996 with a Ph.D. in Computer Science and followed Mumford to the Division of Applied Mathematics at Brown University as a postdoctoral fellow.[7]
Career
Following his postdoctoral fellowship, Zhu lectured briefly in Stanford University's Computer Science Department. In 1998, he joined Ohio State University as an assistant professor in the Departments of Computer Science and Cognitive Science. In 2002, Zhu joined the University of California, Los Angeles in the Departments of Computer Science and Statistics as an associate professor, rising to the rank of full professor in 2006. At UCLA, Zhu established the Center for Vision, Cognition, Learning and Autonomy. His chief research interest has resided in pursuing a unified statistical and computational framework for vision and intelligence, which includes the Spatial, Temporal, and Causal And-Or graph (STC-AOG) as a unified representation and numerous Monte Carlo methods for inference and learning.[8,9]
In 2005, Zhu established an independent non-profit organization in his hometown of Ezhou, the Lotus Hill Institute (LHI). LHI has been involved with collecting large-scale datasets of images and annotating objects, scenes, and activities, having received contributions from many renowned scholars, including Harry Shum. The institute also features a full-time annotation team for parsing image structures, having amassed over 500,000 images to date. Since establishing LHI, Zhu has organized numerous workshops and conferences, along with serving as the general chair for both the 2012 Conference on Computer Vision and Pattern Recognition (CVPR) in Providence, Rhode Island, where he presented Ulf Grenander with a Pioneer Medal, and the 2019 CVPR held in Long Beach, California.[10] In July 2017, Zhu founded DMAI in Los Angeles as an AI startup engaged in developing a unified cognitive AI platform.[11]
In September 2020, Zhu returned to China to join Peking University to lead its Institute for Artificial Intelligence, joining another Chinese AI expert in the US and long-time acquaintance, Microsoft's former head of artificial intelligence and research, Harry Shum. Shum was also appointed by Peking University in August to chair the academic committee of the Institute of Artificial Intelligence. Zhu is working on setting up a new and separate AI research institute, the Beijing Institute for General Artificial Intelligence (BIGAI). According to the introduction, based on the "small data for big task" paradigm, BIGAI focuses on advanced AI technology, multi-disciplinary integration, and international academic exchange to nurture the new generation of young AI talents.[12] The institute is expected to gather professional researchers, scholars, and experts to put Zhu's theoretical framework of artificial intelligence into practice, jointly promoting Chinese original AI technologies and building a new generation of general AI platforms.
Research and work
Zhu has published over three hundred articles in peer-reviewed journals and proceedings across four phases.
Statistical models in Marr's framework
In the early 1990s, Zhu, along with collaborators in the pattern theory group, developed advanced statistical models for computer vision. Focusing on creating a unifying statistical framework for the early vision representations presented in David Marr's posthumously published work titled Vision, they first formulated textures in a new Markov random field model called FRAME, utilizing a minimax entropy principle to introduce discoveries in neuroscience and psychophysics to Gibbs distributions in statistical physics.[13] They subsequently proved the equivalence between the FRAME model and the micro-canonical ensemble, which they designated the Julesz ensemble. This work earned the Marr Prize honorary nomination during the International Conference on Computer Vision (ICCV) in 1999.[15,14]
During the 1990s, Zhu also developed two new classes of nonlinear partial differential equations (PDEs). One class, designed for image segmentation, is called region competition, and this work connecting PDEs to statistical image models received the Helmholtz Test of Time Award at ICCV 2013.[16] The other class, called GRADE (Gibbs Reaction and Diffusion Equations), was published in 1997 and employs a Langevin dynamics approach for inference and learning Stochastic gradient descent (SGD).[17]
In the early 2000s, Zhu formulated textons using generative models with sparse coding theory and integrated both texture and texton models to represent primal sketch.[19,18] Alongside Ying Nian Wu, Zhu advanced the study of perceptual transitions between regimes of models in information scaling and proposed a perceptual scale space theory to extend the image scale space.[20]
Stochastic and-or graph grammar paradigm
From 1999 until 2002, with his Ph.D. student Zhuowen Tu, Zhu developed a data-driven Markov chain Monte Carlo (DDMCMC) paradigm to traverse the entire state-space by extending the jump-diffusion work of Grenander-Miller. With another Ph.D. student, Adrian Barbu, he generalized the cluster sampling algorithm (Swendsen-Wang) in physics from Ising/Potts models to arbitrary probabilities. This advancement in the field made the split-merge operators reversible for the first time in the literature and achieved 100-fold speedups over Gibbs sampler and jump-diffusion. This accomplishment led to the work on image parsing that won the Marr Prize in ICCV 2003.[23,21,22]
In 2004, Zhu moved to high level vision by studying stochastic grammar, a method dating back to the syntactic pattern recognition approach advocated by King-Sun Fu in the 1970s. Zhu developed grammatical models for a few key vision problems, such as face modeling, face aging, clothes, object detection, rectangular structure parsing, and the sort. He wrote a monograph with Mumford in 2006 titled A Stochastic Grammar of Images.[24] In 2007, Zhu and co-authors received a Marr Prize nomination. The following year, Zhu received the J.K. Aggarwal Prize from the International Association of Pattern Recognition for "contributions to a unified foundation for visual pattern conceptualization, modeling, learning, and inference." Zhu has extended the and-or graph models to the spatial, temporal, and causal and-or graph (STC-AOG) to express the compositional structures as a unified representation for objects, scenes, actions, events, and causal effects in physical and social scene understanding problems.
Cognition and visual commonsense
Since 2010, Zhu has collaborated with scholars from cognitive science, AI, robotics, and language to explore what he calls the "Dark Matter of AI"—the 95% of the intelligent processing not directly detectable in sensory input. Together they have augmented the image parsing and scene understanding problem by cognitive modeling and reasoning about functionality (functions of objects and scenes, the use of tools), intuitive physics (supporting relations, materials, stability, and risk), intention and attention (what people know, think, and intend to do in social scene), causality (the causal effects of actions to change object fluents), and utility (the common values driving human activities in video).[26] The results are disseminated through a series of workshops.[29,27]
There are numerous other topics Zhu has explored during this period, including formulating AI concepts such as tools, container, liquids; integrating three-dimensional scene parsing and reconstruction from single images by reasoning functionality, physical stability, situated dialogues by joint video and text parsing; developing communicative learning; and mapping the energy landscape of non-convex learning problems.[30,28]
Small-data paradigm for general AI
In a widely circulated public article written in Chinese in 2017, Zhu referred to popular data-driven deep learning research as a "big data for small task" paradigm that trains a neural network for each specific task with massive annotated data, resulting in uninterpretable models and narrow AI. Instead, Zhu advocated for a "small data for big task" paradigm to achieve general AI.[31]
At the 2023 meeting of the Chinese People's Political Consultative Conference's National Committee, Zhu stated that, in the wake of ChatGPT's release, China should make artificial general intelligence a strategic goal, analogous to the pursuit of nuclear, missile, and satellite technology by the Two Bombs, One Satellite project of the 1960s.[32] In February 2024, the Beijing Institute for General Artificial Intelligence (BIGAI), operating under Zhu's leadership, unveiled what they referred to as the world’s first artificial intelligence (AI) child named "Tong Tong," who possesses her own emotions and intellect and is capable of assigning tasks to herself independently, demonstrating a level of autonomy previously unseen in virtual entities.[33]
Books
S.C. Zhu co-authored "A Stochastic Grammar of Images" with D.B. Mumford, which was published by now Publishers Inc. in 2007. In 2019, Zhu co-authored "Monte Carlo Methods" with A. Barbu, published by Springer Nature, and authored "AI: The Era of Big Integration – Unifying Disciplines within Artificial Intelligence," published by DMAI, Inc. With Y.N. Wu, Zhu co-authored "Computer Vision: Statistical Models for Marr's Paradigm," published by Springer Nature in 2023, as well as "Concepts and Representations in Vision and Cognition," a draft taught for over 10 years and prepared for Springer for 2020.
Papers
In 1996, S.C. Zhu and A. Yuille published "Region competition: unifying snakes, region growing, and Bayes/MDL for multiband image segmentation" in IEEE Transactions on Pattern Analysis and Machine Intelligence. S.C. Zhu and D. Mumford published "Prior learning and Gibbs reaction-diffusion" in IEEE Transactions on Pattern Analysis and Machine Intelligence in 1997. In 1998, S.C. Zhu, Y. Wu, and D. Mumford published "FRAME: filters, random fields, and minimax entropy towards a unified theory for texture modeling" in the International Journal of Computer Vision. This was followed in 2000 by "Equivalence of Julesz Ensemble and FRAME models" by Y. N. Wu, S. C. Zhu, and X. W. Liu, also published in the International Journal of Computer Vision.
In 2002, Z. Tu and S.-C. Zhu published "Image Segmentation by Data Driven Markov Chain Monte Carlo" in IEEE Transactions on Pattern Analysis and Machine Intelligence. Z. Tu, X. Chen, A. Yuille, and S.-C. Zhu published "Image parsing: unifying segmentation, detection, and recognition" in 2003. In 2005, A. Barbu and S.-C. Zhu published "Generalizing Swendsen-Wang to Sampling Arbitrary Posterior Probabilities" in IEEE Transactions on Pattern Analysis and Machine Intelligence, and S.C. Zhu, C. Guo, Y. Wang, and Z. Xu published "What are Textons?" in the International Journal of Computer Vision. In 2006, S.C. Zhu and D. Mumford published "A Stochastic Grammar of Images" in Foundations and Trends in Computer Graphics and Vision. C. Guo, S.-C. Zhu, and Y. Wu published "Primal sketch: Integrating Texture and Structure" in Computer Vision and Image Understanding in 2007. In 2008, Y.N. Wu, C.E. Guo, and S.C. Zhu published "From Information Scaling of Natural Images to Regimes of Statistical Models" in the Quarterly of Applied Mathematics.
In 2015, B. Zheng, Y. Zhao, J. Yu, K. Ikeuchi, and S.C. Zhu published "Scene Understanding by Reasoning Stability and Safety" in the International Journal of Computer Vision, and Y. Zhu, Y.B. Zhao, and S.C. Zhu authored "Understanding Tools: Task-Oriented Object Modeling, Learning and Recognition". In 2016, A. Fire and S.C. Zhu published "Learning Perceptual Causality from Video" in ACM Transactions on Intelligent Systems and Technology, while Y.X. Zhu, C. Jiang, Y. Zhao, D. Terzopoulos, and S.C. Zhu published "Inferring Forces and Learning Human Utilities from Video". In 2018, D. Xie, T. Shu, S. Todorovic, and S.C. Zhu published "Learning and Inferring “Dark Matter” and Predicting Human Intents and Trajectories in Videos" in IEEE Transactions on Pattern Analysis and Machine Intelligence. In 2020, Y. Zhu et al. published "Dark, Beyond Deep: A Paradigm Shift to Cognitive AI with Human-like Commonsense" in Engineering.
External links
External links related to Song-Chun Zhu include his academic page at UCLA. He was also featured in a September 16, 2025 Guardian article titled "‘I have to do it’: Why one of the world’s most brilliant AI scientists left the US for China."
