Less hype and more hardware: SenseTime banks on multimodal AI to regain its edge
SenseTime may have slipped from the spotlight as one of China’s “AI dragons”, but its tilt at multimodal, real-world AI hints at a comeback

Chinese artificial intelligence pioneer SenseTime is betting that its roots in computer vision will help it lead the next phase of AI, as the industry shifts towards multimodal systems and embodied intelligence in the physical world, according to co-founder and chief scientist Lin Dahua.
In an interview with the Post on Wednesday, Lin said the company’s long-standing expertise in vision-based AI put it in a strong position to become a leader in embodied intelligence, robotics and AI agents operating in real-world environments, at a time when there is growing debate about the limits of large language models (LLMs).
“Our strategic approach is somewhat similar to Google’s in the United States, which primarily focuses on multimodal AI including the latest Nano Banana Pro. They also start with vision capabilities as the core, then add language abilities to create real multimodal systems,” said Lin, who is also an associate professor of information engineering at the Chinese University of Hong Kong.