MeshTransformer
Research code for CVPR 2021 paper “End-to-End Human Pose and Mesh Reconstruction with Transformers”
Image/Video Transformation
Image and video have become the language people use to communicate on the Internet. Multimedia content connects people and appeals to the young. This project aims at deep image and video transformation to generate high-quality…
OCR and Document Understanding
We have been developing SOTA technologies and industry-leading product solutions for following scenarios: (1) Universal OCR to detect and recognize any text in image/PDF; (2) Universal math OCR to detect and recognize any math expression…
Rich Media Communications
Modern work increasingly relies on online collaboration with real-time communications (RTC). Our research aims to provide real-time, intelligent, and immersive media experiences, with a long-term vision of advancing multimedia technologies in a manner that shapes…