About
I am Yifan Yang, a Senior Research SDE at Microsoft Research Asia (MSRA), based in Shanghai. I joined MSRA in 2021. My research focuses on visual content generation, multimodal foundation models, and general-purpose agentic systems, with a particular emphasis on bridging research innovation, open-source development, and real-world deployment.
I have published over 40 peer-reviewed papers at top-tier venues such as CVPR, ICCV, ECCV, NeurIPS, ICLR, ICML, and AAAI, and have filed more than ten international patents. My academic service includes serving as an Area Chair for NeurIPS, ICML, and ICLR, and as a Senior Program Committee member for AAAI. I received my bachelor’s degree from Tongji University and my master’s degree from Peking University, and I am currently pursuing a Ph.D. at Shanghai Jiao Tong University.
I have been deeply involved in the development of Microsoft’s Phi model family, including Phi-3 and Phi-4, with several of my techniques transferred into core Microsoft products, including Office and Azure. My work on LLM2CLIP enhances cross-modal representation learning by leveraging large language models; it has been integrated into the Phi-4-mini pretraining pipeline and received the AAAI 2026 Outstanding Paper Award and the WAIC 2026 Youth Outstanding Paper Nomination Award. My recent research also explores self-evolving agents, commercial visual content generation, video world models, text-to-audio-video generation, and structured multimodal reasoning.
Google Scholar (opens in new tab)
|
GitHub (opens in new tab)
If you are interested in internship opportunities or research collaborations, feel free to reach out at
đ§ yifanyang@microsoft.com
Selected Open-Source Projects
My open-source work spans self-evolving agents, multimodal representation learning, world models, and efficient visual generation. As of August, 2026, the projects highlighted below have collectively attracted 19,000 GitHub stars.
- SkillOpt: Executive Strategy for Self-Evolving Agent Skills (opens in new tab) â A text-space optimizer that treats reusable natural-language agent skills as trainable artifacts. It improves a frozen target model through trajectory-driven edits, bounded textual updates, validation gating, and slow/meta optimization, while adding zero model calls at deployment.
Open-source impact: As of August, 2026, SkillOpt has received 16,300 GitHub stars and 1,531 forks, ranked #3 among Python repositories of the day on Trendshift, shipped as a PyPI package, and was integrated by community projects including gbrain, gbrain-evals, and darwin-skill. It has also been featured by Microsoft Research and covered by VentureBeat, Synced (ćşĺ¨äšĺż), The Decoder, and other technology media.
Research result: SkillOpt is best or tied-best in all 52 evaluated modelâbenchmarkâharness settings across six benchmarks, seven target models, and direct-chat, Codex, and Claude Code execution modes.
Paper (opens in new tab) ¡ Project Page (opens in new tab) ¡ Microsoft Research Blog
- LLM2CLIP (opens in new tab) â Uses large language models as powerful textual teachers for CLIP, strengthening visual representations and textâimage alignment for long, dense, and multilingual descriptions. The released models have reached 300K cumulative Hugging Face downloads; the technology became Phi-4-mini’s visual-pretraining approach and received the AAAI 2026 Outstanding Paper Award and the WAIC 2026 Youth Outstanding Paper Nomination Award.
Paper (opens in new tab) ¡ Project Page (opens in new tab) ¡ Model Zoo (opens in new tab)
- Resource2Skill (opens in new tab) â Distills human-created multimodal resourcesâincluding tutorial videos, articles, code, and reference artifactsâinto reusable executable skills that agents can browse, compose, and run in real software environments. Across seven practical authoring domains, it improves the average overall score by +11.9 percentage points over no-skill agents and outperforms strong harness baselines in 26 of 28 main modelâdomain evaluation cells.
Paper (opens in new tab) ¡ Project Page (opens in new tab) ¡ Dataset and Skill Libraries (opens in new tab)
- SkillLens (opens in new tab) â A systematic framework for studying model-generated agent skills across the complete lifecycle from raw experience generation to skill extraction and skill consumption. It provides unified pipelines and reproducible evaluation across SWE-bench, ALFWorld, SpreadsheetBench, BFCL, and SEAL-0.
- World-R1 (opens in new tab) â Reinforces 3D constraints in text-to-video generation through camera-aware latent initialization, 3D-aware rewards, and periodic decoupled training, improving geometric consistency and camera control without changing the base video model architecture. ICML 2026.
Paper (opens in new tab) ¡ Project Page (opens in new tab) ¡ Dataset (opens in new tab)
- Latent Spatial Memory (opens in new tab) â Introduces a persistent 3D scene memory represented directly as latent tokens for video world models, avoiding repeated RGB rendering and re-encoding. The released Mirage system reports 10.57Ă faster generation and 55Ă lower 3D-cache memory.
- RAS: Region-Adaptive Sampling for Diffusion Transformers (opens in new tab) â A training-free sampling strategy that dynamically allocates computation to visually complex regions while reusing intermediate results in simpler regions, accelerating diffusion-transformer inference with minimal quality loss. CVPR 2026.
- BizGenEval: A Systematic Benchmark for Commercial Visual Content Generation (opens in new tab) â Evaluates image-generation models on real-world commercial design tasks involving dense text, precise layouts, multiple visual elements, and strict semantic constraints. It covers five document types and four capability dimensions through 20 evaluation tasks, 400 curated prompts, and 8,000 human-verified checklist questions.
Paper (opens in new tab) ¡ Project Page (opens in new tab) ¡ Dataset (opens in new tab) ¡ Leaderboard (opens in new tab)
- AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation (opens in new tab) â Provides task-oriented prompts and fine-grained evaluation for jointly generated audio and video, measuring audiovisual quality and synchronization, semantic alignment, task-specific capabilities, physical plausibility, and holistic quality. ICML 2026.
Paper (opens in new tab) ¡ Project Page (opens in new tab) ¡ Dataset (opens in new tab)
First-author and Corresponding-author Publications
* denotes co-first author; â denotes corresponding author. Publication titles link to arXiv or official publication pages.
- LLM2CLIP: Powerful Language Model Unlocks Richer Cross-Modality Representation (opens in new tab)
Weiquan Huang, Aoqi Wu, Yifan Yangâ , et al.
AAAI 2026 â Outstanding Paper Award đ ¡ WAIC 2026 Youth Outstanding Paper Nomination Award đ - OmniSch: A Multimodal PCB Schematic Benchmark for Structured Diagram Visual Reasoning (opens in new tab)
Taiting Lu, Kaiyuan Lin, Yuxin Tian, Yubo Wang, Muchuan Wang, Yifan Yangâ , et al.
ECCV 2026 - World-R1: Reinforcing 3D Constraints for Text-to-Video Generation (opens in new tab)
Weijie Wang, Xiaoxuan He, Youping Gu, Yifan Yangâ , et al.
ICML 2026 - AVGen-Bench: A Task-Driven Benchmark for Multi-Granular Evaluation of Text-to-Audio-Video Generation (opens in new tab)
Ziwei Zhou, Zeyuan Lai, Rui Wang, Yifan Yangâ , et al.
ICML 2026 - Video-in-the-loop: Span-grounded Long Video QA with Interleaved Reasoning (opens in new tab)
Chendong Wang, Donglin Bai, Yifan Yangâ , et al.
ICML 2026 - A Large Language Model Powered Integrated Circuit Footprint Geometry Understanding (opens in new tab)
Yida Wang, Taiting Lu, Runze Liu, Lanqing Yang, Yifan Yangâ , et al.
ICML 2026 - HiTVideo: Hierarchical Tokenizers for Enhancing Text-to-Video Generation with Autoregressive Large Language Models (opens in new tab)
Ziqin Zhou, Yifan Yangâ , et al.
AAAI 2026 - VidGuard-R1: AI-generated Video Detection and Explanation via Reasoning Multimodal Language Models and Reinforcement Learning (opens in new tab)
Kyoungjun Park, Yifan Yangâ , et al.
ICLR 2026 - Region-adaptive Sampling for Diffusion Transformers (opens in new tab)
Ziming Liu*, Yifan Yang*, et al.
CVPR 2026 - Diffusion²: Turning 3D Environments into Radio Frequency Heatmaps (opens in new tab)
Kyoungjun Park, Yifan Yangâ , et al.
CVPR 2026 Findings - Zoomer: Adaptive Image Focus Optimization for Black-box Multimodal Large Language Models (opens in new tab)
Jiaxu Qian, Chendong Wang, Yifan Yangâ , et al.
Transactions on Machine Learning Research (TMLR), 2025 - VoLUT: Efficient Volumetric Streaming Enhanced by LUT-based Super-resolution (opens in new tab)
Chendong Wang, Anlan Zhang, Yifan Yangâ , et al.
MLSys 2025 - LoRaSC: Expressive and Generalizable Low-rank Adaptation for Large Models via Slow Cascaded Learning (opens in new tab)
Siwei Li, Yifan Yangâ *, et al.
EMNLP 2024 Findings - VIGOR: Reviving Cloud Gaming Sessions (opens in new tab)
Zhaoyuan He, Yifan Yang*, et al.
ACM CoNEXT 2024 - Nerve: Real-time Neural Video Recovery and Enhancement on Mobile Devices (opens in new tab)
Zhaoyuan He, Yifan Yang*, et al.
Proceedings of the ACM on Networking (CoNEXT), 2024 - Attentive Mask CLIP (opens in new tab)
Yifan Yang*, et al.
ICCV 2023 - ImageBrush: Learning Visual In-Context Instructions for Exemplar-Based Image Manipulation (opens in new tab)
Yasheng Sun*, Yifan Yang*, et al.
NeurIPS 2023 - Directional Self-supervised Learning for Heavy Image Augmentations (opens in new tab)
Yalong Bai*, Yifan Yang*, et al.
CVPR 2022
Technical Reports and Major Preprints
- SkillOpt: Executive Strategy for Self-Evolving Agent Skills (opens in new tab)
Yifan Yangâ , Ziyang Gong, Weiquan Huang, Qihao Yang, Ziwei Zhou, et al.
arXiv preprint arXiv:2605.23904, 2026 â first and corresponding author - SciDiagramEdit: Learning to Edit Scientific Diagrams from Paper Revisions (opens in new tab)
Yasheng Sun, Zezi Zeng, Yifan Yangâ , Chong Luo, Wenyi Wang, Ziwei Liu, JĂźrgen Schmidhuber
arXiv preprint arXiv:2607.15272, 2026 - RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources (opens in new tab)
Yijia Fan, Zonglin Di, Zimo Wen, Yifan Yangâ , et al.
arXiv preprint arXiv:2606.29538, 2026 - Latent Spatial Memory for Video World Models (opens in new tab)
Weijie Wang, Haoyu Zhao, Yifan Yangâ , Feng Chen, Zeyu Zhang, et al.
arXiv preprint arXiv:2606.09828, 2026 - From Raw Experience to Skill Consumption: A Systematic Study of Model-Generated Agent Skills (opens in new tab)
Zisu Huang, Jingwen Xu, Yifan Yangâ , Ziyang Gong, Qihao Yang, et al.
arXiv preprint arXiv:2605.23899, 2026 - ReasonGen-R1: Chain-of-Thought for Autoregressive Image Generation Models through Supervised Fine-tuning and Reinforcement Learning (opens in new tab)
Yu Zhang, Yunqi Li, Yifan Yangâ , et al.
arXiv:2505.24875, under review at ECCV 2026 - MM-WebAgent: A Hierarchical Multimodal Web Agent for Webpage Generation (opens in new tab)
Yan Li, Zezi Zeng, Yifan Yangâ , et al.
arXiv preprint arXiv:2604.15309, 2026 - BizGenEval: A Systematic Benchmark for Commercial Visual Content Generation (opens in new tab)
Yan Li, Zezi Zeng, Ziwei Zhou, Xin Gao, Muzhao Tian, Yifan Yangâ , et al.
arXiv preprint arXiv:2603.25732, 2026 - Phi-4-mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs (opens in new tab)
Abdelrahman Abouelenin, Atabak Ashfaq, et al., Yifan Yang
arXiv Technical Report, 2025 - Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone (opens in new tab)
Marah Abdin, Sam Ade Jacobs, et al., Yifan Yang
arXiv Technical Report, 2024
Collaborative Preprints
- Ideas Have Genomes: Benchmarking Scientific Lineage Reasoning and Lineage-Grounded Idea Generation (opens in new tab)
Yifan Zhou, Qihao Yang, Yan Li, âŚ, Yifan Yang, et al.
arXiv preprint arXiv:2607.08758, 2026 - Covering Human Action Space for Computer Use: Data Synthesis and Benchmark (opens in new tab)
Miaosen Zhang, Xiaohan Zhao, Zhihong Tan, Zhou Huoshen, Yijia Fan, Yifan Yang, et al.
arXiv preprint arXiv:2605.12501, 2026 - EmbodiSkill: Skill-Aware Reflection for Self-Evolving Embodied Agents (opens in new tab)
Ruofei Ju, Xinrui Wang, Xin Ding, Yifan Yang, et al.
arXiv preprint arXiv:2605.10332, 2026 - MemCompiler: Compile, Don’t InjectâState-Conditioned Memory for Embodied Agents (opens in new tab)
Xin Ding, Xinrui Wang, Yifan Yang, Hao Wu, Shiqi Jiang, et al.
arXiv preprint arXiv:2605.07594, 2026
Collaborative Publications
- A Disease-Centric Vision-Language Foundation Model for Precision Oncology in Kidney Cancer (opens in new tab)
Yuhui Tao, Zhongwei Zhao, Zilong Wang, âŚ, Yifan Yang, et al.
Nature Communications 17, 7313 (2026) - WeaveBench: A Long-Horizon, Real-World Benchmark for Computer-Use Agents with Hybrid Interfaces (opens in new tab)
Wanli Li, Bowen Zhou, Yunyao Yu, Zhou Xu, Yifan Yang, Dongsheng Li, Caihua Shan
EMNLP 2026 - OmniLayout: A Schematic-Coupled Multimodal Benchmark for Constraint-Aware Geometric Reasoning in PCB Layout (opens in new tab)
Taiting Lu, Kaiyuan Lin, Mingjia Wang, âŚ, Yifan Yang, et al.
EMNLP 2026 - RE-TRAC: REcursive TRAjectory Compression for Deep Search Agents (opens in new tab)
Jialiang Zhu, Gongrui Zhang, Xiaolong Ma, Lin Xu, Miaosen Zhang, Ruiqi Yang, Song Wang, Kai Qiu, Zhirong Wu, Qi Dai, Ruichun Ma, Bei Liu, Yifan Yang, et al.
ICML 2026 - AdaNav: Adaptive Reasoning with Uncertainty for Vision-Language Navigation (opens in new tab)
Xin Ding, Jianyu Wei, Yifan Yang, et al.
ICML 2026 - AMID: Model-Agnostic Dataset Distillation by Adversarial Mutual Information Minimization (opens in new tab)
Aoqi Wu, Junming Liu, Yuwei Zhang, Weiquan Huang, Liang Hu, Yifan Yang, et al.
Proceedings of the ACM Web Conference 2026 - A Comprehensive Ecosystem for Open-Domain Customized Video Generation (opens in new tab)
Jingxu Zhang, Yuqian Hong, Daneul Kim, Kai Qiu, Qi Dai, Jianmin Bao, Yifan Yang, et al.
ICASSP 2026 - Unified Medical Image Pre-training in Language-Guided Common Semantic Space (opens in new tab)
Xiaoxuan He, Yifan Yang, et al.
ECCV 2024 - StreamMind: Unlocking Full Frame-rate Streaming Video Dialogue through Event-gated Cognition (opens in new tab)
Xin Ding, Hao Wu, Yifan Yang, et al.
ICCV 2025 - Efficient and Adaptive Diffusion Model Inference through Lookup Tables on Mobile Devices (opens in new tab)
Qipeng Wang, Shiqi Jiang, Yifan Yang, et al.
IEEE Transactions on Mobile Computing, 2025 - Online Video Quality Enhancement with Spatial-Temporal Look-up Tables (opens in new tab)
Zefan Qu, Xinyang Jiang, Yifan Yang, et al.
ECCV 2024 - ProLongVid: A Simple but Strong Baseline for Long-context Video Instruction Tuning (opens in new tab)
Rui Wang, Bohao Li, Yifan Yang, et al.
EMNLP 2025 - MageBench: Bridging Large Multimodal Models to Agents (opens in new tab)
Miaosen Zhang, Qi Dai, Yifan Yang, et al.
WACV 2025 - Reducio! Generating 1K Video within 16 Seconds Using Extremely Compressed Motion Latents (opens in new tab)
Rui Tian, Qi Dai, Yifan Yang, et al.
ICCV 2025 - Expand Heterogeneous Learning Systems with Selective Multi-Source Knowledge Fusion (opens in new tab)
Gaole Dai, Huatao Xu, Yifan Yang, Rui Tan, Mo Li
AAAI 2026 - Empowering Agentic Video Analytics Systems with Video Language Models (opens in new tab)
Yuxuan Yan, Shiqi Jiang, Ting Cao, Yifan Yang, et al.
USENIX NSDI 2025 - DreamDistribution: Learning Prompt Distribution for Diverse In-distribution Generation (opens in new tab)
Brian Nlong Zhao, Yifan Yang, et al.
ICLR 2025 - Understanding and Improving Training-free Loss-based Diffusion Guidance (opens in new tab)
Yifei Shen, Xinyang Jiang, Yifan Yang, et al.
NeurIPS 2024 - Online Video Super-resolution with Convolutional Kernel Bypass Grafts (opens in new tab)
Jun Xiao, Xinyang Jiang, Yifan Yang, et al.
IEEE Transactions on Multimedia, 2023