Research Intern for Visual Computing Group – Microsoft Research 

Research Intern, Visual Computing Group

Position: Research Intern
Job Type: Full-time Intern
Location: Beijing / Shanghai
Number of Openings: 1-2

Group Introduction

The Visual Computing Group at Microsoft Research Asia conducts cutting-edge research on multimodal foundation models, visual generation, and AI agents for creative and knowledge-work scenarios. Our mission is to advance the next generation of AI systems that can understand, reason, create, and collaborate with users through multimodal interactions.

Our current research spans:

  • Multimodal reasoning and generation
  • Agentic content creation and editing
  • Visual generation and understanding
  • Human-AI collaboration and AI-native productivity experiences
  • Post-training and reinforcement learning for multimodal models
  • Interactive design and controllable generation
  • Data generation, evaluation, and model alignment

Our technologies are actively transferred to Microsoft Office, M365 Copilot, and future AI-native productivity experiences.

Researchers and interns in the group have opportunities to publish at top-tier conferences such as CVPR, ICCV, ECCV, NeurIPS, ICML, and ICLR, while also contributing to real-world product innovation.

Responsibilities

  • Conduct cutting-edge research in multimodal AI, visual understanding and generation, and AI agents.
  • Design, implement, and improve foundation models and agent systems to solve real-world problems.
  • Build datasets, evaluation benchmarks, and experimentation pipelines for model development.
  • Explore post-training, reinforcement learning, reasoning, and alignment techniques for multimodal models.
  • Collaborate closely with research and product teams to translate research innovations into impactful user experiences.
  • Publish research findings at leading academic venues and contribute to the broader research community.

Qualifications

Required

  • Ability to devote most of your effort to the internship during the internship period.
  • Experience in machine learning, deep learning, computer vision, natural language processing, or related fields.
  • Familiarity with one or more of the following areas:

1.Large Language Models (LLMs)

2.Vision-Language Models (VLMs)

3.Diffusion and generative models

4.Reinforcement learning and post-training

5.AI agents and tool use

6.Multimodal reasoning and generation

  • Strong problem-solving and research capabilities.
  • Ability to read and communicate effectively in English.
  • Strong communication and collaboration skills.
  • Written approval from academic advisor.

Preferred

  • Research experience demonstrated through publications, open-source contributions, competition achievements, or impactful projects.
  • Experience with large-scale experimentation, evaluation, data curation, synthetic data generation, or model alignment.
  • Experience developing agentic systems, multimodal applications, or productivity-related AI systems.
  • Experience with distributed training, model deployment, or systems optimization.

Internship Duration

  • Must obtain approval from academic advisor.
  • Minimum internship duration: 3 months.
  • Longer internship periods are strongly encouraged.