Further Notes on Our Recent Research on AI Delegation and Long-Horizon Reliability
Our recent paper, “LLMs Corrupt Your Documents When You Delegate”, has generated discussion about the reliability of AI systems in delegated workflows. We appreciate the interest in this work and want to clarify several important…
Senior Researcher – AI Computer Architecture
The future of AI is being built not just in software, but in the physics of light. Our Future AI Infrastructure (FAI) team at Microsoft Research Cambridge is pioneering new hardware and system technologies to…
Guiding the AI disruption to the Good Place
The true impact of AI does not lie in how well it takes tests or surfs the web, but in how effectively it teaches, coordinates, and operates in a web built for agents rather than…
New fine-tuning of language models: Match meaning, not tokens
Language models are usually trained to predict the next word, but that does not always lead to the best overall answers. We introduce energy-based fine-tuning, a new method that trains models to produce better full…
Introducing Interwhen: Steering reasoning agents with real-time verification
What if AI agents could check their work as they go? This verification method extracts verifiable properties from natural language and evaluates them using symbolic or model-based verifiers. Interwhen, a new open-source library, enables real-time…
Introducing GitHub Agentic Workflows: AI that runs your repo
What if your repo could run itself? GitHub Agentic Workflows bring AI agents directly into repository automation, enabling tasks to run end-to-end inside GitHub Actions. With built-in guardrails and Microsoft-hosted models on Azure, this system…