Machine Learning and Friends Lunch: Amit Sharma, Building Self-Improving AI Systems
Content
Speaker
Amit Sharma (Prescera Labs)
Abstract
As large language models are applied to complex domains, the frontier is shifting from models trained once to systems that improve themselves. This talk decomposes self-improvement into three axes: environment generation, RL training, and harness evolution. For the first two, I will present executable counterfactuals, a framework that operationalizes counterfactual reasoning as code, enabling scalable generation of RL environments with controllable difficulty and verifiable answers. RL training on these environments induces the core behaviors of counterfactual reasoning — abduction, intervention, and prediction — and, unlike supervised fine-tuning, generalizes out of domain to new code structures and to math word problems. On the harness side, I present interwhen, which compiles natural-language policies into formal verifiers that certify each step of an agent's reasoning at runtime, raising the pass⁴ reliability of a 30B open-weight model on τ²-bench Telecom from 32% to 87%. These axes compose into a single self-improvement loop, and present a central research question on how to perform RL jointly with the harness components and self-evolve both the model and its harness.
Speaker Bio
Amit Sharma is the Chief Scientist and co-founder at Prescera Labs, focused on cybersecurity and self-improving AI systems. He spent the past decade at Microsoft Research, where he worked on AI reasoning and causal inference. His work has led to foundational contributions in causal reasoning, with applications in enhancing AI systems' generalization, explainability and reasoning abilities. He developed the DiCE algorithm for counterfactual explanation and refutation methods for evaluating causal estimates, which are widely adopted in both academia and industry. The related open-source libraries, DoWhy for causal inference and DiCE for counterfactual explanations, have been downloaded by millions of users and are used to impact government policy, health outcomes, and business decisions globally. Amit is also the co-founder of PyWhy, an open-source ecosystem involving Carnegie Mellon University, Microsoft, Amazon and others for advancing scalable causal ML tools. His work has received many awards including the 2025 Outstanding Paper (Top-3 finalist) award at TMLR, Outstanding Paper award at ICLR 2026 workshop on logical reasoning, 2023 NASSCOM AI GameChangers award, Best Paper Award at ACM CHI 2021 conference, and the 2012 Yahoo! Key Scientific Challenges Award.