AION
Opinion / analysisAgents & Tool Use · Safety & Alignment · Applications1 source · Oct 6, 2026

What AI gets wrong and what failure teaches us

Jennifer Neville is a partner research manager at Microsoft who’s built a career around understanding and advancing AI for real-world use, and much like the human-AI interactions she’s been studying, her early-career path was multiturn: math, then physics; cognitive science, then work; and finally computer science—despite her best efforts to avoid the field.

Key points

  • In this conversation with Principal Applied Scientist Chad Atalla, she explores the role evaluation plays in pushing the performance boundaries of today’s AI systems to meet user needs and the “surprising failures” that emerge when models are tested beyond traditional benchmarks.
  • Neville also shares practical guidance for working with current AI systems and discusses why looking closely at data matters when results defy expectations, and what decades of AI progress have taught her about predicting what comes next.
  • From an unexpected career trajectory to the frontier of AI interaction and learning, this episode asks a larger question: what can we learn when the path—whether human or artificial—doesn’t unfold the way we expect?
  • STANDARD INTRODUCTION: This is the Microsoft Research Podcast, where Microsoft researchers—driving advancement through fundamental science and technology research—explore the who, how, and what’s next in computing and AI.

Sources (1)

  • [1]What AI gets wrong and what failure teaches us
    Microsoft Research Blog · Oct 6, 04:19 PM
    Jennifer Neville is a partner research manager at Microsoft who’s built a career around understanding and advancing AI for real-world use, and much like the human-AI interactions she’s been studying, her early-career path was multiturn: math, then physics; cognitive science, then work; and finally computer science—despite her best efforts to avoid the field.
    In this conversation with Principal Applied Scientist Chad Atalla, she explores the role evaluation plays in pushing the performance boundaries of today’s AI systems to meet user needs and the “surprising failures” that emerge when models are tested beyond traditional benchmarks.

Extractive summary: sentences quoted from the sources.