Sessions
Explore the keynotes, panels, and conversations shaping the Pittsburgh AI Summit.
Keynote: From Robots to AI: Productivity, Wellbeing, and Life Beyond Work
When a new technology arrives, we ask what it does to jobs. That's the wrong first question. What a decade of research on robots shows is that technology shocks spill far outside the workplace — into health, households, marriage, and household spending.
I start with the last automation wave, because it's the one we can actually measure. In the U.S. and Germany, robot exposure made factories physically safer but left workers worse off in other ways, raising mental-health and substance-related risks (Gihleb, Giuntella, Stella & Wang, Labour Economics 2022). In China, workers didn't smoothly retrain into better jobs — they cut back on training and retired early, and households adjusted their consumption and savings in response (Giuntella, Lu & Wang, Economic Journal 2025). And in American communities where robots displaced men from well-paid manufacturing jobs, marriage rates fell, divorce and cohabitation rose, and fertility shifted out of marriage (Anelli, Giuntella & Stella, Journal of Human Resources). Automation didn't just move jobs. It reshaped families.
What about AI? So far, the data are surprisingly quiet. Two decades of German panel evidence show only modest effects of AI exposure on workers' health and satisfaction (Giuntella, König & Stella, Scientific Reports 2025), and related work finds no measurable effect of early predictive-AI diffusion on U.S. population health.
But generative AI is a different animal, and the observational data haven't caught up. So we ran an experiment with roughly 3,900 white-collar workers in the U.S. and Germany. Three early lessons: adoption is far below capability — simply telling workers they're permitted to use AI raised usage by 37 percentage points; the productivity gains are real but modest, showing up most clearly in the quality of what workers produce; and the same tool lands very differently across workers and across countries, with German workers showing more ambivalence even as they gain from it.
Lightning Talk: Design as Trust Infrastructure: Getting AI From Prototypes to Enterprise Adoption
Trust in an AI product is not established at the demo or the signed contract. It is built in a continuous chain that runs from the first positioning decision to the moment the person doing the work uses the system and wants to keep using it.
The talk follows that chain in order. How early brand and language choices set the expectations a product must later honor. How to make a model's reasoning legible to a skeptical expert user. How to design for the moment the system is wrong. How much automation a user will actually authorize, and how that threshold moves with use. Why the buyer and the daily user are rarely the same person, and what breaks when a product is designed only for the one who signs.
Keynote: Beyond Prediction: How AI Can Support High-Stakes Decisions Under Uncertainty
AI systems are often judged by how accurately they predict the future. But in high-stakes settings, prediction is only the beginning. A utility deciding whether to shut off power during extreme wildfire weather must balance uncertain ignition risk against disruptions to hospitals, businesses, and households. A public agency must decide not only what an algorithm recommends, but also when to trust it, when to seek human judgment, and when changing conditions make yesterday's recommendation unsafe.
This talk explores how to move from predictive AI to trustworthy decision support. Drawing on research and partnerships in electric-grid and wildfire resilience, I will show how AI can communicate what it does not know, connect uncertainty to operational consequences, and keep human expertise meaningfully in the loop. I will also discuss a counterintuitive finding from human–AI collaboration: adding a person does not automatically improve an algorithmic decision. Poorly designed collaboration can create "automation cliffs," where performance drops sharply when either the model or human is over-relied on.
The goal is not to eliminate uncertainty or automate every choice. It is to design systems that help people act responsibly despite uncertainty. Attendees will leave with a practical framework for evaluating AI in consequential settings: What decision is the model supporting? What uncertainty matters? What role should people play? And how will the system adapt when the world changes?
Fireside Chat: The Human Side of AI – Care, Adoption, and What We Protect
Too often, organizations treat care as an individual burden and AI as just another productivity tool. But in Pittsburgh's regulated industries—life sciences, health tech, and beyond—the real challenge isn't whether AI works, but whether people will embrace it.
In this conversation, Melike Konur introduces AI as Care Infrastructure, a framework that reframes AI not as a replacement for humans, but as infrastructure that expands our capacity for what matters most: creativity, relationships, and strategic thinking. Meanwhile, Dr. Shannon Gregg brings her PhD research on diffusion of innovations and change management to reveal why adoption fails when we ignore human behavior—and how to get it right in high-stakes, compliance-driven environments.
Together, they'll explore: What human capacities should AI protect? How do we design for adoption, not just implementation? And how can leaders ensure technology serves people, not the other way around?
Panel: The AI Evaluation Challenge: Bridging Theory and Practice
AI evaluation faces a dual challenge. Research systems map multiple criteria to overall scores (LLM reviewers assessing admissions candidates, clinicians diagnosing patients) yet these approaches rely on preference models that are often violated in the real world. Meanwhile, enterprise AI hits benchmark records but fails in deployment because accuracy alone doesn't reveal why systems succeed or fail.
Madeline addresses the theoretical gap: She presents a robust algorithm for learning evaluator preferences that works even when modeling assumptions break down. With minimal assumptions and theoretical guarantees, her approach learns any preference function without sacrificing performance when linearity holds, validated through synthetic simulations and real-world data.
Samhitha tackles the practical gap: From her work on Scale AI's Enterprise AI team, she knows that benchmark accuracy doesn't guarantee real-world readiness. She shares lessons from building the Ground Truth Verifier, which diagnoses failures by reasoning across source documents, agent outputs, and reference answers to pinpoint root causes, whether flawed model reasoning, poor prompts, retrieval failures, or incorrect ground truth.
Together, they show that reliable AI requires evaluation that's both theoretically grounded and practically diagnostic, building systems that are not just accurate, but transparent and trustworthy.
(Still being finalized <3)






