What working at Google AI, JPMorgan ML/IO, and DAT taught me about when people rely on AI

How I Approach AI Maturity

John Leo Szrejter · Perspective

A five-stage model for deciding when an AI capability has earned more autonomy

Most AI rollouts stall for the same reason: teams debate whether the model is good enough, when the real question is whether the people using it are ready. I started working on this problem as the sole UX researcher in MLIO, JPMorgan Chase's machine learning and AI product space, where I led research on the agent-assist pilot that became EVEE. Usage was low, and nobody could say whether the tool was the problem or the rollout was.

What came out of that work was a simple rule. Before an AI capability moves up a stage, measure three things at the stage it's on: do people know it's there, do they use it, and do they trust it the right amount? Each one gates the next.

I've since extended that rule into the five-stage model below. Each stage has explicit evidence gates, a measurement set, and governance requirements, grounded in human-automation research. The thresholds are working heuristics, calibrated per product and risk level.

Two things that change as AI climbs the ladder

Efficiency is an outcome, not an early gate. At Surface and Suggest, efficiency targets punish the exploration adoption depends on. In the EVEE pilot, handle time fell 30% only after AHT targets were lifted and specialists felt safe using the tool. At Automate and Agent, efficiency becomes the business case: time saved has to exceed time spent supervising, or the autonomy isn't earned.

Design systems make trust repeatable and autonomy efficient. Confidence indicators, explanations, override, and undo should be standard components, so every product earns trust the same way and the awareness, adoption, and trust measures stay comparable [5]. At Automate and Agent, the system becomes a constraint on what the AI creates: it builds from approved components and patterns instead of inventing its own, which keeps output on standard and cuts the review and rework that eat into time saved.

References

  1. A Model for Types and Levels of Human Interaction with Automation, Parasuraman, Sheridan & Wickens, IEEE Transactions on Systems, Man, and Cybernetics, 2000

  2. Trust in Automation: Designing for Appropriate Reliance, Lee & See, Human Factors, 2004

  3. Humans and Automation: Use, Misuse, Disuse, Abuse, Parasuraman & Riley, Human Factors, 1997

  4. Beyond Accuracy: The Role of Mental Models in Human-AI Team Performance, Bansal et al., AAAI HCOMP, 2019

  5. Guidelines for Human-AI Interaction, Amershi et al., CHI, 2019

  6. Productivity Assessment of Neural Code Completion, Ziegler et al., MAPS, 2022

  7. Levels of Autonomy for AI Agents, Feng, McDonald & Zhang, 2025

  8. Model Cards for Model Reporting, Mitchell et al., FAT*, 2019

  9. AI Risk Management Framework 1.0, NIST, 2023

  10. EU AI Act, Article 14: Human Oversight, European Union, 2024

  11. Foundations for an Empirically Determined Scale of Trust in Automated Systems, Jian, Bisantz & Drury, International Journal of Cognitive Ergonomics, 2000

  12. Human-Centered Artificial Intelligence: Reliable, Safe & Trustworthy, Shneiderman, International Journal of Human-Computer Interaction, 2020

See the framework's origin in the [EVEE case study]