McKinsey: 80% of leaders measure AI ROI with pre-deployment baselines
Serge Bulaev
The text suggests that measuring AI impact in companies still feels experimental, but about 80% of leaders who set pre-deployment baselines report meaningful improvements after using AI. Experts recommend recording key workflow metrics before and after AI is added to see clear changes in speed, quality, and reliability. Small test sets may catch early problems, but larger tests and ongoing monitoring appear needed to track real results over time. Continuous measurement is important because workflow changes or model updates might reduce early gains.

Recent industry insights show how to effectively measure AI ROI: many leaders use pre-deployment baselines to see real workflow gains. While many teams struggle with experimental metrics, a clear playbook has emerged. True impact is revealed not by single AI calls, but by tracking improvements across the entire workflow - making it faster, more accurate, and more cost-effective. This pragmatic approach relies on establishing a quantified baseline, consistently tracking operational metrics, and re-evaluating with every model or prompt update.
Pick a Workflow and Capture a Baseline
To measure AI ROI effectively, first select a contained workflow and document its performance before any AI integration. Record key metrics like cycle time, error rates, throughput, and costs. This pre-deployment baseline provides a quantitative foundation for proving the value of subsequent AI-driven improvements.
Experts recommend mapping a single, contained process from start to finish. Before introducing AI, practitioners must record its current cycle time, error rate, throughput, and cost per task. Without this "single workflow, defined baseline," proving GenAI ROI is nearly impossible. However, McKinsey's 2025 State of AI found that more than 80% of respondents are not seeing tangible enterprise-level EBIT impact from gen AI, while the 2026 update says 37% report positive EBIT impact, highlighting the importance of disciplined measurement approaches.
Track Speed, Quality, and Reliability After Launch
After deploying AI, monitor the same baseline metrics to ensure fair comparisons. To maintain consistency, use the same time frames and conditions as the initial measurement. The Enterprise AI Adoption 2026 benchmark highlights five key dimensions for tracking post-launch performance:
- End-to-end cycle time
- Error or defect rate
- Rework or escalation volume
- Weekly active users and retention
- Incident rate, including policy violations
Calculating the impact is straightforward. For example, cycle-time reduction is (Baseline Time − Post-AI Time) ÷ Baseline Time × 100. Applying this simple formula to metrics like error rates provides clear percentage improvements that business leaders can easily understand without needing a background in statistics.
Use Small Eval Sets for Fast Checks, A/B Tests for Real Outcomes
A two-part testing strategy is crucial for balancing development speed with reliable results.
Offline evaluation sets of 20-50 real failure examples allow teams to quickly catch obvious regressions during development, a technique highlighted in Demystifying evals for AI agents. However, as improvements become more subtle, test suites must grow; detecting meaningful accuracy changes with confidence requires significantly larger sample sizes.
Online A/B tests answer a different, vital question: do users value the changes? While offline tests confirm a model's capability, live experiments measure real-world impact. Although they take longer to reach statistical significance, A/B tests are the ultimate guard against deploying features that don't deliver business value.
Keep the Measurement Loop Alive
AI performance is not static. Continuous monitoring is essential because factors like workflow drift, model updates, or seasonal demand can erode initial gains. Experts recommend re-running the original baseline test packet 30 - 60 days after any significant system change. If key metrics like cycle time or error rates begin to degrade, this practice allows teams to quickly roll back or adjust the AI before it negatively affects customers.