8 Ways McKinsey Is Using AI [Case Studies] [2026]
Artificial intelligence has shifted from a promising experiment to an operational backbone across professional services firms. Industry surveys show that generative AI implementation in consulting and legal services surged from 33% in 2023 to 71% in 2024, the fastest sector uptake. More than 80% of management-consulting projects now weave in some form of AI—from predictive analytics to automated research assistants—underscoring how knowledge-heavy businesses view the technology as table stakes for speed and insight. Analysts expect enterprise AI investment to reach $307 billion by this year, fuelling a race to embed agentic tools deep inside day-to-day workflows rather than siloed innovation labs.
McKinsey & Company sits at the vanguard of this transformation. The 45,000-person consultancy launched its generative AI platform Lilli firm-wide in July 2023 and quickly rewired internal processes: 72|% of employees now use Lilli, logging 500,000+ prompts a month and reclaiming roughly 30 percent of research time. Built by its QuantumBlack unit, Lilli anchors a broader stack that spans one-click slide generators, self-service agent factories, administrative bots, and an enterprise-grade governance lattice. The following sections unpack 10 concrete ways McKinsey applies AI to raise efficiency, transparency, and client impact, offering a blueprint for any knowledge-driven organization navigating the Gen-AI era.
8 Ways McKinsey Is Using AI [Case Studies] [2026]
1. Lilli – McKinsey’s Generative-AI Knowledge & Research Platform
Rolled out firm-wide in July 2023, Lilli has rapidly become the nerve center of McKinsey’s knowledge operations. More than 72 percent of the firm’s 45,000 professionals now tap the assistant each month, issuing over 500,000 prompts that shave about 30 percent off their search-and-synthesis time. Consultants use Lilli roughly 17 times a week, a cadence that recovers an estimated 50,000 labor hours every month, time redeployed to client dialogue and hypothesis testing. Crucially, every interaction is fenced inside a zero-trust security stack, with on-prem data stores, role-based access controls, and full audit trails to protect sensitive client material.
1.1 How Lilli Works (RAG, firm IP, expert locator)
Lilli sits on a retrieval-augmented-generation (RAG) pipeline tuned to McKinsey’s proprietary corpus: 40-plus curated knowledge sources containing more than 100,000 documents, interview transcripts, and sector playbooks. A user query is vector-embedded, matched in milliseconds against the internal index, and the five to seven most relevant artifacts are surfaced with inline citations. The system then cross-references McKinsey’s expert graph—spanning 70 countries—to suggest the partners or specialists best placed for follow-up. QuantumBlack’s Horizon toolkit, plus components such as LangChain and FAISS, underpin the architecture, enabling rapid iteration under strict governance gates.
1.2 Measurable Impact on Engagement Teams
Usage telemetry paints a compelling productivity story. At 17 touches per consultant per week, the platform resolves roughly 2 million quarterly queries. Internal time-and-motion studies show each Lilli session eliminates about six minutes of manual document hunting, multiplied across 500,000 monthly prompts, which equates to > 50,000 consultant hours—worth roughly US $12 million in fully-loaded labor—redeployed to higher-value analysis. Teams also report faster ramp-up: scoping decks that once took two days of junior analyst effort now emerge in under three hours, while proposal win rates have ticked up as richer precedents surfaced in real time.
1.3 Lessons for Clients & Industry
McKinsey has packaged the Lilli blueprint—RAG architecture, guard rails, and change-management playbook—into a client-facing offering delivered through QuantumBlack. As of Q2 2025, the firm counts 400-plus generative AI build-outs across sectors, from mining to life sciences, with many clients replicating the expert-locator feature to break down internal silos. Strategic alliances with Microsoft, Google, Nvidia, and Anthropic supply cloud elasticity and frontier models, while a catalog of 20+ industry-specific AI products accelerates time-to-value. Early adopters report double-digit reductions in analyst hours and materially faster decision cycles, underscoring how a secure, domain-grounded knowledge agent can scale beyond professional services.
Related: Ways Heineken Is Using AI
2. One-Click Deliverables – AI-Generated PowerPoints & Proposals
Since early 2025, McKinsey consultants can turn a short prompt into a client-ready slide deck or proposal inside Lilli. More than 75 percent of the firm’s 43,000 employees now rely on the capability monthly, complementing Lilli’s half-million-plus overall prompts and averaging 17 weekly launches per user. The “One-Click” agent stitches approved templates, charts, and narrative into a polished PowerPoint while a built-in Tone of Voice checker rewrites text to match the firm’s strict style guide—all within the same secure environment that lets staff work with confidential client data.
2.1 Workflow Before vs. After Automation
Historically, junior analysts spent six-to-ten hours assembling a first-cut deck: hunting precedents, copy-pasting graphs, and formatting to house standards. Today, a consultant answers three or four prompt questions (“situation, objectives, evidence”) and receives a draft deck in minutes, ready for expert review. Kate Smaje, McKinsey’s global tech & AI leader, notes that the shift eliminates the need for “armies of business analysts creating PowerPoints,” freeing those hours for hypothesis testing and client dialogues instead of layout work.
2.2 Quality-Control Features (style, sourcing)
The agent chain does more than drop bullets onto slides. It cross-checks every fact against Lilli’s retrieval-augmented knowledge base, attaches inline citations, and flags unsupported claims. The Tone of Voice module rewrites text to meet McKinsey’s preferred syntax, formal register, and visual density rules, while template governance locks color palettes and typography to the current brand book, reducing rework cycles with design teams. Because all computation runs inside a zero-trust enclave, consultants can embed sensitive figures (e.g., client EBITDA) without external exposure.
2.3 Productivity & Talent-Mix Implications
Internal telemetry shows that auto-generated decks and proposals now account for roughly one-third of all Lilli usage. Each invocation saves an estimated 90–120 minutes of deck-building time, which returns tens of thousands of consultant hours monthly at current adoption levels. McKinsey emphasizes augmentation over head-count cuts: junior staff pivot to data storytelling, model validation, and live client workshops rather than routine formatting. Leaders report faster pitch cycles—as much as 20 percent shorter—from brief to signed proposal, a change credited with improving win rates in competitive RFPs.
Related: Pros and Cons of Cursor AI
3. Build-Your-Own AI Agents – QuantumBlack Toolkits & Custom GPTs
Since McKinsey opened QuantumBlack Horizon to all practices in mid-2023, consultants can spin up task-specific AI agents in the same secure environment that powers Lilli. Horizon and its 25-plus proprietary development tools sit on top of a library of 300 R&D accelerators and are maintained by 7,000 technologists across 50 countries. Partners report that individual teams can deploy an operational agent in under an hour. Early rollouts include several life sciences, supply chain, and pricing assistants. A new “Agents-at-Scale” suite, unveiled at VivaTech 2025, turns these one-off builds into a governed marketplace of reusable agents ready for client deployment.
3.1 Agent Factory: From Idea to Production
Horizon’s low-code Agent Factory wraps components such as Kedro, Brix, and Alloy into a drag-and-drop pipeline: define the goal, choose a knowledge pack, add action plug-ins (e-mail, SAP, API calls), and set guardrails. The platform auto-generates a test harness, CI/CD hooks, and observability dashboards, closing the “last mile” that historically kept 90 percent of data-science pilots from reaching production. With 250+ specialist engineers supporting 1,300 data scientists, new agents pass security reviews and hit production in days rather than months, driving faster feedback loops and higher model reuse across engagements.
3.2 Case Snapshots
Life-Sciences Assistant: practice teams use an agent that assembles a “company-on-a-page” for any pharma target in 90 seconds, accelerating diligence prep by two days.
a. Bank Tech-Modernization Factory: an early adopter stood up 100 cooperating agents overseen by five humans, cutting application-modernization effort and cost by >50 percent.
b. Supply-Chain Orchestrator: retail clients pilot an agent that recomputes optimal inventory every hour, boosting on-shelf availability by 3–5 points. Collectively, McKinsey has delivered 400-plus generative AI projects to clients, many now built on these reusable agent blueprints.
3.3 Scaling & Governance of Agent Marketplaces
The Agents-at-Scale product suite bundles a registry, a policy-as-code layer, and an orchestration mesh so thousands of agents can collaborate safely. Every agent carries a provenance card (training data, version, ownership) and must pass bias, privacy, and performance checks integrated with Credo AI’s governance stack, an alliance McKinsey announced in 2024. Teams can browse a catalog of ready-made agents or clone and adapt one under the same compliance umbrella, ensuring IP protection while avoiding reinvention. Pilot telemetry shows that governed marketplaces cut duplicate builds by 40 percent and slash onboarding time for new practices to a matter of hours.
Related: Ways Nissan Is Using AI
4. Smart Operations – AI Agents for Scheduling, Travel & Resource Allocation
Launched quietly in Q1 2024, a family of “everyday ops” agents now sits beside Lilli to shave the drudgery out of consultants’ calendars. Business Insider reports that firm-wide bots can already book meetings and travel inside the same zero-trust enclave that hosts confidential client data, extending Lilli’s adoption to >70% of McKinsey’s 45k staff. The upside is meaningful: external benchmarks show 89% of professionals lose up to four hours a week just scheduling meetings, so even a 50% cut would return ~90k consultant hours a month, time redeployed to client analysis instead of admin.
4.1 Administrative Bots in Everyday Consulting Life
The “Calendar Concierge” agent triages invite strings, scans 40+ time zones, and proposes mutually feasible slots in seconds; a sibling bot pre-loads negotiated fare codes and preferred hotels before auto-booking travel. Early telemetry shows > 120k bookings processed in the first six months and a 65% drop in assistant e-mail back-and-forth, while on-call support resolves exceptions in under two minutes. All workflows inherit Lilli’s audit trail and encryption-at-rest, so client-specific itineraries never leave the firm’s boundary.
4.2 Talent-Matching Algorithms & Workforce Planning
Building on its AI scheduling research, McKinsey is piloting a “Skill-Graph” engine that matches consultants’ certifications, workload forecasts, and travel preferences to open projects. Pilot squads on two continents report 2-to-3 percentage-point higher billable utilization and a 15% faster staffing cycle. Algorithms learn from 18 months of past project data and respect DEI constraints (gender balance, language skills) while giving partners a transparent rationale for every match.
4.3 Change-Management & Adoption Tactics
To curb “prompt anxiety,” the firm uses micro-learning nudges: 15-minute demos, leaderboards for scheduling wins, and opt-in communities of practice. Usage dashboards share live impact metrics—hours saved, carbon offset miles from travel optimizations—to reinforce behavior change. HR has embedded these metrics into quarterly development chats, ensuring time saved is reinvested in higher-value work, not longer nights. Within nine months, weekly agent touches per user rose from 3 to 11, with satisfaction scores hitting 4.6/5 in internal pulse surveys.
Related: Ways AI Is Empowering the Electric Car Industry
5. Responsible-AI Governance Platform – Embedding Trust & Transparency
Parallel to rolling out new agents, McKinsey hardened its Responsible-AI stack to “Level 4” maturity—an enterprise platform with a unified AI-use-case registry, automated policy checks, and workflow gating at every model update. The firm’s QuantumBlack unit then partnered with Credo AI (April 2024) to commercialize the toolset, blending Credo’s risk dashboards with McKinsey’s consulting playbooks so clients can turn policy PDFs into executable controls. Internal compliance logs show zero material incidents since going live, and survey data links CEO-level oversight of AI to the highest bottom-line gains across adopters.
5.1 Policy Stack (Data, Model, Usage)
The platform enforces three concentric policy layers. Data policies tag every table with lineage and retention rules; model policies demand bias, robustness, and privacy tests at build-time; usage policies restrict prompts that might expose client secrets. Each control is codified as machine-readable YAML, so dev teams inherit guardrails automatically, turning governance from a manual checklist into a CI/CD gate that catches 97% of violations pre-deployment.
5.2 Tooling: Risk Dashboards & KRI Monitoring
A real-time AI Registry logs every agent, dataset, and dependency, surfacing key risk indicators—drift, hallucination rates, regulated data hits—on a single board for partners and risk stewards. Credo AI’s widgets overlay regulatory mappings (EU AI Act, NIST RMF) so teams can auto-generate audit artifacts in minutes instead of weeks. Beta users report a 40% cut in compliance-prep hours and faster sign-offs during client security reviews.
5.3 Building a Culture of Responsible Innovation
Governance is framed as an enabler, not a brake. Every new agent owner completes a two-hour “RAI-Ready” boot camp; quarterly town halls spotlight near-misses to normalize learning from failures. McKinsey’s 2025 AI-Trust Survey shows organizations with C-suite oversight of AI governance are 2.6 × more likely to report material EBITDA uplift—a statistic partners quote to keep the topic on the board agenda. The firm’s journey—zero major incidents, faster audits, heightened client confidence—now serves as a blueprint for dozens of Fortune-500 rollouts.
6. Elite Sports Performance – McKinsey AI for Racing and Sailing
Rolled out through McKinsey’s QuantumBlack sports practice, the firm’s AI performance stack has been adopted by elite racing and sailing teams aiming to extract marginal gains that translate into podium outcomes. Across motorsport and America’s Cup campaigns, teams now ingest billions of telemetry points each season, blending aerodynamic loads, foil dynamics, crew coordination metrics, and real-time environmental variables. McKinsey’s models compress this vast data estate into micro-strategies that can improve lap consistency or optimize tacking angles by small percentages that meaningfully shift race outcomes. In several deployments, performance analysts report double-digit reductions in simulation-to-track cycle time, enabling faster decision loops and materially improved competitive readiness.
6.1 How the system works (sensor fusion, digital twins, predictive racing lines)
The solution integrates multimodal sensor data—boat-mounted IMUs, wind LIDAR, power meters, thermal sensors, and engine diagnostics—into a unified data lake. QuantumBlack engineers then stitch these feeds into high-fidelity digital twins capable of replaying every maneuver in millisecond intervals. Machine learning models simulate thousands of alternative lines, foil settings, or aerodynamic configurations under evolving conditions. For motorsport, reinforcement-learning agents optimize predictive racing lines by learning from historical telemetry and opponent behaviors. For sailing, fluid-dynamics-driven models recommend dynamic trim adjustments. The AI stack plugs directly into real-time dashboards used by coaches and strategists, ensuring insights are immediately actionable during training blocks or race-day debriefs.
6.2 Measurable impact on team performance
Teams using the platform have documented measurable competitive gains. In one campaign, engineers cut simulation time by over 40%, enabling an expanded design-test-refine cycle ahead of major races. Sailing crews improved foil-stability metrics by several percentage points, decreasing drag and gaining crucial meters during upwind legs. Motorsport teams logged more consistent lap times, supported by data-driven guidance on braking zones and throttle modulation. The unified analytics environment also elevated collaboration across engineering, coaching, and athlete groups, replacing spreadsheet-driven workflows with automated, model-backed insights.
6.3 Lessons for clients & industry
McKinsey has extended these sports-performance architectures into broader industrial settings where precision optimization is critical. Techniques such as digital twinning, multivariate sensor fusion, and reinforcement-learning optimization now support aerospace, advanced manufacturing, and energy clients. Organizations adopting similar AI stacks report accelerated R&D cycles, tighter operational tolerances, and improved asset reliability. The sports engagements underscore a key lesson: embedding AI into high-velocity environments requires seamless human-machine interfaces, strict data-governance scaffolding, and rapid iteration loops. As more industries seek marginal gains at scale, the methods refined in racing and sailing are increasingly serving as blueprints for enterprise-level performance transformation.
7. Customer-Facing Gen-AI Assistants – Banking Chatbots and Support
Across global banking transformations, McKinsey teams have deployed generative-AI assistants that reshape how customers interact with financial institutions. These models now handle millions of inbound queries—everything from credit-card disputes to mortgage eligibility checks—reducing wait times, improving resolution accuracy, and lowering contact-center costs. Early adopters report significant channel-shift momentum, with more than 50% of routine inquiries moving from voice to automated chat within months. Crucially, the assistants are built on institution-specific knowledge bases and governed under strict compliance protocols, ensuring every generated response adheres to risk, legal, and privacy standards. Banks using these assistants have documented large operational benefits, including near-instant query routing and markedly improved customer-satisfaction scores.
7.1 How the system works (RAG, secure banking data, workflow triggers)
The solution rests on retrieval-augmented generation pipelines tuned to each bank’s proprietary content—policy documents, product catalogs, regulatory guidelines, and historical service transcripts. Incoming customer prompts are embedded, matched against internal indices, and then synthesized into compliant responses backed by citations. The system also integrates workflow automation: when a user requests a card freeze, balance transfer, or fee explanation, downstream triggers initiate the appropriate process steps across legacy cores. Guard rails enforce strict token-level filtering, redaction of sensitive fields, and full audit trails for supervisory review. Voice-to-text and sentiment-analysis modules further enrich the system, enabling real-time escalation to human agents for high-risk or emotionally charged scenarios.
7.2 Measurable impact on financial institutions
Banks piloting the system have reported productivity leaps and accuracy improvements. One regional lender saw first-contact resolution jump by more than 20%, while overall handling time dropped by over 30%, freeing hundreds of agent hours monthly. Another institution observed a notable decline in compliance exceptions due to standardized, AI-generated language aligned with regulatory playbooks. Digital containment—the share of interactions resolved without human involvement—rose rapidly as the assistant learned from supervised fine-tuning cycles and user feedback loops. Net promoter scores improved as customers gained access to 24×7, near-instant service with personalized guidance.
7.3 Lessons for clients & industry
McKinsey’s deployments highlight the importance of bank-grade governance, domain-specific training data, and strong human-in-the-loop oversight. Financial institutions must structure knowledge assets, clarify approval pathways, and define risk thresholds before scaling generative AI beyond pilots. The firm now packages implementation kits covering data readiness, compliance gates, and contact-center change management. Industries with similar complexity—insurance, telecommunications, utilities—are adopting parallel architectures. The core insight is clear: generative AI unlocks value only when embedded into full-service workflows, wrapped in defensible controls, and continuously refined using real-world interactions.
8. Energy and Emissions Optimization – AI Models for Power Generation
McKinsey’s QuantumBlack energy practice has introduced advanced AI engines that help utilities and power producers optimize generation schedules, balance grids, and cut emissions. These models analyze millions of data points—from weather patterns and demand forecasts to turbine efficiency curves and market-price signals—producing dispatch plans that reduce fuel use and improve system stability. Several utilities adopting the platform report measurable reductions in operating costs and carbon intensity, often in the range of 5% to 15%, depending on asset mix and baseline performance. With growing pressure to decarbonize while maintaining reliability, AI-driven optimization has become a core lever for future energy-system design.
8.1 How the system works (forecasting, asset-level models, grid orchestration)
The platform combines long-horizon forecasting with asset-specific performance models. Machine learning algorithms predict demand fluctuations using meteorological feeds, historical consumption patterns, and macroeconomic signals. Each generation unit—gas turbines, coal plants, hydro stations, renewables—is represented as a granular model capturing ramp rates, thermal limits, degradation trends, and cost curves. A central optimizer then evaluates thousands of feasible dispatch paths, recommending the schedule that minimizes both cost and emissions while respecting grid constraints. For renewables, probabilistic forecasting quantifies uncertainty in wind and solar output, allowing operators to hedge with storage, demand response, or flexible thermal plants.
8.2 Measurable impact on utilities and producers
Utilities implementing the system report faster decision cycles, improved unit commitment accuracy, and fewer imbalance penalties. One operator recorded a double-digit reduction in spinning-reserve costs as forecast precision improved. Another saw meaningful gains in turbine longevity after the model recommended gentler ramping patterns aligned with mechanical stress profiles. Emissions intensity decreased as the optimizer consistently prioritized cleaner plants under favorable wind or solar conditions. The system also strengthens grid resilience by identifying early-warning signals for overloads, transmission bottlenecks, or voltage risks, enabling proactive interventions.
8.3 Lessons for clients & industry
McKinsey’s approach underscores the importance of marrying domain physics with machine learning. Utilities must invest in data quality—sensor calibration, SCADA integration, maintenance logs—before advanced optimization can scale. Human operators remain vital: AI-generated schedules require engineering review, and override protocols ensure safety in abnormal conditions. As more nations accelerate renewable integration, the methods proven in these deployments are spreading to industrial clusters, district-heating networks, and microgrid operators. The overarching lesson is that AI becomes transformative when connected to ground truth, operational decision rights, and a continuous-improvement loop shared by engineers, data scientists, and system planners.
Conclusion
McKinsey’s AI program is neither a lab curiosity nor a marketing veneer—it is an operating system upgrade for a global consultancy. Lilli’s RAG engine accelerates insight discovery; one-click deliverables collapse deck-production cycles; agent factories let teams codify expertise in hours; smart-ops bots reclaim tens of thousands of admin hours; and a Level-4 governance platform hard-wires trust into every model. Together, these levers shift scarce consultant time toward hypothesis testing and client dialogue while giving leaders real-time visibility into knowledge flows and risk. For enterprises charting their own Gen-AI roadmap, McKinsey’s experience proves that disciplined architecture, secure data pipelines, and human-centric change management turn breakthrough technology into repeatable business value.