From a 3,000-Person Pilot to 100,000 Users: The BBVA AI Scale Case Study

How BBVA moved from a controlled ChatGPT Enterprise pilot to broad adoption through workflow selection, peer champions, leadership training, governance, and layered measurement.

AIZIGOO
From a 3,000-Person Pilot to 100,000 Users: The BBVA AI Scale Case Study

Enterprise AI stories are often reduced to an impressive demo or a single time-saving number. In a highly regulated global bank, however, turning an experiment into a new way of working requires more than model quality. The organization must decide where to start, which data may enter the system, who validates employee-built tools, and how usage differs from business impact.

According to public materials from BBVA and OpenAI, BBVA's ChatGPT Enterprise rollout began with roughly 3,000 employees in 2024 and expanded to more than 100,000 employees worldwide by 2026. Reported results include weekly active use above 70% among deployed employees, about three hours saved per employee per week, and efficiency gains of up to roughly 80% in selected workflows.

Those numbers can make the formula look simple: give everyone a strong AI tool. The more reusable lesson is the operating system around the tool—build trust in a bounded pilot, let domain teams discover use cases, and have leaders and control functions create a safe path to scale.

First, understand what the numbers are

The adoption and productivity figures in this article come from BBVA and its technology provider, OpenAI. They are not an independent randomized study or external audit. Time saved includes employee-reported data, and the “up to 80%” result describes selected workflows. This analysis therefore focuses on operating lessons rather than treating the figures as causal proof of company-wide financial impact.

PeriodPublicly described stageObservable signal
May 2024Initial rollout to about 3,000Cross-country, cross-function pilot
November 2024Five months in83% routine use; nearly 3,000 GPTs created
2025Expanded to 11,000 licenses20,000+ GPTs; about 4,000 used frequently
June 2026More than 100,000 users globally70%+ weekly active use; about 3 hours saved weekly
A small AI pilot expanding through training and review into standard work across many departments

1. A safe learning environment came before enterprise scale

In 2024, BBVA first gave ChatGPT Enterprise to roughly 3,000 employees across countries and business areas. The goal was not merely to count usage. It was to test productivity and innovation in real work while providing a shared enterprise boundary, rather than leaving employees to experiment independently in consumer tools.

Five months later, BBVA said 83% of licensed employees had incorporated the tool into their routine and had created nearly 3,000 custom GPTs. The meaningful leading indicator was not login count. It was whether domain knowledge had started to become repeatable tools.

An organization copying this approach should avoid selecting only AI enthusiasts. Mix legal, operations, customer support, finance, and technology—functions with different data and failure costs—so that scaling problems appear while the blast radius is still small.

2. The central AI team did not build every use case

A central team cannot encode every local workflow without losing context and creating a long queue. BBVA allowed employees to build custom GPTs and developed a network of AI champions and advanced “wizard” users. They ran hands-on sessions, helped peers apply the technology, and surfaced useful work inside each function.

Leadership participated as users, not only sponsors. Public material says 250 senior leaders, including the CEO and chair, received dedicated training. When leaders use the system directly, discussions can move from “adopt AI” to concrete blockers such as approval, data access, evaluation, and ownership.

A 2026 BBVA article describes an employee knowledge assistant in Mexico and Spain that handles more than 34,000 queries a month against a knowledge base of over 2,500 documents. It illustrates why a repeatable job with a known audience and trusted internal sources scales better than generic chat usage alone.

3. The important number is not 20,000 GPTs—it is the workflows that survived

OpenAI's 2025 case material says employees created more than 20,000 custom GPTs and around 4,000 were used frequently. That is not a clean success rate. Some tools may have been personal, temporary, overlapping, or experimental, and the public material does not define “frequent.”

The operating principle is still useful: keep idea creation broad, but keep the gate to becoming a standard tool narrow. Ask whether people return to it, whether outputs can be checked against a source, whether a human can detect failure, and whether it avoids replacing high-impact judgment.

Domain experts narrowing many AI ideas into a small set of validated reusable workflows
StageGate questionEvidence to retain
IdeaIs the recurring problem explicit?User, task, current time
PilotDoes AI reduce the real bottleneck?Before/after time, errors, samples
ValidationDoes it meet quality and control needs?Test set, approvals, failure rules
StandardizationCan another team reproduce quality?Owner, version, usage guide
OperationsDoes it stay safe after change?Usage, quality, incidents, regression tests

4. Efficiency was treated as one layer, not the whole outcome

The most concrete public workflow is an internal assistant in Peru used by more than 3,000 employees. It reportedly reduced average query handling time from about 7.5 minutes to about one minute; the case material describes this as roughly an 80% efficiency improvement. Faster handling is meaningful, but a complete scorecard must also examine accuracy, freshness, follow-up queries, and downstream cost from a wrong answer.

The reported three hours saved per employee per week needs similar care. Multiplying it by 100,000 people to manufacture a giant savings figure would assume equal use and that every saved hour becomes cash—neither is established. Read productivity in four layers:

  • Adoption: how many eligible employees become repeat users?
  • Process: how much do search, drafting, and handling time change?
  • Quality: do accuracy, edits, rework, and customer outcomes improve?
  • Risk: are privacy, bias, bad advice, and approval bypass controlled?

If quality or risk deteriorates, a faster process is not ready to scale. Conversely, a modest time saving may be valuable when it prevents a consequential omission or broadens access to expert knowledge.

5. Governance became a rail for scaling, not a final wall

Security, legal, and compliance are unavoidable in banking. A notable feature of the BBVA case is that these functions were aligned early rather than appearing only as a last-minute approval board. The organization combined an enterprise environment, enablement, responsible-AI principles, and workflow controls so employees could learn inside known boundaries.

BBVA has also published responsible-AI principles that connect value creation, human oversight, fairness, transparency, privacy, security, and accountability to development and daily use. Principles alone do not make a system safe, but they can become operational gates.

Security, legal, compliance, and human approval reviewing AI workflows through controlled gates
  • Separate public-information work from restricted internal data.
  • Require human review when outputs affect customer rights or financial decisions.
  • Keep evidence sources and freshness visible.
  • Name an owner, change record, and shutdown authority.
  • Feed errors, complaints, and model changes into regression tests.

Good governance is not a wall against every experiment. It is a set of rails that tells low-risk experiments how far they may travel. Clear risk tiers and approval rules can let domain teams learn faster inside the permitted zone.

A 90-day adaptation for another organization

You do not need BBVA's scale to test the underlying system.

Weeks 1–2: choose work and establish a baseline

Collect two high-frequency, checkable tasks from each participating function. Record current time, error and rework rates, data used, and final approver. Exclude high-impact decisions—credit, hiring, payment, or customer rights—from the first pilot, or place them behind mandatory human approval.

Weeks 3–6: run a mixed pilot with peer champions

Provide one approved environment to users from different functions and assign one champion per team. Measure completed work, edit reasons, and failures—not prompt count. Share stopped experiments as well as successful ones each week.

Weeks 7–10: narrow the reusable candidates

Keep workflows that are repeatedly used and measurably useful. Document input templates, evidence rules, forbidden actions, human approval, sample quality checks, and ownership. Merge overlapping personal tools into one managed version.

Weeks 11–13: apply a scale gate

Review adoption, time, quality, and risk together. Expand only workflows that pass, and turn failures into new evaluation cases. Archive or retire unused tools so the catalog does not become a graveyard of experiments.

Frontline users, AI champions, evaluators, and leaders improving workflows in a continuous evidence loop

A measurement board worth copying

LayerLeading indicatorOutcome indicatorWarning signal
AdoptionWeekly active and repeat usersRetention by functionLogins rise alone
EfficiencyHandling time, search stepsThroughput, lead timeRework rises
QualitySample accuracy, evidence linksErrors, complaints, repeat contactsConfident wrong answers
ReuseValidated tools and ownersCross-team reuseDuplicate tools explode
RiskHuman review and blocked actionsIncidents and violationsApproval bypass
LearningChampion activity and trainingFailures become testsOnly wins are shared

What this case does not prove

The case is strong evidence that large-scale adoption can be organized, but it does not prove that:

  1. all reported time savings became lower cost or higher revenue;
  2. every custom GPT is accurate, safe, and maintained over time;
  3. peak gains in selected workflows generalize across the bank;
  4. one model or vendor alone can reproduce the organizational change; or
  5. high-impact customer decisions can be automated without human review.

BBVA also entered this phase with years of digital transformation, AI Factories, and substantial data and technology capacity. Smaller organizations should narrow scope and invest proportionally more in security review, external assurance where appropriate, and explicit stop conditions.

Conclusion: the thing that scaled was the learning system

The attention-grabbing number is 100,000 users. The most reusable lesson is how BBVA decided what came after the first 3,000. Multiple functions learned in a bounded environment, peer champions spread practice, leaders trained as users, control functions defined scale conditions, and repeatedly useful workflows moved toward standardization.

Managing AI by seat count can grow activity without changing work. Managing it through a problem, owner, baseline, quality and risk gates, and a feedback loop can turn a small pilot into an organizational learning system. The pattern is simple to state and hard to execute: learn broadly, standardize narrowly on evidence, then expand again under control.

Sources

Product capabilities, public metrics, and company policies can change. Performance figures are published by BBVA and its provider and should not be treated as independent financial or causal validation. Verify current security, privacy, regulatory, and internal-control requirements before deployment.