Use Cases Are the Beginning, Not the Goal
How concrete AI projects can gradually build organizational capabilities.
Why companies start with use cases
The pattern has become familiar. A company holds an AI workshop. Digital sticky notes are collected, working groups produce longlists, and the ideas are then sorted in a matrix by expected value and feasibility. Summarization, knowledge search, forecasting, document review, customer communication. At the end there may be 40 use cases, five favorites, and the instruction to build a pilot quickly.
One reason for writing this text was an article by former colleagues at McKinsey. In Capturing Central Europe’s AI opportunity, Czímer, Van der Veken, and Zetek (2026) argue that companies should not treat AI as a side initiative, but redesign business domains and workflows around AI. More recent McKinsey publications develop the same point further: value emerges less from additional tools than from changed workflows, operating models, responsibilities, and capabilities (Schmitz et al. 2026; De Smet et al. 2026; Weddle 2026). I agree with the core observation, but I would not read it primarily as a call for a broad transformation program. The more interesting question, in my view, is how companies can move from concrete applications to reliable organizational capabilities.
This starting point is reasonable. The question “Where can we use AI?” translates an abstract technology into an operational context. It creates a common object of discussion for business functions, IT, data science, and leadership. It also leads more quickly to testable assumptions than an AI strategy that consists mainly of target pictures and maturity models.
At the same time, I see a recurring problem. Some companies treat the use-case list as the strategy. They collect ideas and build prototypes without clarifying how processes, roles, data access, controls, and operational ownership have to change. In these cases, the number of pilots increases while the organization’s ability to use AI reliably changes only marginally.
In my view, companies should therefore begin with concrete AI use cases, but understand this starting point as part of a longer learning process. Well-chosen use cases can gradually trigger changes in processes, responsibilities, and infrastructure. Transformation then emerges from working on concrete problems, rather than independently of them.
Why this starting point is useful
AI only becomes assessable in a concrete work situation. “We use a large language model” says little. “We shorten the review of incoming contract drafts, flag uncertain passages, and have them approved by a lawyer” describes a process, users, an expected effect, and a boundary of automation.
Good use cases serve several functions. First, they create a shared basis for discussion, because a business function does not need to understand model architecture in order to talk about waiting times, errors, or recurring manual work. They make expected value testable, because processing time, quality, throughput, or error rates can be compared with a baseline. They also reveal dependencies: the concrete case shows whether the required data is accessible, current, and legally usable. Finally, responsibilities have to be clarified. It becomes specific who uses the result, who intervenes when errors occur, and who decides on changes. Technical, regulatory, and organizational limits become visible while the work is being done.
This concreteness matters because AI adoption is increasing quickly, but not evenly. According to OECD (2026), 20.2 percent of firms in the covered countries used AI in 2025; among large firms the share was 52.0 percent, while among small firms it was 17.4 percent. The average therefore hides very different starting conditions. A generic transformation recipe would be unconvincing for that reason alone.
Productivity effects also do not arise simply because a model is available. In a large field study with 5,172 customer support workers, Brynjolfsson, Li, and Raymond (2025) show that a generative AI assistant increased the number of resolved cases per hour by an average of 15 percent. The effects were especially large for less experienced workers and small for highly experienced workers. This is a strong finding for a specific, workflow-embedded use of AI. It is not evidence of an automatic productivity increase in every company.
Where the use-case approach reaches its limits
A prototype usually tests whether a model can perform a task in principle. A production system also requires decisions about data access, workflow integration, behavior under uncertainty or failure, quality and cost control, and responsibility after the project phase.
Many initiatives stall at this transition. Typical signs include manual data exports, individual prompts without versioning, unclear roles and permissions, or missing integration into the systems actually used in day-to-day work. Continuous monitoring of quality and cost is often missing as well, as is a named responsibility for operations and further development. Under these conditions, a technically working proof of concept is not yet a reliable application.
A demo can hide all of this. A small team uploads selected example documents, manually corrects difficult inputs, and presents the best outputs. This is legitimate for early learning. It becomes problematic when this controlled situation is mistaken for operational maturity.
A recent McKinsey survey illustrates the gap: McKinsey & Company (2025) reports that almost two thirds of respondents said their organization had not yet scaled AI enterprise-wide. This is a consulting survey based on self-reports, not a neutral census. It also does not prove that every company should scale AI broadly as quickly as possible. But the finding fits a familiar pattern: usage and experimentation grow faster than organizational embedding.
The Germany-focused study Produktivität. Neu gedacht., published in June 2026, differentiates these intermediate stages further. McKinsey & Company (2026) reports that, of 80 surveyed CxOs, 26 percent placed their company in the pilot phase, 47 percent reported active scaling, and 11 percent described their organization as “AI-first”. The authors state that the survey is not representative; 51 percent of participants work in companies with more than five billion euros in revenue. The numbers therefore should not be read as a picture of the German economy as a whole. What is more relevant is the conceptual gap between rolling out pilots and actually anchoring AI in core processes and decisions.
This gap also appears in newer McKinsey articles. De Smet et al. (2026) distinguish three horizons of enablement, automation, and reinvention, and report from a survey of 750 employees that leaders are more likely to see enterprise-wide value when workflows are redesigned rather than when tools are merely provided. Schmitz et al. (2026) make the same point in organizational terms: many companies accelerate existing tasks without adapting decision paths, governance, teams, and capabilities. For my argument, the exact percentage is less important than the recurring diagnosis: the bottleneck is often not access to models, but the redesign of work.
The German study also calls for AI to be understood as a “new operating system” of the company. As an image, this makes the organizational reach of AI visible. As a practical instruction, it is too broad for me. Not every company and not every process has to become “AI-first”. What matters is whether a concrete application creates a better process and whether the capabilities built in the process can be used elsewhere. Transformation should not become an end in itself.
It is also important that not every pilot has to move into production. A deliberately stopped experiment can be valuable if it disproves a testable assumption, reveals a risk, or creates reusable knowledge. What matters is a transparent decision about whether to continue, adapt, or stop the initiative.
Lessons from analytics transformations
The tension between individual projects and transformation is not new. Data warehouses were meant to create a reliable data foundation. Business intelligence was meant to make decisions more transparent. Data lakes promised flexible access, and advanced analytics promised better forecasts. Many organizations invested in platforms, dashboards, and central teams, only to find that decisions did not automatically become more data-driven.
One explanation lies in the interaction between technology and organization. Brynjolfsson and Hitt (2000) show that the value of IT investments depends substantially on complementary investments in processes, capabilities, and organizational structures. Mikalef et al. (2020) describe big-data analytics as an organizational capability in which technical, human, and management-related resources interact.
At the same time, large transformation programs without concrete applications were not a guarantee of change either. A central data platform can be technically impressive and still irrelevant for business functions. A center of excellence can formulate standards without improving a single operational process. Vial (2019) defines “digital transformation” not as the mere introduction of new technology, but as changes triggered by the combination of digital technologies. The outcome should therefore not be confused with the technology project.
In my view, this means that companies should not begin by building the most comprehensive AI platform possible. A concrete use case can serve as a practical test bed. It shows whether data is available, whether decisions change, whether employees use the system meaningfully, and whether the organization can sustain operations. At the same time, a successfully completed case should leave behind components, standards, or experience that can be reused in further applications.
Additional requirements created by generative AI
Classical analytics systems often provided information or forecasts. Generative AI reaches further into knowledge work: it drafts text, condenses documents, prepares decisions, generates code, and communicates with users. AI agents go one step further. They can call tools, query data, create records, or coordinate multi-step workflows.
In an article for The European Business Review, Niessing, Feldmann, and Bücker (2026) described this transition as a shift from pilots to pipelines. The term pipeline is not only technical here. It refers to a designed workflow with roles, inputs, checks, handovers, and measurable outputs. For the argument in this text, the key point is that agentic AI should not be understood as an additional chatbot next to the process, but as part of a deliberately designed workflow.
This increases the organizational reach. An isolated chatbot processes inputs and produces answers. An agent that reads customer data, prepares a goodwill decision, and triggers a booking in an operational system needs identities, permissions, logging, thresholds, and a defined escalation path. In that case, errors can propagate across several connected systems.
The frequently used term “human in the loop” refers to the involvement of a person in an automated process. As an operating concept, this phrase is not sufficient. It has to be specified which person reviews which outputs, based on which information, and within which time frame. The decision authority, detection of systematic errors, and investigation of incidents also have to be defined. Human oversight is therefore a designed process.
The NIST framework published by Tabassi (2023) organizes AI risk management into the four functions govern, map, measure, and manage, and treats governance as an ongoing cross-cutting task. This logic is useful beyond formal compliance: risks have to be assigned to a context, made measurable, treated, and organizationally owned. This cannot simply be attached to a finished model afterwards.
Criteria for a robust AI use case
Many prioritization matrices evaluate only expected value and technical feasibility. That is sufficient for an initial sorting, but it favors ideas that are easy to demonstrate. In my view, a good AI use case should answer eleven questions:
- Problem: What concrete problem is being solved, and for whom?
- Process: Which decision or work step changes?
- Baseline: How good, fast, or expensive is the current process?
- Benefit: How often does the process occur, how much time or money does it currently consume, and what improvement would make the use case worthwhile in economic or organizational terms?
- Data: Which data is required, and is the system allowed to use it?
- Quality: Which tests and metrics will be used to assess performance?
- Risk: Which errors are likely, and which would have serious consequences?
- Control: Where do humans decide, where does the system automate, and how is escalation handled?
- Operations: Who is responsible for cost, monitoring, incidents, and further development?
- Integration: How does the result enter the real workflow without media breaks?
- Reuse: Which data access patterns, components, standards, or capabilities remain available for further cases?
A demo use case is aimed at proving a technical possibility. It often works with selected example data and ends at the model output. This can be useful in an early exploration phase. A robust use case additionally has to clarify how the result enters a real decision, how it is evaluated against a baseline, and who carries business responsibility.
The term ownership refers here to ongoing business responsibility. It goes beyond sponsorship and budget approval. IT can operate a system and data science can evaluate it. But the responsibility for whether it works sensibly in the business context, and what consequences its output may have, must be anchored in the relevant process.
Use cases as a starting point for capability building
By use-case first, I mean an approach in which the work begins with a small number of concrete applications. At the same time, the shared technical and organizational capabilities that will be needed later should already be considered. This avoids long lead times for a comprehensive target architecture without isolating individual projects from each other. In my view, a small portfolio of different cases is useful when they promise testable value and build shared capabilities at the same time.
The following figure summarizes this logic. A use case begins with a concrete problem, is moved through evaluation and integration into a production application, and ideally leaves behind a reusable capability. That capability then feeds back into further cases.
The concept of a domain is helpful here. A domain is a coherent area of business responsibility, such as sales, customer service, supply chain, engineering, or claims handling. Czímer, Van der Veken, and Zetek (2026) argue that such domains are large enough to be financially relevant while still bounded enough to redesign workflows end to end. This fits the use-case-first logic: an individual use case should not be optimized in isolation, but should show within a domain which data, roles, interfaces, and controls are needed on a durable basis.
One example is a sequence of three applications. The first production case establishes governed access to internal documents. The second uses the same identity and permission architecture and adds systematic evaluation. The third connects the model to an operational system through an API and introduces audit logs. In this way, a platform capability gradually emerges as a generalized solution to problems already encountered. This logic also fits the argument by Singla et al. (2026) and Czímer, Van der Veken, and Zetek (2026) that successful companies do not roll out many arbitrary use cases in parallel, but select a few economically relevant domains and build reusable capabilities within them.
The feedback effect is central. Beyond its immediate outcome, each use case should make the organization more capable in a bounded area. Depending on the case, this may mean reliable data and API access, roles and permissions, evaluation and approval standards, or governed prompt and model management. More integrated applications also require human-in-the-loop and escalation processes, monitoring, audit logs, and clear operational responsibility. AI literacy, meaning the ability to use and assess AI systems appropriately in the respective work context, is also one of these reusable capabilities.
Not every capability has to be fully developed in the first project. But every project should explicitly state which capabilities it needs, which it builds, and who will maintain them afterwards. Otherwise the same foundational questions are answered again in every pilot.
What this means for leadership and organization
Leadership should not define every individual use case centrally, but its role also should not be limited to budget approval. Leadership sets the decision logic: the relevance of the problems addressed, the acceptable risks, the reusable capabilities to be built, and the criteria for continuing or stopping a pilot.
For the organization, this implies shared responsibility. Business functions own the problem and the process. Data and AI teams develop and evaluate the technical solution. IT and security ensure integration and operations. Privacy, legal, and worker representation should be involved early where the case requires it. Central governance should create minimum standards and transparency without forcing every small exploration through the same approval process as a high-risk production system.
For employees, generic prompt training is not enough. People working with AI in a real process need to understand the task, the limits of the system, and their own decision authority. AI literacy is therefore role-specific. A claims handler needs different competencies from a product owner, a developer, or a board member.
Infrastructure should also not begin with a maximal platform vision. Standardization is useful where several real cases have the same problem. A shared model access layer, identity management, or evaluation service should therefore demonstrably make several use cases easier to implement. Reuse emerges through deliberate product decisions, not automatically through the centralization of technology.
Conclusion
Concrete use cases are a useful starting point because value, risks, and organizational prerequisites can be examined at the level of a real problem. But durable AI use requires more than collecting applications and demonstrating technical feasibility. Process integration, quality assurance, responsibilities, and possible operations have to be considered already during a pilot.
In my view, an AI strategy should therefore connect manageable problems with measurable changes in the process. At the same time, every project should examine which components, standards, and competencies will remain available for further applications. A single use case does not create organizational transformation. A sequence of well-chosen and productively embedded use cases can gradually build the capabilities required for it.