AI Accelerates Everything Except the One Thing Innovation Requires

February 26, 2026
Urquhart Wood

AI Accelerates Everything Except the One Thing Innovation Requires

The “one thing” is knowing what’s worth building. AI can process data, generate insights, and produce recommendations faster than any team. But knowing what your customers are trying to accomplish and where they struggle requires a kind of learning that comes from doing the work, not reading the output. Recent research helps explain why, and a closer look at what discovery actually requires reveals something beyond the research.

What the Research Found

Researchers at the Wharton School recently published a study on how AI reshapes human reasoning. Shaw and Nave ran three preregistered experiments with over 1,300 participants and nearly 10,000 individual trials. They gave participants reasoning problems with optional access to an AI assistant. Without participants knowing, the AI was programmed to sometimes provide correct answers and sometimes provide confidently stated wrong answers.

When the AI was right, participants performed significantly better than those working without AI. No surprise. But when the AI confidently provided wrong answers, participants adopted those wrong answers roughly 80% of the time. Their accuracy dropped well below what they would have achieved on their own, with no AI at all.

The researchers call this cognitive surrender: adopting AI outputs with minimal scrutiny, bypassing the deliberate thinking that would otherwise catch mistakes.

But here is the finding that matters most. Participants with AI access reported significantly higher confidence in their answers, even though half the AI outputs were wrong. People felt more certain while being less accurate. And they had no idea it was happening.

The study used logic and reasoning puzzles, not innovation decisions. Nobody has yet replicated this experiment in the context of customer research, so we have to exercise some intellectual humility about the implications for customer research. But the conditions that produced cognitive surrender in the lab are the same conditions innovation professionals navigate every day: time pressure, complexity, ambiguity, and the desire for clear answers where certainty is hard to come by.

In the Wharton study, the tasks had correct answers that participants could, in principle, have recognized. They didn’t. In customer research, the truth lives with the customer. It may or may not be in the AI database, but without doing the discovery work with customers, the team has no way to know. There’s nothing to contradict what the AI says. The only error signal comes later, when the product fails in the market and the investment has already been made. The conditions for uncritical adoption are, if anything, more favorable.

A Gap That Was Already There

Even the best teams carry assumptions about how well they understand their customers. That is not a failure. It is a natural consequence of being close to the work.

Three recent independent studies, using different questions and samples, point to the same pattern: leaders believe their organizations understand customers far more than customers believe they are understood.

  • In a PwC study, a majority of business leaders said their company understands customer needs well, while only 18% of customers felt that most companies do.
  • In a Gartner study, 77% of executives reported high confidence in their organization’s understanding of customers, compared with just 27% of customers who felt companies understand them.
  • In Salesforce research, 78% of business leaders said their company understands customers’ needs and expectations, versus 30% of customers who felt companies actually do.

This gap existed before generative AI. Now layer the cognitive surrender research on top of it. Teams already overestimate their understanding of customers. AI produces articulate, authoritative outputs that reinforce that overconfidence. And the research tells us people adopt those outputs with less scrutiny while feeling more confident. The result: faster, more articulate answers to the wrong question. And everyone feels great about it.

Two Kinds of Knowledge

Most organizations have enormous quantities of operational data: how customers use current products, which features they adopt, where they drop off, what they rate highly. AI is extraordinarily good at analyzing this kind of data.

But operational knowledge is not innovation knowledge. Operational data tells you what customers do. It does not tell you what customers are trying to accomplish, how they measure success, or where they struggle independent of any particular product. Where they are struggling is where the opportunity for new value creation lies.

There is an old joke about a man searching for his lost keys under a streetlight. When asked where he dropped them, he points across the street. “Then why look here?” “Because this is where the light is.”

When teams point AI at operational data and ask, “What do our customers need?”, the AI delivers fast, comprehensive, impressively presented answers. The analysis is real. The patterns are valid. But the conclusions are limited by the same data the organization has always had, just processed more efficiently. The light is bright where your operational data lives. But your innovation opportunities are somewhere else entirely.

Before the Discovery Work Begins

There is a problem that precedes the data quality question entirely.

Before any discovery work begins, someone has to make a series of strategic choices that no dataset contains. Who is the actual target customer? What job defines the market we want to enter or create? And which discovery approach is appropriate given your firm’s objectives, competitive position, and risk tolerance? These are not questions with universal answers. They are strategic choices the firm must make, and the answers determine the discovery process downstream.

Depending on the firm’s objectives and situation, the right discovery strategy might focus on the core functional job of a primary customer, or it might trace the full chain of people involved in a purchase and consumption process, or it might specifically seek customers who are overserved by current solutions and ready for a simpler alternative. Each strategy produces a different discovery design and a fundamentally different set of outputs. Selecting the wrong one does not produce imprecise results. It produces precise answers to irrelevant questions.

AI cannot make these choices, because the inputs that determine them are not in any customer dataset. They live inside the firm: in its capabilities, its competitive position, its risk appetite, and its strategic objectives. None of that is available to a model. So AI will make these choices anyway, implicitly, by defaulting to whatever framing appears most commonly in its training data for a given category. The output will look authoritative regardless of whether the underlying framing matches the firm’s actual situation.

The job hierarchy problem compounds this. In jobs-and-outcomes discovery, the level at which a job is defined is not a feature of the customer’s reality. It is a function of the firm’s objectives mapped onto that reality. Define the job at too high a level of abstraction and the innovation space becomes too broad to act on. Define it too narrowly and the real opportunity disappears from view. The correct level is determined by what the firm is trying to build and why. AI has no access to either input. It will default to whatever level of abstraction is most common in its training data for a given category, and nothing in the output reveals whether that default is strategically appropriate or not.

This is a different limitation than the data quality concerns discussed below. Better training data does not solve it. A more powerful model does not solve it. It cannot be solved by any model, because the required inputs are strategic judgments internal to the firm, and the output gives no signal when those judgments have been made incorrectly, or not made at all.

The Black Box Problem

There is a second risk. AI can produce job and outcome statements that look reasonable and actionable, and for well-established product categories in established markets, those outputs are often directionally useful. But the further you move from what already exists toward what customers need next, the less reliable that output becomes, and from AI outputs alone, you have no way to know where that line is.

No frontier model provider discloses what its training data contains. For well-established product categories, there is relevant customer information in that data, and Harvard Business School researchers found that AI outputs can approximate human survey responses for those categories.

But survey responses reflect the assumptions of whoever designed the survey, not the reality of what customers would say if asked what they are trying to accomplish and where they struggle. Approximating survey responses is not the same as producing valid jobs-and-outcomes knowledge. And for genuinely novel concepts, the same researchers found that accuracy dropped to near random chance.

Carnegie Mellon researchers found a further problem: LLMs conflate perspectives from different people in different situations into single responses. Even when the broad direction is right, the apparent specificity is manufactured. You cannot tell whose reality you are looking at.

How many CEOs want to make multi-million dollar innovation bets on information they cannot verify?

But There Is a Bigger Issue

Set those concerns aside for a moment. Even if AI could produce perfectly accurate customer insight, there is a more fundamental problem that rarely gets discussed.

Customer discovery is not just a data collection method. It is how teams learn.

Think about how AI fits into complex professional work. For routine tasks with clear inputs and known success criteria, AI can increasingly run entire tasks without human involvement. But when the task itself is determining what matters, humans are essential at both ends. Only humans can decide what questions to ask, which customers matter, and what direction to pursue. And only humans can interpret what the findings mean and decide what to do about them. AI is powerful in the middle, processing, synthesizing, accelerating. But when the question of what matters is still open, the judgment at both ends is human work.

Most people use the term “customer discovery” loosely. In Lean JTBD OS®, it has a precise meaning: the work of learning directly from customers what they are trying to accomplish, how they measure success, and where they struggle, so that the team can determine which problems are worth solving, and how to address them, before committing resources.

AI can contribute to this work. It can help conduct industry analysis, prepare for interviews, generate possible job map steps, analyze transcripts, identify patterns across conversations, and accelerate the synthesis of findings. What it cannot do is the core interaction: sitting with a real customer, asking the right questions, hearing what is not said, and building the understanding that only comes from doing the work firsthand. And AI does not have access to the strategic context, the risk appetite, or the leadership vision that determines which opportunities a firm chooses to pursue.

What Happens in the Conversation

In Lean JTBD OS discovery interviews, listening, analysis, and judgment happen at the same time. The interviewer is making real-time decisions about what to probe further, how to frame the next question for clarification, what the customer actually means, and what matters. That is a significant part of the analysis. It is not a step that can be separated from the conversation and handed to a machine. At least not yet.

Carnegie Mellon researchers also studied whether LLMs can replace human participants in qualitative research. They found that LLM-generated responses lack what they called “palpability and contextual depth,” the lived texture that only a real person brings to a conversation. The lead researcher concluded that there are “nuances that human participants contribute that you cannot possibly get out of LLM-based agents, no matter how good the technology is.”

In customer research, those nuances matter. They are the difference between a plausible summary of what customers might need and a firsthand account of what a specific customer is actually struggling to accomplish.

There is also something that gets lost when teams delegate this work to AI, even if AI performed flawlessly. Teams that conduct or observe discovery interviews together hear the hesitation in a customer’s voice. They watch their own assumptions break in real time. They walk out of those conversations with a shared understanding of what customers actually need, not a summary of what someone else concluded. That shared understanding is one of the most valuable outcomes of good discovery work. No report can replicate it, because the learning happens in the room.

Conducting Jobs-to-Be-Done Interviews Does Not Need to Be Onerous

Everything above might suggest that the alternative to AI-generated customer insight is a massive, expensive research program. Not true.

Three essential questions define the core customer insight every organization needs before making a confident innovation decision: WHO is the target customer? WHAT job are they trying to get done? WHERE do they struggle (what desired outcomes remain unmet)?

Notice what makes these three sufficient in some projects and not in others. The first is usually an internal strategic choice. And in some situations, the stakes of the struggle are already established before discovery begins, by the organization’s mandate, by prior evidence, or by the nature of the project. When that is true, interviews only need to carry what remains: the job and where customers struggle. But in a typical market bet, the stakes are precisely what you do not know. Then two more answers must come from customers themselves: WHY do they struggle, and WHAT are the consequences if the need remains unmet? Consequences are what separate a genuine unmet need from a mere annoyance, and they are the evidence an investment decision stands on.

This is what Minimally Viable Discovery (MVD) formalizes: a strategic objective frame plus five customer-story questions, with a clear division of labor. The frame carries what the firm already knows or must choose. Interviews carry what only customers can answer. The questions organize the inquiry; what makes MVD powerful is the discipline of capturing the answers as job and desired outcome statements, obtained through collaboration with customers, independent of any solution. A focused set of interviews, conducted by someone who knows what type of customer inputs to obtain and how to get them, can reveal jobs and outcomes (customers’ true needs) that no amount of operational data or AI analysis will surface.

In my work helping companies drive growth through innovation, I have seen teams achieve excellent results with a handful of well-run interviews and these few highly actionable customer inputs. Whether the minimum is sufficient for a given project depends on the organization’s business objectives and market situation. When the stakes are higher, the market is unfamiliar, or deeper precision is required, additional discovery domains are added accordingly. The discipline scales. Keep it as simple as possible, only as complex as necessary.

Lean JTBD OS is discovery infrastructure built on the language of Outcome-Driven Innovation® pioneered by Tony Ulwick at Strategyn, where I worked for seven years. I use ODI language with his permission for its precision and clarity.

Lean JTBD OS is designed to precede and complement methodologies like Design Thinking, Lean Startup, and Agile, not compete with them. Iteration is the right tool for refining solutions. JTBD is the right tool for revealing which problems are worth solving first.

The goal is not to give teams a better way to use AI. The goal is to give innovation teams and their leadership the theory, the language, and the tools they need to own the front end and back end of innovation themselves. To know which questions to ask before AI is involved, and to know what the answers mean after AI has done its work. That capability does not depend on any particular technology. It belongs to the team. And it becomes more valuable, not less, as AI gets more powerful, because the better AI gets at producing plausible outputs, the more critical it becomes that someone on the team can evaluate that output.

If your team doesn’t know the answers to those three questions from customers themselves, you’re guessing. And when you guess about what customers need, you risk building something they won’t buy.

Where This Leaves Us

“How should we use AI?” is a good question. AI can analyze operational data with extraordinary speed and precision and even generate good job and outcome statements. What it has not shown it can do is the human work at both ends: deciding what questions matter, sitting with a real customer to understand their reality, and interpreting what you learn. That is the work of discovery. That is where the learning lives. And that is how teams develop the judgment to know what is worth building.

AI has made it faster and easier to build than ever before. It has not made it faster to understand what is worth building. That understanding comes from doing the work. Speed without direction is not innovation. It is expensive motion.

Humans lead the discovery. AI accelerates the execution. Customers are the source of truth.

The leaders who see this will not abandon AI. They will build the discovery discipline underneath it, so that AI amplifies real understanding rather than accelerates assumptions. Not because their teams lack talent, but because the gap between operational knowledge and jobs-and-outcomes knowledge is structural. Most teams have never been shown the distinction. That is not a failure of effort or intent. It is an industry-wide gap, and closing it requires a discipline.

References

  1. Shaw, S. D., & Nave, G. (2026). Shaw, S. D., & Nave, G. (2026). Thinking—Fast, Slow, and Artificial: How AI is Reshaping Human Reasoning and the Rise of Cognitive Surrender. Working paper, The Wharton School, University of Pennsylvania. Available at SSRN, abstract 6097646.
  2. PwC (2023). Customer understanding perception gap research.
  3. Gartner (2022). Executive confidence versus customer perception study.
  4. Salesforce (2022). Business leader confidence versus customer assessment research.
  5. Kapania, S., Agnew, W., Eslami, M., Heidari, H., & Fox, S. (2025). “Simulacrum of Stories”: Examining Large Language Models as Qualitative Research Participants. Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). ACM.
  6. Brand, J., Israeli, A., & Ngwe, D. (2025). Using Gen AI for Early-Stage Market Research. Harvard Business Review.

Categories

Forbes Interviews Urko