Writing archive

Aug 3, 2026 · Google

Google Had the Chatbot. It Evaluated It Like Search.

Google had conversational AI before ChatGPT, but judged it against Search's standards, economics, and risks. The Innovator's Dilemma explains why that was rational—and why it delayed the product people actually wanted.

Dark teal layered arcs hold copper bands as a small cobalt plane opens into a pale aqua field

Google did not miss conversational AI because it lacked the technology.

It missed the first mass-market moment because it understood the technology through the product it already had.

That distinction is the heart of Clayton Christensen's The Innovator's Dilemma. Successful companies often fail to lead a disruption not because they are badly managed, but because they apply the standards, customers, economics, and risk controls that made the existing business successful.

On August 2, 2026, Cheng Lou pointed to a revealing admission from Google Chief Scientist Jeff Dean: Google had an internal chatbot before ChatGPT, but did not initially see enough value beyond what people could get from Search. Tibo Sottiaux, a former DeepMind researcher, replied with a firsthand recollection: he had worked on a system called LMChat roughly a year before ChatGPT, but Google was too nervous to release it and DeepMind was prevented from shipping products that could disrupt Google.

An X exchange in which Cheng Lou recalls Google's pre-ChatGPT internal bot and Tibo Sottiaux says he worked on the LMChat team

Sottiaux's claim about the codename and internal product constraints is a public firsthand account, not a company-confirmed history. But the broader story does not depend on that post alone. Google's own publications and Dean's own words establish the central facts.

Google had built the technology. It had internal users. It knew the interface was compelling. What it did not have was an organizational context in which the product could be judged on its own terms.

Google Was Early

Google's lead was not hidden in hindsight.

In January 2020, Google Research introduced Meena, a 2.6-billion-parameter open-domain chatbot designed to converse about almost anything. The team explicitly declined to release an external demo because of safety and bias concerns.

In May 2021, Google announced LaMDA, describing free-flowing conversation as a path to “entirely new categories of helpful applications.” By January 2022, its researchers had published work on making LaMDA safer, more factual, and higher quality. In August 2022, Google began giving small groups of US users access to LaMDA through AI Test Kitchen.

OpenAI launched ChatGPT on November 30, 2022, explicitly as a research preview intended to collect feedback about its strengths and weaknesses. Google opened early access to Bard on March 21, 2023.

The gap was not years of missing research. It was a few consequential months between a constrained experiment and a general-purpose product that anyone could try.

In a long-form interview with Dwarkesh Patel, Dean explained the decision more directly. Google had an internal version of Meena before ChatGPT. During the pandemic, employees used it as a kind of lunch companion. Yet the company viewed the system from a Search perspective: a search product should return factual information with extremely high reliability, while language models hallucinated, made mistakes, and could produce offensive content.

Dean then identified what Google had underestimated. The chatbot was useful for tasks people would never give a search engine: drafting a note to a veterinarian, summarizing text, or creating new material. Users did not adopt chatbots merely as a different route to the same web pages. They used them to do different work.

That is the critical admission.

Google evaluated a new market through the performance criteria of the old one.

This Is What the Innovator's Dilemma Looks Like

“Disruption” is often used as a synonym for any dramatic technological change. Christensen's theory is narrower and more useful.

The Christensen Institute's definition describes a process in which a product begins in simpler, cheaper, or more accessible applications that established companies have little incentive to serve, then improves and moves upmarket. The disruptive product is often worse on the dimensions the incumbent's best customers value. Its advantage appears along different dimensions for different users or jobs.

Early chatbots were plainly worse than Google Search at many search tasks. They hallucinated. They rarely showed sources. Their answers were difficult to audit. They could produce toxic or embarrassing responses. If the question was “Can this replace Google Search without degrading Google's promise of reliable information?”, caution was the correct answer.

But that was the wrong market test.

The better questions were:

  • Can this help someone begin from a blank page?
  • Can it transform, summarize, explain, or draft instead of merely retrieve?
  • Can conversation make computing accessible to people who do not know the right query, syntax, or software workflow?
  • Can a public release generate the usage data needed to discover applications the company cannot predict internally?

On those dimensions, the chatbot did not need to outperform Search. It needed to make a different set of jobs possible.

The product looked inadequate when placed on Search's performance curve and extraordinary when placed on its own.

That is why the Innovator's Dilemma is difficult. The incumbent's mistake is usually defensible in the language of quality.

Search Was More Than a Product

Google's decision cannot be separated from the economics around it.

Alphabet's 2022 annual report recorded $162.45 billion of revenue from “Google Search & other.” Search had a mature advertising system, an enormous distribution surface, well-understood user behavior, and a reputation built around retrieving dependable information from the web.

A general-purpose chatbot challenged every part of that system at once.

It changed the interface from a list of links to a generated answer. It made the placement and economics of advertising uncertain. It concentrated responsibility for an answer inside Google's product rather than distributing it across sources. It increased reputational exposure when the system was wrong, offensive, or confidently strange. And it risked teaching users to begin important tasks somewhere other than the search box.

Sebastian Mallaby recently described this on NPR's Planet Money as a “triple innovator's dilemma”: the chatbot threatened Google's reliability promise, its advertising model, and its politically exposed market position.

None of those concerns were imaginary.

That is precisely why they were powerful enough to delay the product.

Inside a startup, an unreliable chatbot with unclear monetization can be a promising wedge. Inside Google, the same artifact can look like a lower-quality version of a hugely profitable service carrying unacceptable downside.

The technology did not change between those two descriptions.

The organization around it did.

The Missing Capability Was Permission

Large companies often respond to this problem by creating a research lab, incubator, or innovation team. Google had several of the best in the world. That solved the problem of invention.

It did not solve the problem of permission.

A disruptive unit needs more than technical independence. It needs permission to:

  • serve users the core business does not yet understand;
  • ship against a different quality curve;
  • use staged access and explicit limitations rather than inherit the incumbent's release bar;
  • discover a business model instead of proving the incumbent model in advance;
  • and cannibalize existing behavior before a competitor does it.

Without those permissions, the new unit can produce breakthroughs but still has to return to the incumbent for a launch decision. The invention is separate. The incentives are not.

Sottiaux's claim that DeepMind was blocked from shipping products that could disrupt Google is therefore more than an anecdote about corporate nervousness. If accurate, it describes the exact boundary at which an innovation lab stops being structurally independent: it may explore the future so long as the future does not threaten the present.

Dario Amodei's Similar Story Is Actually Different

Dario Amodei has told a superficially similar story about Anthropic.

In the summer of 2022, Anthropic had an early version of Claude before ChatGPT's release. According to TIME's account based on interviews with Amodei, the company chose not to release it because he feared triggering a race before Anthropic had done enough safety work. He later said he suspected that was the right decision, while acknowledging that it was not clear-cut.

The chronology is similar. The mechanism is not.

Anthropic did not have a dominant search product or a $162 billion search-advertising line to protect. Withholding Claude may have cost it the first-mover advantage, but it was not preserving an incumbent product from cannibalization. It was a safety judgment by a young lab whose stated strategy depended on developing more controllable systems.

This distinction matters because not every delayed release is evidence of the Innovator's Dilemma. Safety can be substantive rather than a cover for inertia. Google's own concerns about factuality and harmful output were also real.

The diagnostic question is not simply, “Who had a model first?”

It is, “Which existing customers, metrics, margins, and liabilities determined what the model was allowed to become?”

For Google, Search supplied the answer.

For Anthropic, the stated constraint was safety.

Both decisions carried an opportunity cost. Only one was made inside the gravitational field of a dominant incumbent business.

OpenAI Changed the Evaluation Environment

OpenAI's decisive product move was not only building a capable model. It was creating a public environment in which the model's value could be discovered despite its defects.

The launch post did not claim ChatGPT was ready to replace a search engine, professional writer, tutor, or programmer. It called the product a research preview and invited users to expose its strengths and weaknesses.

That framing converted imperfection from a reason not to ship into a reason to learn.

Once people began using the system, the market supplied evidence no internal benchmark could produce. Users demonstrated new jobs, developed prompting practices, exposed failure modes, changed their habits, and created demand for capabilities that had previously looked like demos.

Google's internal use had already provided a clue. Employees chatted with Meena because it was pleasant and useful in a way Search was not. But internal enthusiasm could still be treated as an experiment. A public product made the behavior strategically legible.

ChatGPT did not merely beat Google to launch.

It changed the evidence Google could use to justify launching.

How an Incumbent Should Respond

The lesson is not that incumbents should release unsafe products or ignore the quality standards their customers trust.

It is that a disruptive product cannot be required to satisfy the full operating model of the business it may replace before it is allowed to find its own market.

An incumbent facing this problem needs a separate decision system:

1. Define the new job before comparing products

Do not ask whether the new product is better than the incumbent in aggregate. Identify what people can now accomplish that the incumbent was not designed to do.

2. Give the new product its own metrics

Measure learning, task completion, creation, iteration, and new-user adoption—not only factual retrieval, incumbent revenue, or migration from the existing product.

3. Make cannibalization an explicit mandate

If the new unit needs permission from the business it may disrupt, the incumbent has not created structural independence. Leadership has only moved the conflict into a later meeting.

4. Separate a safety floor from an incumbent quality ceiling

Some behaviors must block release. Others can be managed through limited access, clear product framing, citations, human review, monitoring, and rapid iteration. A new product needs a defensible safety threshold, not automatic inheritance of every expectation attached to a mature one.

5. Learn in the market before the market is obvious

Disruptive use cases often look small, unserious, or unprofitable at first. By the time they satisfy the incumbent's planning model, a new entrant may already own the user habit and feedback loop.

The Wrong Question Was Rational

Google has not disappeared. It responded, shipped Bard, reorganized its AI efforts, and built Gemini into a major competitor. The Innovator's Dilemma predicts a difficult competitive response; it does not guarantee extinction.

But the lost moment still matters.

Google had years of conversational research, an internal chatbot people enjoyed, LaMDA, enormous compute, global distribution, and many of the researchers who created the foundations of modern language models.

The missing resource was not intelligence.

It was a way to value a product that looked worse when judged as Search and better when judged as something new.

That is the question leaders should carry from this episode.

When a new system appears inside a successful company, asking “Is it better than our current product?” feels rigorous.

It may also be the mechanism that protects the current product from the future.

The more useful question is:

What becomes valuable if we stop requiring the future to look like the business we already understand?

Sources