A Brief History of AI: From the Turing Test to ChatGPT
Ask most people when artificial intelligence began, and they’ll say 2022 — the year ChatGPT launched. That’s a reasonable guess from the outside, but it’s wrong in a way that matters: the actual turning points in AI’s 70-year history are almost all invisible from where the public was standing at the time. Understanding what those turning points actually were — not the sanitized one-liners, but the specific technical failures and specific technical wins — is the fastest way to tell which parts of today’s AI hype are durable and which parts are marketing.
1950s: a test that was never meant to be a finish line
In 1950, Alan Turing published “Computing Machinery and Intelligence,” proposing what became known as the Turing Test: if a human judge, conversing by text, can’t reliably tell a machine from a person, the machine passes. It’s worth being blunt about this today — the Turing Test is a bad benchmark by modern standards. It measures conversational mimicry, not understanding, and a system can pass it through tricks (evasion, deflection, mimicking a non-native speaker to excuse errors) that have nothing to do with intelligence. Turing himself framed it as a replacement for the unanswerable question “can machines think?”, not as an engineering target — later generations of AI marketing flattened that nuance away.
The field’s founding moment came in 1956 at the Dartmouth Summer Research Project, where John McCarthy coined “artificial intelligence” and the organizers proposed — in writing — that “every aspect of learning… can in principle be so precisely described that a machine can be made to simulate it,” and that meaningful progress could happen in a single summer. It didn’t. But the framing stuck, and the first fifteen years of AI research were built almost entirely on symbolic reasoning: representing knowledge as logical propositions and searching through them — systems like Newell and Simon’s Logic Theorist (1956), which proved mathematical theorems, and later SHRDLU (1970), which could follow natural-language instructions inside a simulated world of toy blocks. These systems worked precisely because their worlds were small and fully specified. The moment you needed common sense, ambiguity tolerance, or knowledge that hadn’t been explicitly hand-coded, symbolic AI had no mechanism to cope.
1970s: the first AI winter had a specific cause
By the mid-1970s, two failures made the gap between promise and delivery impossible to paper over. Machine translation — heavily funded during the Cold War for translating Russian scientific text — turned out to require real-world knowledge and disambiguation that rule-based systems couldn’t provide; the 1966 ALPAC report concluded it wasn’t worth continued funding, and US government money for the field largely dried up. In the UK, the 1973 Lighthill Report reached a similarly damning verdict on AI’s progress relative to its promises, and British funding collapsed too. This is the first AI winter: not vague disillusionment, but two specific government reports killing two specific national funding streams after two specific technical failures.
1980s: expert systems, and a winter with a named mechanism
AI’s first real commercial wave came through expert systems — programs like MYCIN (medical diagnosis) and, most successfully, XCON (also called R1) at Digital Equipment Corporation, which configured VAX computer orders and reportedly saved DEC on the order of tens of millions of dollars a year in the mid-1980s by catching configuration errors humans missed. This was real, deployed value — not vaporware. Corporations spent heavily replicating it.
What killed the wave had a specific name researchers still use: the knowledge-acquisition bottleneck. Every fact an expert system knew had to be manually elicited from a human expert and hand-encoded by a “knowledge engineer” as an if-then rule. This didn’t scale — doubling a system’s coverage roughly doubled its maintenance burden, rules interacted with each other in ways that got harder to predict as the rule base grew, and the systems were completely brittle outside their exact domain. By the late 1980s the specialized Lisp-machine hardware industry built around this wave collapsed commercially, and the second AI winter set in, lasting into the mid-1990s.
1990s–2000s: the unglamorous approach that actually worked
While symbolic AI stalled, statistical machine learning quietly took over the parts of AI that actually shipped: spam filters, fraud detection, and early web search ranking, using methods like support vector machines and, later, random forests and gradient-boosted trees — algorithms that learn a decision boundary from labeled examples rather than from hand-written rules. None of this generated headlines. All of it worked, and worked reliably, in production, at scale — which is exactly why it displaced symbolic AI in industry even before deep learning existed.
Meanwhile, the specific mathematical machinery for today’s neural networks — backpropagation, formalized by Rumelhart, Hinton, and Williams in a 1986 paper — had already existed for over two decades before it became relevant. Having the right algorithm was never the bottleneck. Having enough labeled data and enough compute to make that algorithm worth running was.
2012: a benchmark result that was hard to argue with
In 2012, Krizhevsky, Sutskever, and Hinton entered a convolutional neural network called AlexNet into the ImageNet Large Scale Visual Recognition Challenge. The result wasn’t a marginal win: AlexNet’s top-5 error rate was 15.3%, against 26.2% for the second-place entry — which used the field’s best hand-engineered computer vision features. An eleven-point gap on a mature, fiercely contested benchmark almost never happens; it’s what ended the “is deep learning actually better” debate in computer vision within about a year.
What made it possible wasn’t a new algorithm — backpropagation and convolutional networks were both already known. It was three converging resources: ImageNet’s 1.2 million labeled training images (a scale no one had trained on before), two consumer GTX 580 GPUs repurposed to run the network’s roughly 60 million parameters in parallel, and training refinements (ReLU activations instead of sigmoid/tanh, and dropout regularization) that made networks this deep actually trainable without collapsing. This is the pattern worth remembering: the theory was old, the resources were new, and the resources are what mattered.
2017: the architectural detail that actually explains the win
Google researchers’ 2017 paper “Attention Is All You Need” introduced the transformer, and it’s worth being precise about what its “attention” mechanism actually does, because the popular explanation (“it weighs the importance of words”) skips the part that matters. For every token, the model computes a query vector, and compares it — via dot product — against key vectors from every other token in the sequence, producing a set of relevance weights that are then used to compute a weighted sum of value vectors. That’s the whole mechanism: query, key, value, weighted sum, repeated across multiple attention “heads” in parallel.
Here’s the part that actually explains why transformers displaced the previous dominant architecture (recurrent neural networks, or RNNs): an RNN processes a sequence step by step, where the computation at position t depends on the completed computation at position t-1. That dependency chain makes it structurally impossible to parallelize an RNN’s core computation across the sequence during training — you’re stuck waiting on the previous step no matter how much GPU compute you throw at it. The transformer’s attention computation, by contrast, is one big matrix multiplication across the whole sequence at once, which GPUs are extremely good at parallelizing. The transformer isn’t obviously “smarter” per token than an RNN — it wins because it lets you convert raw compute directly into training speed at a scale RNNs structurally cannot match.
This connects to an argument researcher Rich Sutton made in a widely-cited 2019 essay called “The Bitter Lesson”: across AI’s history, methods that scale efficiently with more compute and data have consistently beaten methods built on more clever, human-crafted structure, once enough compute became available. The transformer is close to a textbook case. Nearly every large language model since 2017 — GPT, Gemini, Claude, Llama — runs on some variant of this same architecture, not because nobody has tried alternatives, but because none have yet beaten its combination of quality and parallelizable training efficiency at scale.
2022: the specific mechanism behind “suddenly working,” not just packaging
GPT-3 was publicly documented by OpenAI in 2020 — two years before ChatGPT. It’s tempting to say the difference in 2022 was purely a friendlier chat interface, but that undersells a real technical step in between: RLHF, reinforcement learning from human feedback. Base language models like GPT-3 are trained to predict the statistically likely next token, which makes them fluent but not necessarily helpful, honest, or safe to talk to directly — they’ll confidently continue a prompt in whatever direction is statistically probable, including wrong or harmful directions. OpenAI’s InstructGPT work (2022) fine-tuned the base model using human rankings of which responses were actually good, training a reward model and then optimizing the language model against it. ChatGPT is downstream of that specific technique, not just a UI wrapped around GPT-3. The interface mattered for adoption; RLHF is what made the thing behind the interface usable for ordinary conversation in the first place.
What’s actually different about AI today — and where the evidence gets shaky
The honest, unhyped version is: mostly scale, not a new kind of intelligence. Larger models trained on more data with more compute reliably do things smaller versions of the same architecture couldn’t — this is well established. Where the story gets more contested is the specific claim that these gains appear as sudden “emergent capabilities” — skills that seem to snap into existence once a model crosses some size threshold, rather than improving gradually. This claim has real pushback in the research literature: a widely discussed 2023 paper by Schaeffer, Miller, and Koyejo argued that at least some measured “emergence” is a mirage created by the choice of scoring metric — switching from an all-or-nothing accuracy score to a smoother, partial-credit metric on the same models and the same tasks made several supposedly “emergent” jumps disappear into smooth, predictable curves. The honest state of the field right now is that some genuine emergent behavior likely exists, some of it is a measurement artifact, and researchers don’t yet agree on the ratio — which is a more useful thing to know than the flattened headline version of either claim.
That unresolved argument about what “AI” is actually doing under the hood is also the right note to end this piece on, because the next one in this series covers something people conflate even more than they conflate expert systems with intelligence: the difference between machine learning, deep learning, and large language models — three nested categories, not three names for the same thing.
Is ChatGPT the first real AI?
No — AI research goes back to the 1950s. What's specific to ChatGPT is that it's the first large language model wrapped in RLHF (reinforcement learning from human feedback) tuning and a free chat interface, so it behaved helpfully enough for ordinary people to use directly. GPT-3, the base model underneath that lineage, had already existed since 2020 as a research artifact.
What was an AI winter, in concrete terms?
A period where AI research funding collapsed because a specific technical approach hit a wall it couldn't get past — not a vague loss of interest. The first winter (mid-1970s) followed the failure of machine translation projects and the 1966 ALPAC report that killed US government funding for it. The second (late 1980s–mid-1990s) followed the collapse of the commercial expert-systems market once companies realized the maintenance cost of hand-coding rules didn't scale.
Was AlexNet really that much better, or is this exaggerated in retrospect?
It's not exaggerated — the gap was unusually large for a machine learning benchmark. AlexNet's top-5 error rate on ImageNet 2012 was 15.3%; the second-place entry, using traditional hand-engineered computer vision features, scored 26.2%. An 11-point gap in a mature, heavily-contested benchmark is a genuinely rare result, which is why it ended the debate over deep learning almost overnight within the computer vision field.