Artificial intelligence turned 70 in 2026, counting from the summer of 1956, when researchers met at Dartmouth College to ask whether machines could learn. The idea is far older than that meeting, because people imagined thinking machines in myths long before the first computer existed. Medieval engineers built programmable automata, while philosophers wrote down the rules of logic that computers still follow today. This guide traces the full history of artificial intelligence, from those early ideas to the generative AI era that ChatGPT opened. Along the way, it explains why each era rose and fell, which is the part of the evolution of artificial intelligence that most timelines leave out.
Quick answer: The history of AI as a scientific field began at the 1956 Dartmouth workshop, whose 1955 proposal by John McCarthy gave the discipline its name. Alan Turing laid the theoretical groundwork in 1950, and symbolic AI led the field until the late 1980s. Statistical machine learning took over in the 1990s, deep learning broke through in 2012, and generative AI reached the mainstream in November 2022.
Key Takeaways
- The term “artificial intelligence” first appeared in the August 1955 proposal for the 1956 Dartmouth workshop, written by John McCarthy, Marvin Minsky, Nathaniel Rochester and Claude Shannon.
- No single person invented AI: Turing framed the question, McCarthy named the field, while Allen Newell and Herbert Simon built Logic Theorist, widely called the first AI program.
- AI survived two funding collapses, called AI winters, in roughly 1974 to 1980 and 1987 to 1993, each following promises that the computers of the day could not keep.
- Every revival came from a new approach: expert systems in the 1980s, statistical learning in the 1990s, deep learning after 2012 and transformer models after 2017.
- Every AI system in use today, including ChatGPT, Gemini and Claude, is still narrow AI, since artificial general intelligence remains a research goal.
What Is Artificial Intelligence?
Artificial intelligence is the branch of computer science that builds machines able to perform tasks that normally require human intelligence. Those tasks include reasoning, learning, understanding language, recognizing images and making decisions. McCarthy later described the field as “the science and engineering of making intelligent machines,” and researchers have pursued that goal with very different methods. As a result, the AI timeline reads less like a straight line than a series of rival schools taking turns at the front.
Two schools shaped most of that story over seven decades. Symbolic AI tried to encode intelligence as explicit rules that a programmer could read and check. Connectionism, by contrast, tried to grow intelligence from networks of simple artificial neurons that learn from data, and most of today’s systems belong to this second camp. For a full definition of the field and its branches, see our guide to what artificial intelligence is.
The Origins of Artificial Intelligence: Thinking Machines Before Computers
The origins of artificial intelligence stretch back more than two thousand years, long before electronic computers existed. Early thinkers asked two questions that AI still tries to answer: can a machine act on its own, and can reasoning be reduced to mechanical steps?
Ancient Myths, Automata and Artificial Beings
Greek mythology already featured artificial beings such as the mechanical servants of Hephaestus, the god of metalwork. Another myth described Talos, a bronze giant who guarded the island of Crete by circling it three times a day. Hero of Alexandria turned some of these stories into real devices in the first century CE, describing automata driven by water, air pressure and falling weights. In 1206, the engineer Ismail al-Jazari completed The Book of Knowledge of Ingenious Mechanical Devices, which documented programmable water clocks and musical automata.

One of al-Jazari’s musical machines let its operator change the drum patterns by moving pegs on a rotating cylinder. These devices did not think, yet they proved that a machine could follow a stored sequence of instructions without a human guiding each move. That idea of programmable behavior sits at the root of every computer program written since.
Formal Logic and Mathematical Reasoning
The second thread of AI’s origins came from logic and mathematics rather than engineering. Aristotle’s syllogisms, written in the fourth century BCE, set out rules for drawing valid conclusions from premises. Euclid’s Elements showed how a whole body of knowledge could be derived step by step from a few axioms. In the ninth century, the Persian mathematician Muhammad ibn Musa al-Khwarizmi wrote systematic methods for solving equations, and the word “algorithm” comes from the Latin form of his name.
Later European philosophers pushed the idea of mechanical reasoning further. In the late 1200s and early 1300s, Ramon Llull designed paper wheels that combined concepts to generate new statements. René Descartes argued in the 1600s that animal bodies worked like machines, although he doubted any machine could use language flexibly. Gottfried Leibniz went furthest by building a mechanical calculator in the 1670s. He also dreamed of a universal symbolic language, the characteristica universalis, paired with a calculus ratiocinator that could settle disputes by calculation.
Babbage, Lovelace and Mechanical Calculation
The nineteenth century turned mechanical calculation into serious engineering. Charles Babbage designed the Analytical Engine in the 1830s, a general-purpose mechanical computer with separate memory and processing units, although he never finished building it. Ada Lovelace translated and annotated a paper on the engine in 1843, noting that it could manipulate symbols of any kind, including musical notes. She also wrote that the engine “has no pretensions whatever to originate anything,” a doubt Alan Turing would answer directly a century later.
In 1854, George Boole published The Laws of Thought, which reduced logical reasoning to an algebra with two values, true and false. Boolean logic became the language of digital circuits, so every AI system today still runs on Boole’s arithmetic at its lowest level.
The Foundations of Modern AI (1936-1955)
The development of artificial intelligence rests on a burst of work in the 1930s and 1940s. That work defined what a computer is, how information can be measured and how a network of artificial neurons might compute.
Alan Turing, the Turing Machine and the Turing Test
Alan Turing, a British mathematician, described an abstract “universal machine” in 1936 that could carry out any calculation expressible as rules. That concept, now called the Turing machine, became the theoretical model for every programmable computer. During the Second World War, Turing helped break German codes at Bletchley Park, where he saw how machines could speed up reasoning tasks.

In 1950, Turing published “Computing Machinery and Intelligence” in the journal Mind. The paper opens with the question “Can machines think?” before replacing it with a practical test called the imitation game. In that game, a judge exchanges typed messages with a hidden human and a hidden machine, and the machine passes if the judge cannot reliably tell them apart. The Turing Test gave early researchers a concrete target, while the paper also answered objections such as Lovelace’s claim that machines cannot originate anything.
Cybernetics, Information Theory and Electronic Computers
Several neighboring fields fed into AI during the same decade. Norbert Wiener’s Cybernetics (1948) studied feedback and control in both animals and machines. Claude Shannon’s 1948 paper on information theory showed how information could be measured in bits, and his 1950 paper described how a computer might play chess. Meanwhile, the first stored-program computers, including the Manchester Baby in 1948 and the Ferranti Mark 1 in 1951, gave researchers real hardware to test their ideas.

McCulloch, Pitts and the First Artificial Neurons
The history of neural networks starts in 1943, when neurophysiologist Warren McCulloch and logician Walter Pitts proposed a mathematical model of a neuron. Their model neuron fires when its inputs cross a threshold, and they showed that networks of such units could compute logical functions. In 1949, psychologist Donald Hebb suggested that connections between neurons strengthen when they fire together, which offered an early theory of learning. Two years later, Marvin Minsky and Dean Edmonds built SNARC, a machine with 40 artificial neurons that simulated a rat learning a maze.
The Birth of Artificial Intelligence: The 1956 Dartmouth Workshop
Most historians date the birth of artificial intelligence as an academic field to the summer of 1956, when researchers gathered at Dartmouth College in Hanover, New Hampshire.
What Happened at the Dartmouth Workshop?
In August 1955, John McCarthy, then a young mathematics professor at Dartmouth, wrote a funding proposal with three senior partners. They were Marvin Minsky of Harvard, Nathaniel Rochester of IBM and Claude Shannon of Bell Labs. The Dartmouth proposal rested on a daring conjecture that “every aspect of learning or any other feature of intelligence can in principle be so precisely described that a machine can be made to simulate it.” The Rockefeller Foundation funded the project, and the workshop ran for about eight weeks.
The meeting produced no single breakthrough, since participants came and went instead of working as one team. Its real effect was social, because it gave scattered researchers a shared name, a shared agenda and a community. Its attendees went on to lead the major AI laboratories at MIT, Carnegie Mellon and Stanford, which is why 1956 counts as the year AI became a discipline.
Who Coined the Term Artificial Intelligence?
John McCarthy coined the term “artificial intelligence” in the 1955 Dartmouth proposal. He chose the phrase partly to separate the new field from cybernetics and automata theory, which he considered too narrow. McCarthy created the LISP programming language in 1958, which served as the standard language of AI research for three decades. He also founded the Stanford AI Laboratory in 1963, so many people call him the father of artificial intelligence.

Logic Theorist: The First AI Program
Allen Newell, Herbert Simon and programmer Cliff Shaw brought a working program to Dartmouth called Logic Theorist. They built it in 1955 and 1956 at the RAND Corporation and the Carnegie Institute of Technology, now Carnegie Mellon University. Logic Theorist proved 38 of the first 52 theorems in chapter two of Whitehead and Russell’s Principia Mathematica, and one of its proofs was more elegant than the original. It relied on heuristic search, trying promising paths first instead of checking every possibility, an idea that shaped AI for decades.
Other programs have a claim to being first in narrower senses. Christopher Strachey wrote a checkers program that ran on the Ferranti Mark 1 in 1952, and Arthur Samuel began his own checkers work at IBM around the same time. Logic Theorist, however, was the first program built specifically to imitate human problem solving, which is why most histories give it the title.
The Golden Years of Early AI (1956-1974)
The two decades after Dartmouth brought fast progress and even faster predictions. Government money, mostly from the US Advanced Research Projects Agency (ARPA, later DARPA), flowed to MIT, Stanford, Carnegie Mellon and the Stanford Research Institute with few conditions attached.
Symbolic AI, Search and Problem Solving
Early AI research treated intelligence as the manipulation of symbols according to rules. Newell and Simon followed Logic Theorist with the General Problem Solver in 1957, which aimed to solve any problem expressed as goals and allowed moves. In 1961, James Slagle’s SAINT program at MIT solved calculus integration problems at the level of a first-year college student. Researchers also built semantic networks, which stored knowledge as linked concepts, along with planning systems that worked out sequences of actions.
A different response to uncertainty appeared in 1965, when Lotfi Zadeh at UC Berkeley introduced fuzzy logic. Fuzzy logic let programs reason with degrees of truth rather than strict true or false values, and engineers later used it in washing machines, cameras and train controls.
Game Playing and the Beginnings of Machine Learning
Games made ideal test beds for early AI because the rules were clear and success was easy to measure. Arthur Samuel’s checkers program at IBM improved by playing against itself, marking the start of the history of machine learning. In 1959, Samuel popularized the term “machine learning” to describe that ability. His program eventually beat a respectable amateur, which impressed audiences because nobody had programmed the winning strategy directly.
Natural Language Processing: ELIZA and SHRDLU
Early work on natural language processing produced some of the most memorable programs of the era. Joseph Weizenbaum’s ELIZA, written at MIT between 1964 and 1966, imitated a psychotherapist by turning the user’s statements into questions. ELIZA understood nothing, yet many users confided in it, which disturbed Weizenbaum enough that he became a lasting critic of AI. Terry Winograd’s SHRDLU (1968 to 1970) followed typed commands to move blocks in a simulated world. It impressed visitors, although it only worked inside that tiny “microworld.”

Robotics and Computer Vision: Shakey the Robot
The Stanford Research Institute (later SRI International) built Shakey between 1966 and 1972, the first mobile robot that could reason about its own actions. Shakey combined a camera, a range finder and a planning program to push boxes between rooms. Its research also produced the A* search algorithm, which navigation software still uses. In Japan, Waseda University completed WABOT-1 in 1973, a full-scale humanoid robot that could walk and grip objects.

The Perceptron and Early Neural Networks
Frank Rosenblatt, a psychologist at Cornell, introduced the perceptron in 1958, a learning algorithm for a single layer of artificial neurons. His Mark I Perceptron machine learned to recognize simple shapes from camera images. Newspapers reported that the Navy expected the machine to walk, talk and see, so the hype ran well ahead of the hardware.

Optimism spread across the whole field during this period. Herbert Simon predicted in 1965 that “machines will be capable, within twenty years, of doing any work a man can do.” Predictions like this set expectations that the technology of the day could not meet.
The First AI Winter (1974-1980)
The first AI winter was a period of sharp funding cuts that lasted from roughly 1974 to 1980. Early AI programs worked on toy problems but failed on real ones, and funders on both sides of the Atlantic lost patience at almost the same time.
Why AI Funding Declined
Three reports and policy changes did most of the damage to AI funding and research:
- The ALPAC report (1966): A US government committee found machine translation slower, less accurate and more expensive than human translation. The ALPAC report ended most US funding for translation research.
- The Mansfield Amendment (1969): Congress required military research funding to serve a direct defense mission, which pushed DARPA away from open-ended research. In 1974, DARPA cancelled a major speech understanding program at Carnegie Mellon after it missed its goals.
- The Lighthill report (1973): Mathematician Sir James Lighthill told the UK Science Research Council that AI had failed to achieve its “grandiose objectives.” His report led Britain to cut AI funding at most universities except Edinburgh, Essex and Sussex.
Neural network research suffered a separate blow in 1969. Marvin Minsky and Seymour Papert published Perceptrons, which proved that a single-layer perceptron could not learn simple functions such as XOR. Many readers took the book as proof that neural networks were a dead end, so connectionist funding nearly vanished for over a decade.
The Technical Limits of Early AI
The first AI winter was not only a political event, because early AI hit real technical walls:
- Limited computer power: Computers of the 1970s had a tiny fraction of the memory and speed that language or vision tasks required.
- Combinatorial explosion: Search programs that solved small puzzles ground to a halt as the number of possible moves grew exponentially.
- Commonsense knowledge: Programs knew nothing about the everyday world, so they could not follow a simple children’s story.
- Moravec’s paradox: Tasks humans find hard, such as algebra, proved easy for computers, while effortless human skills like walking proved extremely hard. Roboticist Hans Moravec described this pattern in 1988.
The AI Boom of the 1980s: Expert Systems
AI returned in the 1980s through a narrower, more commercial idea. Instead of chasing general intelligence, researchers captured a human specialist’s knowledge in a single domain and packaged it as software. These programs were called expert systems, and for a few years they turned AI into a business.
Knowledge Representation and Knowledge Engineering
An expert system has two main parts: a knowledge base of if-then rules gathered from human experts, plus an inference engine that applies those rules. Knowledge engineers spent months interviewing specialists to turn their judgment into rules a computer could follow, and three systems defined the approach:
- DENDRAL (from 1965): Built at Stanford by Edward Feigenbaum, Joshua Lederberg and colleagues, it identified chemical compounds from mass spectrometry data.
- MYCIN (early 1970s): Also from Stanford, it recommended antibiotics for blood infections and matched specialists in tests, although hospitals never adopted it.
- XCON (1980): Carnegie Mellon built it for Digital Equipment Corporation (DEC) to configure computer orders, and it reportedly saved DEC tens of millions of dollars a year.

Success at DEC convinced large companies to set up their own AI groups. A new industry of AI software firms and specialized hardware makers quickly grew around them, fueling the AI boom of the early 1980s.

Japan’s Fifth Generation Computer Project
In 1982, Japan’s Ministry of International Trade and Industry launched the Fifth Generation Computer Systems project. This ten-year plan aimed to build massively parallel computers that reasoned with the logic programming language Prolog. The project alarmed Western governments, which answered with programs of their own. The United States started the Strategic Computing Initiative in 1983, and the United Kingdom launched the Alvey Programme that same year.
The Neural Network Revival: Hopfield, Backpropagation and Connectionism
Neural networks came back during the same decade that expert systems peaked. In 1982, physicist John Hopfield showed that a certain type of network could store and recall memories, now called the Hopfield network. Geoffrey Hinton and Terry Sejnowski followed with the Boltzmann machine in 1985. In Japan, Kunihiko Fukushima had already built the Neocognitron in 1980, a layered vision network that inspired later image models.
The largest step came in 1986, when David Rumelhart, Geoffrey Hinton and Ronald Williams published a paper in Nature on backpropagation. The method trained networks with several layers, solving the XOR problem that Perceptrons had highlighted. Our explainer on how backpropagation adjusts a network’s weights walks through the process with real numbers. In 1989, Yann LeCun at Bell Labs used backpropagation to train a convolutional network that read handwritten zip codes.
Probabilistic Reasoning and Bayesian Networks
Expert systems handled uncertainty poorly, because real-world evidence is rarely black and white. In 1988, Judea Pearl published Probabilistic Reasoning in Intelligent Systems, which introduced Bayesian networks as a way to reason with probabilities. This work pulled AI toward mathematics and statistics, and it earned Pearl the Turing Award in 2011.
The Second AI Winter (1987-1993)
The second AI winter began in 1987, when the market for specialized AI hardware collapsed, and it lasted into the early 1990s. The expert-system boom ended for reasons that echo the first winter, since vendors sold the technology as more general and durable than it was.
AI Hype Meets Commercial Reality
The hardware market failed first, starting with companies such as Symbolics and Lisp Machines Inc., which had sold expensive workstations built to run LISP. In 1987, cheaper desktop computers from Apple and IBM matched their power, and that market collapsed almost overnight.
Expert systems also turned out to be brittle, failing badly on any case their rules did not cover. They were costly to maintain as well, because every change in a business meant writing more rules by hand. Government support faded at the same time, as DARPA scaled back the Strategic Computing Initiative and Japan ended its Fifth Generation project in 1992.
Critics inside the field pushed in new directions. At MIT, Rodney Brooks argued in the late 1980s that robots should react to the world directly instead of planning from symbolic models, and his idea later shaped the Roomba vacuum. Meanwhile, the label “AI” grew so unfashionable that many researchers renamed their work informatics, knowledge-based systems or computational intelligence.
AI in the 1990s and 2000s: The Statistical Turn
AI recovered in the 1990s by becoming less ambitious in public and more rigorous in private. Researchers stopped writing rules by hand and started building systems that learned patterns from data, borrowing methods from statistics and optimization. Much of this work shipped inside ordinary products without the AI label.
Machine Learning Replaces Hand-Written Rules
The 1990s became a turning point in the history of machine learning, as statistical methods took over speech recognition, spam filtering and data mining. Support vector machines, introduced by Corinna Cortes and Vladimir Vapnik in 1995, became a standard classification tool. Gerald Tesauro’s TD-Gammon reached near world-champion backgammon in 1992 by playing against itself. Richard Sutton and Andrew Barto then set out the theory of reinforcement learning in their 1998 textbook.
In Germany, Sepp Hochreiter and Jürgen Schmidhuber published the long short-term memory (LSTM) network in 1997, which later powered speech and translation systems. The field also adopted the idea of intelligent agents, programs that perceive an environment and act to reach goals. Stuart Russell and Peter Norvig built their 1995 textbook Artificial Intelligence: A Modern Approach around that idea, and it became the most widely used AI textbook worldwide. To see how learning-based methods differ from earlier rule-based systems, read our comparison of AI vs machine learning vs deep learning.
Deep Blue Defeats Garry Kasparov
In May 1997, IBM’s Deep Blue beat world chess champion Garry Kasparov in a six-game rematch, after Kasparov had won their first match in February 1996. Deep Blue did not learn in the modern sense. Instead, custom chips let it evaluate up to 200 million positions per second, guided by an evaluation function that grandmasters helped tune. The match still changed public perception, because a machine had beaten the best human at a game long treated as a symbol of intelligence.

Autonomous Systems and the DARPA Grand Challenge
Robotics advanced through public competitions such as the DARPA Grand Challenge. In 2004, DARPA offered a prize for a driverless vehicle that could cross about 150 miles of the Mojave Desert, and no team finished. A year later, Stanford’s car “Stanley,” led by Sebastian Thrun, won the second challenge by using machine learning to read its sensor data. The 2007 Urban Challenge moved the contest into a mock city, and many of its engineers went on to build today’s self-driving programs.

Big Data, the Netflix Prize and ImageNet
The 2000s produced the raw material that deep learning would need, which was data at internet scale. In 2006, Netflix offered a million dollars to anyone who could improve its movie recommendations by 10 percent. A team that blended many machine learning models claimed the prize in 2009. That same year, Fei-Fei Li and her colleagues released ImageNet, a database that eventually held more than 14 million labeled images. ImageNet became the benchmark that deep learning would break wide open in 2012.
The Deep Learning Revolution (2006-2016)
Deep learning is machine learning with neural networks that contain many layers, and it moved to the center of AI between 2006 and 2016. Three forces came together to make the deep learning revolution possible: better training methods, far more data and much faster hardware.
GPUs, Big Data and Deep Neural Networks
In 2006, Geoffrey Hinton and his colleagues at the University of Toronto showed how to train deep networks one layer at a time. A year later, NVIDIA released CUDA, which let researchers run general calculations on graphics processing units (GPUs). GPUs handled the large matrix operations inside neural networks many times faster than ordinary processors. As a result, training jobs that once took months could finish in days.

AlexNet and ImageNet (2012)
The breakthrough arrived in 2012, when Alex Krizhevsky, Ilya Sutskever and Geoffrey Hinton entered a deep convolutional network called AlexNet in the ImageNet challenge. According to their NeurIPS paper, AlexNet achieved a top-5 error rate of 15.3 percent, compared with 26.2 percent for the next best entry. Within two years, almost every leading computer vision team had switched to deep learning, and speech recognition soon followed.
Large technology companies moved quickly to hire the researchers behind these results. Google brought Hinton on board in 2013, while Facebook opened its AI research lab under Yann LeCun that December. Google then bought the London lab DeepMind in 2014, and in 2018 Hinton, LeCun and Yoshua Bengio shared the Turing Award for their work on deep learning.
Watson, Siri and AI in Everyday Life
AI reached ordinary consumers during the same period. In 2011, IBM’s Watson beat two champions on the quiz show Jeopardy!, while Apple introduced Siri on the iPhone 4S. Amazon launched Alexa in 2014, and recommendation engines, photo tagging and voice assistants soon put AI in front of hundreds of millions of people. For modern developments in how machines parse user inputs, explore our guide to how does AI visual search work.
AlphaGo and Reinforcement Learning
In March 2016, DeepMind’s AlphaGo defeated Lee Sedol, one of the world’s best Go players, by four games to one. Go has far more possible positions than chess, so brute-force search alone could never work. AlphaGo combined deep neural networks with reinforcement learning and tree search, as its authors described in Nature. Its successor, AlphaZero, learned chess, shogi and Go from scratch in 2017 without any human game data.

The Rise of Generative AI and Large Language Models (2017-2022)
Generative AI refers to models that create new text, images, audio, video or code instead of only classifying data. It grew out of a single architecture introduced in 2017, then reached the general public five years later.
Transformers and the Attention Mechanism
In June 2017, eight researchers, most of them at Google, published “Attention Is All You Need,” which introduced the transformer architecture. Transformers use an attention mechanism that lets a model weigh how each word in a passage relates to every other word. They also process whole sequences in parallel, which suited GPUs well and made much larger models practical to train. Our guide to how AI works explains tokens, embeddings and attention step by step.

From GPT-1 to GPT-3: The Growth of Large Language Models
Large language models (LLMs) are transformer networks trained to predict the next piece of text across enormous collections of writing. OpenAI released GPT-1 in 2018, and Google released BERT the same year, a model that reads words in context and reached Google Search in 2019. GPT-2 followed in 2019, while GPT-3 arrived in 2020 with 175 billion parameters. GPT-3 could write essays, answer questions and generate code from short prompts, and its performance kept improving with scale, a pattern researchers called scaling laws.
Generative image models advanced during the same years. Ian Goodfellow introduced generative adversarial networks (GANs) in 2014, and diffusion models later powered DALL-E 2, Midjourney and Stable Diffusion in 2022. Our comparison of generative AI vs predictive AI explains how these creative models differ from classic forecasting systems.
AlphaFold and AI in Science
Deep learning also proved its value in scientific research. In 2020, DeepMind’s AlphaFold 2 predicted the three-dimensional shapes of proteins with accuracy close to laboratory methods, cracking a 50-year-old problem in biology. The team later released predicted structures for over 200 million proteins, which researchers now use in drug discovery.
ChatGPT Brings Generative AI to the Mainstream
On November 30, 2022, OpenAI released ChatGPT, a chatbot built on a GPT-3.5 model tuned with reinforcement learning from human feedback (RLHF). It reached an estimated 100 million users within about two months, one of the fastest adoption curves of any consumer app at the time. For most people, ChatGPT answers the question of when AI became popular, since it put a capable language model in every web browser.
AI Today: 2023 to 2026
The years since ChatGPT have moved faster than any earlier period in the evolution of AI. Four developments define the current era: multimodal models, reasoning models, AI agents and a new wave of regulation.
GPT-4 and Multimodal AI
OpenAI released GPT-4 in March 2023, which accepted images as well as text and scored highly on many professional exams. Competitors followed with Anthropic’s Claude, Google’s Bard (renamed Gemini in 2024) and Meta’s Llama models, released with open weights from Llama 2 onward. By 2024, multimodal AI that handles text, images, audio and video inside one model had become standard.
Reasoning Models and Open-Weight Competition
In September 2024, OpenAI introduced o1, a model that works through problems step by step before answering. This approach improved results in mathematics, science and coding. In January 2025, the Chinese lab DeepSeek released DeepSeek-R1, an open-weight reasoning model that rivaled o1 on reasoning benchmarks. Its launch showed that frontier AI was no longer limited to a handful of American companies.
AI Agents
The next step moved AI from answering questions to completing tasks. AI agents pair a language model with tools such as web search, code execution and business software. With those tools, agents can plan and carry out multi-step work with limited human input, and coding, research and support agents spread across businesses during 2025 and 2026. Usage kept climbing as well, since TechCrunch reported that ChatGPT reached 900 million weekly active users in February 2026. Our guide to agentic RAG explains how these multi-step systems work.
AI Governance, Safety and the 2024 Nobel Prizes
Regulators moved faster in this era than in any earlier one. The European Union’s AI Act, the first comprehensive AI law, entered into force on August 1, 2024. Its bans on unacceptable-risk uses applied from February 2025, while rules for general-purpose AI models followed in August 2025. In the United States, the NIST AI Risk Management Framework (2023) became a common reference for companies managing AI risk. To understand how enterprises adapt their compliance and risk postures, review our analysis of AI contextual governance. Furthermore, post-classical computational frontiers are evolving in parallel; see our report on the latest breakthroughs in quantum computing.
Science also honored the field in October 2024. John Hopfield and Geoffrey Hinton won the Nobel Prize in Physics for foundational work on neural networks. Demis Hassabis and John Jumper of DeepMind shared the Nobel Prize in Chemistry with David Baker for protein structure work. Neural network research, which funders had largely abandoned after 1969, had earned science’s highest honor.
Why AI Keeps Booming and Busting: Five Lessons from AI History
Across seventy years, the same pattern keeps repeating: a new method produces striking results, researchers and investors overpromise, then the method hits a wall. Funding dries up until a different approach takes over, and five lessons explain why that cycle keeps returning.
- Predictions outran hardware. Early researchers had the right ambitions but computers millions of times too weak, so neural networks proposed in 1943 only worked at scale once GPUs arrived.
- Data mattered as much as algorithms. Backpropagation existed by 1986, yet deep learning had to wait until the internet and ImageNet supplied enough labeled examples.
- Hand-written knowledge does not scale. Symbolic programs and expert systems broke whenever the world changed, whereas systems that learn from data can simply be retrained.
- Each winter ended with a new method. General problem solvers gave way to expert systems, which in turn yielded to statistical learning before deep learning took the lead.
- Narrow success gets mistaken for general intelligence. ELIZA, Deep Blue and modern chatbots all seemed smarter than they were, raising expectations the next wave had to meet.
These lessons still apply to today’s models, which face limits of their own, including hallucinations, high energy use and a shrinking supply of fresh training data. Artificial intelligence history suggests that the next big gains may come from new approaches rather than from scale alone.
Artificial General Intelligence and the Future of AI
Artificial general intelligence (AGI) means a system that matches human ability across almost any intellectual task, and no such system exists today. Every program in this history, from Logic Theorist to the latest reasoning models, is narrow AI that performs well within the limits of its training. Our guide to the types of artificial intelligence explains the difference between narrow AI, AGI and superintelligence.
Experts disagree widely about when AGI might arrive. In a 2023 survey of 2,778 published AI researchers, the aggregate forecast gave a 50 percent chance that machines could beat humans at every task by 2047. Many respondents gave much later dates, according to the survey report “Thousands of AI Authors on the Future of AI”. History supports caution on both sides, since 1960s researchers promised human-level AI within a generation, yet deep learning then advanced faster than most experts expected.
The near-term future looks clearer than the long-term one. AI agents will take on longer tasks, smaller models will run directly on phones and laptops, and domain-specific models will spread across medicine, law and science. Governance will shape the pace as much as technology, which makes AI ethics, safety and transparency part of the history still being written.
Artificial Intelligence Timeline: Key Milestones by Decade
The table below summarizes the full AI history timeline, from early foundations to the current generative AI era.

| Period | Year | Milestone |
|---|---|---|
| Before 1900 | 1206 | Al-Jazari documents programmable automata in The Book of Knowledge of Ingenious Mechanical Devices |
| Before 1900 | 1843 | Ada Lovelace publishes her notes on Babbage’s Analytical Engine |
| Before 1900 | 1854 | George Boole publishes The Laws of Thought |
| 1930s | 1936 | Alan Turing describes the universal Turing machine |
| 1940s | 1943 | McCulloch and Pitts publish the first mathematical model of an artificial neuron |
| 1940s | 1948 | Norbert Wiener publishes Cybernetics; Claude Shannon publishes information theory |
| 1950s | 1950 | Turing proposes the Turing Test in “Computing Machinery and Intelligence” |
| 1950s | 1951 | Minsky and Edmonds build SNARC, an early neural network machine |
| 1950s | 1955 | McCarthy’s Dartmouth proposal coins the term “artificial intelligence” |
| 1950s | 1956 | The Dartmouth workshop founds AI as a field; Logic Theorist is presented |
| 1950s | 1958 | Rosenblatt introduces the perceptron; McCarthy creates LISP |
| 1950s | 1959 | Arthur Samuel popularizes the term “machine learning” |
| 1960s | 1965 | DENDRAL expert system project begins at Stanford; Zadeh introduces fuzzy logic |
| 1960s | 1966 | ELIZA completed at MIT; Shakey project starts; ALPAC report published |
| 1960s | 1969 | Minsky and Papert publish Perceptrons |
| 1970s | 1973 | Lighthill report leads to UK funding cuts; WABOT-1 completed in Japan |
| 1970s | 1974 | First AI winter begins as DARPA and UK funding fall |
| 1980s | 1980 | XCON enters use at DEC; expert-system boom begins |
| 1980s | 1982 | Japan launches the Fifth Generation project; Hopfield networks introduced |
| 1980s | 1986 | Rumelhart, Hinton and Williams popularize backpropagation |
| 1980s | 1987 | Lisp machine market collapses; second AI winter begins |
| 1980s | 1988 | Judea Pearl publishes his work on Bayesian networks |
| 1990s | 1995 | Support vector machines introduced; Russell and Norvig publish AI: A Modern Approach |
| 1990s | 1997 | Deep Blue defeats Kasparov; LSTM network published |
| 2000s | 2005 | Stanford’s Stanley wins the DARPA Grand Challenge |
| 2000s | 2006 | Hinton’s team revives deep neural networks; Netflix Prize launches |
| 2000s | 2009 | ImageNet dataset released |
| 2010s | 2011 | IBM Watson wins Jeopardy!; Apple launches Siri |
| 2010s | 2012 | AlexNet wins ImageNet and starts the deep learning boom |
| 2010s | 2014 | GANs introduced; Google acquires DeepMind |
| 2010s | 2016 | AlphaGo defeats Lee Sedol |
| 2010s | 2017 | Transformer architecture introduced in “Attention Is All You Need” |
| 2010s | 2018 | GPT-1 and BERT released; Turing Award for deep learning pioneers |
| 2020s | 2020 | GPT-3 released; AlphaFold 2 solves protein structure prediction |
| 2020s | 2022 | ChatGPT launches on November 30 |
| 2020s | 2023 | GPT-4, Claude and Llama 2 released; multimodal AI spreads |
| 2020s | 2024 | EU AI Act enters into force; reasoning models appear; Nobel Prizes for AI research |
| 2020s | 2025 | DeepSeek-R1 released; AI agents spread across businesses |
| 2020s | 2026 | ChatGPT reaches 900 million weekly active users in February |
Frequently Asked Questions About the History of Artificial Intelligence
When Was AI Invented?
Artificial intelligence became a formal field of research at the 1956 Dartmouth workshop, and the term first appeared in its August 1955 proposal. The core ideas are older, since McCulloch and Pitts modeled artificial neurons in 1943 and Turing asked whether machines could think in 1950.
Who Invented AI?
No single person invented AI, since the work was shared across several pioneers. Alan Turing laid its theoretical foundations, John McCarthy named the field with three co-organizers, and Allen Newell and Herbert Simon built the first AI program.
Who Is the Father of AI?
John McCarthy is most often called the father of artificial intelligence, because he coined the term, organized the Dartmouth workshop and created LISP. Alan Turing is sometimes given the same title for his earlier theoretical work.
Who Coined the Term Artificial Intelligence?
John McCarthy coined the term “artificial intelligence” in the 1955 Dartmouth proposal, choosing it to set the new field apart from cybernetics.
What Was the First AI Program?
Logic Theorist, created by Allen Newell, Herbert Simon and Cliff Shaw in 1955 and 1956, is widely considered the first AI program. Christopher Strachey’s checkers program, which ran on the Ferranti Mark 1 in 1952, was an earlier game-playing program.
Is ChatGPT the First AI?
No, ChatGPT arrived 66 years after the Dartmouth workshop. It was the first AI tool to reach a mass audience so quickly, but it builds on decades of work in neural networks, machine learning and transformers.
How Old Is AI?
As a research field, AI is 70 years old in 2026, counting from the 1956 Dartmouth workshop. The idea of artificial beings is more than two thousand years older, reaching back to Greek myths and ancient automata.
What Are the Main Eras of AI History?
AI history falls into five broad eras: symbolic AI (1956 to 1974), expert systems (1980 to 1987), statistical machine learning (1990s and 2000s), deep learning (2012 to 2017) and generative AI (2017 onward). Two AI winters separate the first three eras.
Why Did the AI Winters Happen?
Both AI winters happened because promises outran what the technology could deliver. The first followed the ALPAC and Lighthill reports plus DARPA funding changes. The second followed the collapse of the Lisp machine market and the high upkeep of expert systems.
When Did Machine Learning and Deep Learning Begin?
Machine learning began in the 1950s with Arthur Samuel’s self-improving checkers program, and he popularized the term in 1959. Deep learning grew from 1980s neural network research but took off in 2012, when AlexNet won the ImageNet challenge.
When Did AI Become Popular?
AI reached mass popularity after OpenAI released ChatGPT on November 30, 2022. Earlier waves of public attention followed Deep Blue’s 1997 chess win, the 2011 launch of Siri and AlphaGo’s 2016 victory.
Conclusion: What the History of AI Teaches Us
The history of artificial intelligence is a story of big ideas that waited decades for hardware, data and methods to catch up. Turing asked the question in 1950, and the Dartmouth group named the field in 1956. Symbolic AI and expert systems rose and fell, while statistical learning quietly rebuilt the field before deep learning carried it into daily life. Each boom brought real progress, and each winter taught researchers to promise less while measuring more.
Knowing this history helps you judge today’s AI claims with a clearer eye. If your team wants to separate what works today from what is still hype, our AI consulting team can help you choose an approach that fits your data and goals. For unfamiliar terms from this timeline, browse the SoftBrixAI glossary of AI terms.
References
- Turing, A. M. (1950). “Computing Machinery and Intelligence.” Mind, 59(236), 433-460. https://academic.oup.com/mind/article/LIX/236/433/986238
- McCarthy, J., Minsky, M., Rochester, N., and Shannon, C. (1955). A Proposal for the Dartmouth Summer Research Project on Artificial Intelligence. https://www-formal.stanford.edu/jmc/history/dartmouth/dartmouth.html
- National Research Council (1966). Language and Machines: Computers in Translation and Linguistics (ALPAC report). https://nap.nationalacademies.org/catalog/9547/language-and-machines-computers-in-translation-and-linguistics
- Lighthill, J. (1973). Artificial Intelligence: A General Survey. Science Research Council. https://www.chilton-computing.org.uk/inf/literature/reports/lighthill_report/p001.htm
- Rumelhart, D., Hinton, G., and Williams, R. (1986). “Learning representations by back-propagating errors.” Nature, 323, 533-536. https://doi.org/10.1038/323533a0
- Krizhevsky, A., Sutskever, I., and Hinton, G. (2012). “ImageNet Classification with Deep Convolutional Neural Networks.” NeurIPS. https://proceedings.neurips.cc/paper_files/paper/2012/file/c399862d3b9d6b76c8436e924a68c45b-Paper.pdf
- Silver, D., et al. (2016). “Mastering the game of Go with deep neural networks and tree search.” Nature, 529, 484-489. https://doi.org/10.1038/nature16961
- Vaswani, A., et al. (2017). “Attention Is All You Need.” https://arxiv.org/abs/1706.03762
- DeepSeek-AI (2025). “DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.” https://arxiv.org/abs/2501.12948
- Grace, K., et al. (2024). “Thousands of AI Authors on the Future of AI.” https://arxiv.org/abs/2401.02843
- The Nobel Prize in Physics 2024. https://www.nobelprize.org/prizes/physics/2024/summary/
- The Nobel Prize in Chemistry 2024. https://www.nobelprize.org/prizes/chemistry/2024/summary/
- Regulation (EU) 2024/1689 (Artificial Intelligence Act). https://eur-lex.europa.eu/eli/reg/2024/1689/oj
- TechCrunch (February 27, 2026). “ChatGPT reaches 900M weekly active users.” https://techcrunch.com/2026/02/27/chatgpt-reaches-900m-weekly-active-users