How we got here  /  30 moments, 6 eras

67 years, one idea.

From rule-based systems to classical machine learning, deep learning, transformers, and today's large language models.

Rule-based systems and early neural nets2Classical machine learning2Deep learning3Transformers2Large language models15Agentic and reasoning models6
Era 01 of 06

Rule-based systems and early neural nets

1958–1966 · 2 events
1958

The Perceptron

Frank Rosenblatt's Perceptron was one of the first algorithms that could learn a linear decision boundary from labeled examples, rather than being explicitly programmed with rules.

neural-networkshistory
The Perceptron
John C. Hay and Albert E. Murray, Mark I Perceptron Operators' Manual, Public domain, Wikimedia Commons
1966

ELIZA

Joseph Weizenbaum's ELIZA simulated a conversation using pattern matching and scripted rules, with no learning or understanding involved, yet convinced many users they were talking to something intelligent.

nlphistory
Era 02 of 06

Classical machine learning

1986–1995 · 2 events
1986

Backpropagation popularized

Rumelhart, Hinton, and Williams showed backpropagation could efficiently train multi-layer neural networks, giving the field a practical way to adjust every weight in a network based on its errors.

traininghistory
Backpropagation popularized
Sky99, CC BY-SA 3.0, Wikimedia Commons
1995

Support Vector Machines mature

SVMs became a dominant approach for classification tasks through the 1990s and 2000s, relying on hand-engineered features rather than learned representations.

classical-mlhistory
Support Vector Machines mature
ZackWeinberg, based on a PNG version by Cyc, CC BY-SA 3.0, Wikimedia Commons
Era 03 of 06

Deep learning

2012–2014 · 3 events
2012

AlexNet

A deep convolutional neural network trained on GPUs dramatically beat the previous state of the art on the ImageNet competition, kicking off the modern deep learning era.

deep-learningcomputer-vision
AlexNet
TheStriker, CC BY-SA 4.0, Wikimedia Commons (EVGA GeForce GTX 580, the GPU model used to train AlexNet)
2013

word2vec

Mikolov et al. showed that simple neural networks could learn dense word embeddings from raw text, where words with similar meanings ended up with similar vectors.

embeddingsnlp
2014

Sequence-to-sequence models with attention

Bahdanau et al. introduced an attention mechanism for neural machine translation, letting a model look back at relevant input words instead of compressing a whole sentence into one fixed vector.

attentionnlp
Era 04 of 06

Transformers

2017–2018 · 2 events
2017

Attention Is All You Need

Vaswani et al. showed that attention alone, without recurrence or convolution, was enough to build state-of-the-art sequence models, and dramatically more parallelizable to train.

transformersattentionlandmark-paper
Attention Is All You Need
dvgodoy, CC BY 4.0, Wikimedia Commons
2018

BERT

Google's BERT applied bidirectional transformer pretraining to language understanding tasks, setting new benchmarks across a wide range of NLP problems.

transformerspretraining
Era 05 of 06

Large language models

2019–2024 · 15 events
2019

GPT-2

OpenAI's GPT-2 demonstrated that a large transformer trained only to predict the next token could generate coherent, general-purpose text across many tasks without task-specific training.

llmgenerative
2020

GPT-3

Scaling the same next-token-prediction approach to 175 billion parameters produced a model that could perform new tasks from just a handful of examples in its prompt, with no retraining.

llmscalingfew-shot
2022

ChatGPT

Fine-tuning a GPT-3.5 model on human feedback (RLHF) and packaging it as a conversational assistant brought LLMs to hundreds of millions of everyday users almost overnight.

llmrlhfchat
2023

LLaMA

Meta released LLaMA, a family of open-weights language models from 7 billion to 65 billion parameters, showing that smaller models trained on more data could match much larger closed models and kicking off a wave of open-weights research and fine-tuning.

llmopen-weights
2023

GPT-4

OpenAI's GPT-4 accepted both text and images as input and passed several professional and academic benchmarks at a human level, becoming the new reference point for frontier model capability.

llmmultimodal
2023

Claude

Anthropic released its first Claude models through its API, trained with Constitutional AI, an approach that uses a set of written principles to guide model behavior instead of relying only on human-labeled preferences.

llmrlhfchat
2023

Direct Preference Optimization

Rafailov et al. introduced DPO, a way to align a language model with human preference data by optimizing a simple classification loss directly on that data, without training a separate reward model or running reinforcement learning.

rlhfalignmenttraining
2023

Function calling and tool use

OpenAI added function calling to its API, letting a model output a structured call to a developer-defined function instead of only free text, a pattern other providers adopted and that became the basis for most agent and tool-use frameworks.

tool-useagents
2023

Llama 2

Meta released Llama 2 with a license that permitted commercial use, a shift from the research-only terms of the original LLaMA that helped establish open-weights models as viable options for production products.

llmopen-weights
2023

Gemini 1.0

Google introduced Gemini as its first natively multimodal model family, built and trained to handle text, images, audio, and video together rather than adding vision onto a text-only model afterward.

llmmultimodal
2023

Mixtral 8x7B

Mistral AI released Mixtral, an open-weights sparse mixture-of-experts model that routed each token to 2 of 8 expert sub-networks per layer, matching or beating much larger dense models while activating only a fraction of its total parameters at inference time.

moeopen-weightsllm
2023

Retrieval-augmented generation goes mainstream

As developers ran into the limits of a model's fixed training data and context window, retrieving relevant documents at query time and inserting them into the prompt, an idea first described by Lewis et al. in 2020, became the standard way to ground LLM applications in external or up-to-date information.

ragretrieval
2024

Gemini 1.5 and the long-context race

Google previewed Gemini 1.5 Pro with a context window of up to 1 million tokens, far beyond prior commercial models, prompting a broader industry push toward long-context models that could hold entire codebases or books in a single prompt.

long-contextllm
2024

Llama 3

Meta released Llama 3, trained on over 15 trillion tokens, which closed much of the performance gap between open-weights and the best closed models of the time and became a widely used base for fine-tuned and specialized models.

llmopen-weights
2024

GPT-4o

OpenAI's GPT-4o was trained end-to-end across text, vision, and audio in a single model, cutting voice-conversation latency to near real-time and making natively multimodal interaction, rather than separate models stitched together, the new default.

llmmultimodal
Era 06 of 06

Agentic and reasoning models

2024–2025 · 6 events
2024

OpenAI o1 and inference-time reasoning

OpenAI released o1, a model trained to generate an extended internal chain of reasoning before answering, trading more computation at inference time for better performance on math, coding, and science problems, a technique often called test-time or inference-time scaling.

reasoninginference-time-scaling
2024

Model Context Protocol

Anthropic open-sourced the Model Context Protocol, a standard way for a language model to connect to external data sources and tools, aimed at replacing one-off custom integrations with a common interface that any compatible client or server could use.

mcptool-useagents
2025

DeepSeek-R1

DeepSeek released R1, an open-weights reasoning model trained largely with reinforcement learning on top of the DeepSeek-V3 base model (itself reportedly pretrained for around $5.6 million), that matched OpenAI's o1 on several benchmarks and showed its full chain of reasoning to users. R1's own additional training cost was never separately disclosed.

reasoningopen-weightsrlhf
2025

Claude 4

Anthropic released Claude Opus 4 and Claude Sonnet 4, hybrid models that could switch between a fast response mode and an extended thinking mode, with a particular focus on sustained multi-step coding and agentic tasks.

reasoningagentsllm
2025

GPT-5

OpenAI released GPT-5 as a system that automatically routes each request between a fast-response mode and a deeper reasoning mode depending on the task, rather than requiring users to pick between separate model families.

reasoningllm
2025

Gemini 3

Google released Gemini 3, its next-generation multimodal and reasoning model family, including a DeepThink variant for extended reasoning, continuing the shift toward models that scale performance with more inference-time computation.

reasoningmultimodalllm