The Perceptron
Frank Rosenblatt's Perceptron was one of the first algorithms that could learn a linear decision boundary from labeled examples, rather than being explicitly programmed with rules.

From rule-based systems to classical machine learning, deep learning, transformers, and today's large language models.
Frank Rosenblatt's Perceptron was one of the first algorithms that could learn a linear decision boundary from labeled examples, rather than being explicitly programmed with rules.

Joseph Weizenbaum's ELIZA simulated a conversation using pattern matching and scripted rules, with no learning or understanding involved, yet convinced many users they were talking to something intelligent.
Rumelhart, Hinton, and Williams showed backpropagation could efficiently train multi-layer neural networks, giving the field a practical way to adjust every weight in a network based on its errors.
SVMs became a dominant approach for classification tasks through the 1990s and 2000s, relying on hand-engineered features rather than learned representations.
A deep convolutional neural network trained on GPUs dramatically beat the previous state of the art on the ImageNet competition, kicking off the modern deep learning era.

Mikolov et al. showed that simple neural networks could learn dense word embeddings from raw text, where words with similar meanings ended up with similar vectors.
Bahdanau et al. introduced an attention mechanism for neural machine translation, letting a model look back at relevant input words instead of compressing a whole sentence into one fixed vector.
Vaswani et al. showed that attention alone, without recurrence or convolution, was enough to build state-of-the-art sequence models, and dramatically more parallelizable to train.

Google's BERT applied bidirectional transformer pretraining to language understanding tasks, setting new benchmarks across a wide range of NLP problems.
OpenAI's GPT-2 demonstrated that a large transformer trained only to predict the next token could generate coherent, general-purpose text across many tasks without task-specific training.
Scaling the same next-token-prediction approach to 175 billion parameters produced a model that could perform new tasks from just a handful of examples in its prompt, with no retraining.
Fine-tuning a GPT-3.5 model on human feedback (RLHF) and packaging it as a conversational assistant brought LLMs to hundreds of millions of everyday users almost overnight.
Meta released LLaMA, a family of open-weights language models from 7 billion to 65 billion parameters, showing that smaller models trained on more data could match much larger closed models and kicking off a wave of open-weights research and fine-tuning.
OpenAI's GPT-4 accepted both text and images as input and passed several professional and academic benchmarks at a human level, becoming the new reference point for frontier model capability.
Anthropic released its first Claude models through its API, trained with Constitutional AI, an approach that uses a set of written principles to guide model behavior instead of relying only on human-labeled preferences.
Rafailov et al. introduced DPO, a way to align a language model with human preference data by optimizing a simple classification loss directly on that data, without training a separate reward model or running reinforcement learning.
OpenAI added function calling to its API, letting a model output a structured call to a developer-defined function instead of only free text, a pattern other providers adopted and that became the basis for most agent and tool-use frameworks.
Meta released Llama 2 with a license that permitted commercial use, a shift from the research-only terms of the original LLaMA that helped establish open-weights models as viable options for production products.
Google introduced Gemini as its first natively multimodal model family, built and trained to handle text, images, audio, and video together rather than adding vision onto a text-only model afterward.
Mistral AI released Mixtral, an open-weights sparse mixture-of-experts model that routed each token to 2 of 8 expert sub-networks per layer, matching or beating much larger dense models while activating only a fraction of its total parameters at inference time.
As developers ran into the limits of a model's fixed training data and context window, retrieving relevant documents at query time and inserting them into the prompt, an idea first described by Lewis et al. in 2020, became the standard way to ground LLM applications in external or up-to-date information.
Google previewed Gemini 1.5 Pro with a context window of up to 1 million tokens, far beyond prior commercial models, prompting a broader industry push toward long-context models that could hold entire codebases or books in a single prompt.
Meta released Llama 3, trained on over 15 trillion tokens, which closed much of the performance gap between open-weights and the best closed models of the time and became a widely used base for fine-tuned and specialized models.
OpenAI's GPT-4o was trained end-to-end across text, vision, and audio in a single model, cutting voice-conversation latency to near real-time and making natively multimodal interaction, rather than separate models stitched together, the new default.
OpenAI released o1, a model trained to generate an extended internal chain of reasoning before answering, trading more computation at inference time for better performance on math, coding, and science problems, a technique often called test-time or inference-time scaling.
Anthropic open-sourced the Model Context Protocol, a standard way for a language model to connect to external data sources and tools, aimed at replacing one-off custom integrations with a common interface that any compatible client or server could use.
DeepSeek released R1, an open-weights reasoning model trained largely with reinforcement learning on top of the DeepSeek-V3 base model (itself reportedly pretrained for around $5.6 million), that matched OpenAI's o1 on several benchmarks and showed its full chain of reasoning to users. R1's own additional training cost was never separately disclosed.
Anthropic released Claude Opus 4 and Claude Sonnet 4, hybrid models that could switch between a fast response mode and an extended thinking mode, with a particular focus on sustained multi-step coding and agentic tasks.
OpenAI released GPT-5 as a system that automatically routes each request between a fast-response mode and a deeper reasoning mode depending on the task, rather than requiring users to pick between separate model families.
Google released Gemini 3, its next-generation multimodal and reasoning model family, including a DeepThink variant for extended reasoning, continuing the shift toward models that scale performance with more inference-time computation.