Large Language Models Explained

What is really inside a chatbot. Large language models learn to predict the next word from huge amounts of text. How they are trained, why they sound smart, and where they slip.

Rows of server racks with blinking lights inside a data center

Language models are trained in data centers: Carl Lender from Sunrise, USA, CC BY 2.0

A very well-read autocomplete

At heart, a large language model does one thing: given some text, it predicts what comes next. Text is split into tokens, which are whole words or pieces of words. The model looks at the tokens so far and produces a probability for every possible next token. One is picked, added to the text, and the process repeats, which is how a full answer appears a little at a time.

That sounds simple, but predicting text well requires picking up grammar, facts, styles of reasoning and a great deal about how people communicate. To guess the next word in a chemistry textbook or a legal contract, the model has to absorb a lot about chemistry or law.

Video: Large Language Models explained briefly (3Blue1Brown), embedded from YouTube.

How they are trained

Training happens in stages. In pretraining, a Transformer network with billions of adjustable numbers, called parameters, reads an enormous amount of text gathered from books, websites, code and other sources. Each time it guesses a next token wrong, its parameters are adjusted slightly. This step takes huge clusters of specialized chips running for weeks or months.

The result is knowledgeable but not yet a good assistant. A second stage, fine-tuning, teaches it to follow instructions and hold a helpful conversation, often using examples written by people and feedback in which humans or other models rate its answers. Developers also train in safety behaviors, such as declining clearly harmful requests.

Why they make mistakes

An LLM produces text that is likely, not text that is guaranteed true. When it lacks reliable information it can still write a confident, fluent answer, a failure known as hallucination. It also has a context window, a limit on how much text it can consider at once, and its built-in knowledge stops at a training cutoff unless it is connected to search or other tools.

Models can reflect biases in their training data and can be thrown off by oddly worded questions. That is why good practice is to treat answers as a strong first draft and to check anything that matters.

What they are good at

LLMs shine at drafting and editing, summarizing long documents, explaining ideas at different levels, translating, brainstorming and writing or reviewing code. Many can now also take in images and audio, and newer reasoning models spend extra computation working through a problem step by step before answering. Used with care, they are among the most flexible tools ever built for working with words.

Media credits
  • Language models are trained in data centers: Carl Lender from Sunrise, USA, CC BY 2.0

Text written by Strawberry Lemonadai.

← Neural Networks ExplainedHow AI Image Generators Work →

Shop

The Strawberry Lemonadai collection

Tees and hoodies in the Strawberry Lemonadai colors, printed to order in the USA. Use code FIRST15 from the newsletter for 15% off your first order.

Watch

Strawberry Lemonadai Shorts

Quick, accurate, under a minute. Every Short is original and made only for this channel.

YouTube @StrawberryLemonadAI
Channel launching soon. Subscribe on the newsletter to hear first.
Get notified
A computer beat the world chess champion in 1997Photos: IBM Deep Blue at Computer History Museum (9361685537).jpg - Anton Chiang from Cupertino, CA, USA (CC BY 2.0) via Wikimedia Commons | Garry Kasparov (37097592314).jpg - Gage Skidmore from Peoria, AZ, United States of America (CC BY-SA 2.0) via Wikimedia Commons | Chess game Staunton No. 6 perfil view 8.jpg - Wilfredor (CC0) via Wikimedia Commons | Chess game Staunton No. 6.jpg - Wilfredor (CC0) via
An AI solved a 50 year biology puzzle and won a Nobel PrizePhotos: Human Oxy-Hemoglobin Protein.jpg - PDB code 2DN1 1.25 a resolution crystal structures of human (CC0) via Wikimedia Commons | Ribbon diagram of the DED.jpg - BQUB16-Oibanez (CC BY-SA 4.0) via Wikimedia Commons | Laboratory pipettes.jpg - J.N. Eskra (CC BY-SA 4.0) via Wikimedia Commons | Use of a Multichannel Pipette for High-Throughput Liquid Handling in a Biosafety Cabinet.jpg - Siduduziwe Nxumal
ChatGPT hit an estimated 100 million users in 2 monthsPhotos: BalticServers data center.jpg - BalticServers.com (CC BY-SA 3.0) via Wikimedia Commons | Diverse people using phones.jpeg - Rawpixel Ltd (CC BY 2.0) via Wikimedia Commons | Crowd of people with phones.jpg - Rawpixel Ltd (CC BY 2.0) via Wikimedia Commons | Sam Altman CropEdit James Tamim.jpg - TechCrunch (CC BY 2.0) via Wikimedia Commons | Ilya Sutskever and Sam Altman in TAU.jpg - Eladkarmel (CC B
The T in ChatGPT comes from one 2017 paperPhotos: Mountain View (CA, USA), Charleston Road, Google-Fahrräder -- 2022 -- 2901.jpg - Dietmar Rabich (CC BY-SA 4.0) via Wikimedia Commons | Mountain View (CA, USA), Charleston Road, Abstellplatz für Google-Fahrräder -- 2022 -- 2899.jpg - Dietmar Rabich (CC BY-SA 4.0) via Wikimedia Commons | BalticServers data center.jpg - BalticServers.com (CC BY-SA 3.0) via Wikimedia Commons | Пример кода на Java.jpg
The word robot was invented for a 1920 playPhotos: Karel Čapek podepisuje první výtisky Povětroně, Pestrý týden 27.1.1934.jpg - Unknown authorUnknown author (Public domain) via Wikimedia Commons | Karel Čapek 30.léta.jpg - re-photo by David Sedlecký (Public domain) via Wikimedia Commons | Plakat za predstavo R.U.R Rossmus Universal Robots v Narodnem gledališču v Mariboru 28. oktobra 1933.jpg - Unknown authorUnknown author (Public domain) via Wikim
Newsletter

The weekly AI digest

The AI stories that actually matter this week, explained in plain English, plus one useful tool tip. No hype, no jargon, unsubscribe anytime.