Explain Like I'm a Specific Expert
Explain how a large language model (like GPT or Claude) learns and generates text to each of these three audiences: 1. **An experienced…
For the software engineer Start with the part you'll find suspicious and let me argue the other way. Yes, the training objective is literally "given tokens 1..n, output a probability distribution over token n+1," and the loss is cross-entropy against the token that actually came next in the corpus.Read the full answer
How Large Language Models Work: Three Explanations For the Experienced Software Engineer Think of an LLM as a single, very large pure function: f(tokens[0..n]) → probability distribution over tokens[n+1]. Text is chunked into tokens (roughly word fragments, from a vocabulary of about 100k).Read the full answer