LLM Fundamentals

1. LLM Architecture — High Level

An LLM (Large Language Model) is a neural network trained on massive amounts of text to understand and generate language. Modern LLMs are primarily based on the Transformer architecture, where input text is tokenized, converted into embeddings, processed through multiple transformer layers containing attention and feed-forward networks, and finally converted into probabilities for the next token. During generation, the model repeatedly predicts the next token until the response is complete. Interview line: “At a high level, an LLM takes tokenized input, converts it into vector representations, processes them through transformer layers using attention and neural networks, and predicts output tokens probabilistically.”

2. Transformers

A Transformer is the architecture behind most modern LLMs. Its major components are token embeddings, positional information, self-attention, feed-forward networks, and normalization layers. Unlike older RNNs, Transformers can process tokens in parallel during training and use attention to understand relationships between tokens regardless of their distance in the text. Models such as GPT are essentially Transformer-based architectures optimized for language generation.

3. Attention / Self-Attention

Self-attention allows each token to determine which other tokens in the input are important for understanding its meaning. It uses Query, Key, and Value representations: the query asks what information is relevant, keys represent available information, and values contain the actual information to aggregate. For example, in “The bank approved the loan because it had sufficient funds,” attention helps the model understand what “it” refers to based on surrounding context. Multi-head attention performs this process from multiple learned perspectives.

4. Tokens & Tokenization

Tokenization converts text into smaller units called tokens before an LLM processes it. A token may be a complete word, part of a word, punctuation, or sometimes a space-related unit depending on the tokenizer. The model doesn't directly understand text; it processes token IDs that are converted into numerical vectors. Tokenization matters because context windows, API costs, latency, and model limits are generally measured in tokens rather than words.

5. Context Window

A context window is the maximum amount of tokenized information an LLM can consider in a single request, including the system prompt, user input, retrieved documents, conversation history, and generated output depending on the model's limits. A larger context window allows more information to be supplied, but simply having a large window doesn't mean every piece of information will be equally useful. That's why production systems still use RAG, retrieval, reranking, summarization, and context compression to provide the most relevant information.

6. Temperature / Top-p

Temperature controls how random or deterministic the model's token selection is: lower temperature makes outputs more predictable, while higher temperature increases diversity. Top-p (nucleus sampling) restricts token selection to the smallest group of tokens whose combined probability reaches a specified probability mass. For production applications such as structured extraction, we generally prefer lower randomness, while creative generation can tolerate higher randomness. Interview distinction: temperature changes the probability distribution's sharpness; top-p limits the candidate probability mass.

7. Pretraining vs Fine-tuning vs Prompting

Pretraining is the initial large-scale training where a model learns general language patterns, knowledge, and capabilities from huge datasets. Fine-tuning further trains an existing model on a specific dataset to change or specialize its behavior, such as following a particular style or performing a domain task. Prompting doesn't modify the model's weights; instead, we provide instructions and context at inference time. For example, if I need an LLM to answer questions using changing legal documents, RAG + prompting is usually more appropriate than fine-tuning because the knowledge changes frequently.

8. Inference

Inference is the process of using a trained model to generate a prediction or response. In an LLM application, inference typically means sending tokens to the model, running them through the neural network, calculating probabilities for the next token, selecting a token according to the decoding strategy, and repeating the process until completion. In production, inference performance is heavily influenced by model size, input/output tokens, hardware, batching, caching, quantization, and model serving infrastructure.

9. Reasoning Models vs Normal LLMs

A traditional LLM is generally optimized to generate the next token effectively, while reasoning models are specifically trained or optimized to spend more computation on complex multi-step reasoning before producing the final answer. Reasoning models can be particularly useful for difficult mathematics, coding, planning, or complex decision-making, but they can introduce additional latency and cost. The important interview point is: use a reasoning model when the task actually benefits from deeper reasoning; don't automatically use the most expensive model for every task.