Pre-training
The initial phase in which a model learns general patterns from a large dataset before being adapted to specific tasks or domains. In LLMs, this usually relies on self-supervised objectives such as predicting missing or subsequent tokens, updating weights to compress regularities in language and the training corpus.
Pre-training creates the statistical base on which prompting, in-context learning, and fine-tuning operate. Broad coverage does not guarantee current facts, truth, consent over the data, or uniform competence across domains. The corpus, objective, and filtering process determine a meaningful share of the biases and gaps carried into the final model.