A law of next-token prediction in large language models.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41116396.
- Also identified by DOI 10.1103/5rn3-49lc.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Large language models (LLMs) have been widely employed across various application domains, yet their black-box nature poses significant challenges to understanding how these models process input data internally to make predictions. In this paper, we introduce a precise and quantitative law that governs the learning of contextualized token embeddings through intermediate layers in pretrained LLMs for next-token prediction. Our findings reveal that each layer contributes equally to enhancing prediction accuracy, from the lowest to the highest layer-a universal phenomenon observed across a diverse array of open-source LLMs, irrespective of their architectures or pretraining data. We demonstrate that this law offers different perspectives and actionable insights to inform and guide practices in LLM development and applications, including model scaling, pretraining tasks, and interpretation.