Text embedding models yield detailed conceptual knowledge maps derived from short multiple-choice quizzes.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41876490.
- Also identified by DOI 10.1038/s41467-026-69746-w and PMC identifier 13013618.
- Licence recorded as CC BY.
- The licence permits redistribution, so the abstract is shown in full and the full text is available from the publisher.
Abstract
Real-world conceptual knowledge is complex, multifaceted, and substantially over-simplified in most laboratory studies. Here we develop a mathematical framework, based on natural language processing models, for tracking and characterizing the acquisition of real-world conceptual knowledge. Our approach embeds each concept in a high-dimensional representation space where nearby coordinates reflect similar or related concepts. We test our approach using behavioral data from participants who answered small sets of multiple-choice quiz questions interleaved between watching two Khan Academy course videos. We apply our framework to the videos' transcripts and the text of the quiz questions to quantify the content of each moment of video and each quiz question. We use these embeddings, along with participants' quiz responses, to track how the learners' knowledge changed after watching each video and predict their success on individual quiz questions. Our findings show how a small set of quiz questions may be used to obtain rich and meaningful insights into what each learner knows, and how their knowledge changes over time as they learn.
Medical subject headings
- Knowledge
- Learning
- Educational Measurement
- Concept Formation