Progressive discretization for generative retrieval: A self-supervised approach to high-quality DocID generation.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 40516381.
- Also identified by DOI 10.1016/j.neunet.2025.107663.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Generative retrieval is a novel retrieval paradigm where large language models serve as differentiable indices to memorize and retrieve candidate documents in a generative fashion. This paradigm overcomes the limitation that documents and queries must be encoded separately and demonstrates superior performance compared to traditional retrieval methods. To support the retrieval of large-scale corpora, extensive research has been devoted to devising a discrete and distinguishable document representation, namely the DocID. However, most DocIDs are built under unsupervised circumstances, where uncontrollable information distortion will be introduced during the discretization stage. In this work, we propose the Self-supervised Progressive Discretization framework (SPD). SPD first distills document information into multi-perspective continuous representations in a self-supervised way. Then, a progressive discretization algorithm is employed to transform the continuous representations into approximate vectors and discrete DocIDs. The self-supervised model, approximate vectors, and DocIDs are further integrated into a query-side training pipeline to produce an effective generative retriever. Experiments on popular benchmarks demonstrate that SPD builds high-quality search-oriented DocIDs that achieve state-of-the-art generative retrieval performance.
Medical subject headings
- Information Storage and Retrieval
- Neural Networks, Computer
- Supervised Machine Learning