A Non-negative Deep VAE: the Generalized Gamma Belief Network.
basic_science · Level V
Where this comes from
- Record sourced from PubMed, PMID 41941806.
- Also identified by DOI 10.1109/TPAMI.2026.3680869.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Gamma belief network (GBN), widely viewed as deep probabilistic topic models, has demonstrated its potential for uncovering multi-layer interpretable latent representations from text corpora. Its notable performance in document modeling largely arises from the expressive nature of gamma-distributed latent variables, which naturally capture sparsity, nonnegativity, skewness, heavy-tailed pattens, and from their seamless extension to multi-layer hierarchical structures. However, existing GBN and its variations are constrained by linear generative model, thereby limiting their expressiveness and applicability. To address this limitation, we introduce Generalized Gamma Belief Network (Generalized GBN), which extends original linear generative model to a more expressive non-linear generative model. Since parameters of Generalized GBN no longer possess an analytic conditional posterior, we further propose an upward-downward Weibull inference network to approximate posterior distribution of latent variables. The parameters of both generative model and inference network are jointly trained within variational inference framework. In addition, we provide theoretical analyses that demonstrate the effectiveness of Generalized GBN in modeling data variability and achieving disentangled representations. The former benefit arises from its hierarchical latent-variable structure, while the latter stems from its inherent ability to model sparsity. Finally, we conduct comprehensive experiments on both expressivity and disentangled representation learning tasks to evaluate the performance of Generalized GBN against Gaussian variational autoencoders serving as strong baseline models.