Is a (satellite) image worth a thousand data points? Comparing machine learning approaches to predict environmental and social inequalities in England.
other
Where this comes from
- Record sourced from PubMed, PMID 42627871.
- Also identified by DOI 10.1371/journal.pone.0356472.
- No licence information is recorded for this record.
- Because redistribution is not established, this page shows the abstract only. Follow the links below for the full text.
Abstract
Accurately mapping environmental and social inequalities at fine spatial scales is critical for urban policy, yet the data required are often costly and infrequently updated. Vision foundation models can extract information directly from satellite imagery, offering a rapid and scalable alternative. We compare three modelling pipelines for predicting two contrasting indicators of urban inequality - air pollution and house prices - across England on a fine hexagonal grid: regression models trained on structured features of form and function (census, land cover and urban morphology), models trained on 128-dimensional image embeddings from a geospatial foundation model and a hybrid of the two, each evaluated with and without coarse regional context. Structured features achieve the best overall accuracy, but image embeddings become competitive once regional context is added - most clearly for air pollution, where image-only models reach a median R2 of 0.78 (0.85 with regional context), indicating that the embeddings capture genuine image signal. For house prices the picture is more cautious: image-only models achieve a median R2 of around 0.58; however, much of the embeddings' apparent gain reflects coarse spatial location rather than image content. Off-the-shelf satellite embeddings, while not yet surpassing data-intensive approaches, are a promising low-cost complement, particularly for rapid, large-scale, or repeated analyses and in settings where traditional data are limited.
Medical subject headings
- Satellite Imagery
- Machine Learning