Generalized Bayes Quantification Learning under Dataset Shift

Jacob Fiksel; Abhirup Datta; Agbessi Amouzou; Scott Zeger

doi:10.1080/01621459.2021.1909599

Generalized Bayes Quantification Learning under Dataset Shift

Jacob Fiksel, Abhirup Datta, Agbessi Amouzou, Scott Zeger

Bloomberg School of Public Health

Research output: Contribution to journal › Article › peer-review

1 Scopus citations

Abstract

Quantification learning is the task of prevalence estimation for a test population using predictions from a classifier trained on a different population. Quantification methods assume that the sensitivities and specificities of the classifier are either perfect or transportable from the training to the test population. These assumptions are inappropriate in the presence of dataset shift, when the misclassification rates in the training population are not representative of those for the test population. Quantification under dataset shift has been addressed only for single-class (categorical) predictions and assuming perfect knowledge of the true labels on a small subset of the test population. We propose generalized Bayes quantification learning (GBQL) that uses the entire compositional predictions from probabilistic classifiers and allows for uncertainty in true class labels for the limited labeled test data. Instead of positing a full model, we use a model-free Bayesian estimating equation approach to compositional data using Kullback–Leibler loss-functions based only on a first-moment assumption. The idea will be useful in Bayesian compositional data analysis in general as it is robust to different generating mechanisms for compositional data and allows 0’s and 1’s in the compositional outputs thereby including categorical outputs as a special case. We show how our method yields existing quantification approaches as special cases. Extension to an ensemble GBQL that uses predictions from multiple classifiers yielding inference robust to inclusion of a poor classifier is discussed. We outline a fast and efficient Gibbs sampler using a rounding and coarsening approximation to the loss functions. We establish posterior consistency, asymptotic normality and valid coverage of interval estimates from GBQL, which to our knowledge are the first theoretical results for a quantification approach in the presence of local labeled data. We also establish finite sample posterior concentration rate. Empirical performance of GBQL is demonstrated through simulations and analysis of real data with evident dataset shift. Supplementary materials for this article are available online.

Original language	English (US)
Pages (from-to)	2163-2181
Number of pages	19
Journal	Journal of the American Statistical Association
Volume	117
Issue number	540
DOIs	https://doi.org/10.1080/01621459.2021.1909599
State	Published - 2022

Keywords

Bayesian
Compositional data
Estimating equations
Machine learning
Quantification

ASJC Scopus subject areas

Statistics and Probability
Statistics, Probability and Uncertainty

Access to Document

10.1080/01621459.2021.1909599

Cite this

@article{d884e32febd54b4db5b7bafba0a2b07d,

title = "Generalized Bayes Quantification Learning under Dataset Shift",

abstract = "Quantification learning is the task of prevalence estimation for a test population using predictions from a classifier trained on a different population. Quantification methods assume that the sensitivities and specificities of the classifier are either perfect or transportable from the training to the test population. These assumptions are inappropriate in the presence of dataset shift, when the misclassification rates in the training population are not representative of those for the test population. Quantification under dataset shift has been addressed only for single-class (categorical) predictions and assuming perfect knowledge of the true labels on a small subset of the test population. We propose generalized Bayes quantification learning (GBQL) that uses the entire compositional predictions from probabilistic classifiers and allows for uncertainty in true class labels for the limited labeled test data. Instead of positing a full model, we use a model-free Bayesian estimating equation approach to compositional data using Kullback–Leibler loss-functions based only on a first-moment assumption. The idea will be useful in Bayesian compositional data analysis in general as it is robust to different generating mechanisms for compositional data and allows 0{\textquoteright}s and 1{\textquoteright}s in the compositional outputs thereby including categorical outputs as a special case. We show how our method yields existing quantification approaches as special cases. Extension to an ensemble GBQL that uses predictions from multiple classifiers yielding inference robust to inclusion of a poor classifier is discussed. We outline a fast and efficient Gibbs sampler using a rounding and coarsening approximation to the loss functions. We establish posterior consistency, asymptotic normality and valid coverage of interval estimates from GBQL, which to our knowledge are the first theoretical results for a quantification approach in the presence of local labeled data. We also establish finite sample posterior concentration rate. Empirical performance of GBQL is demonstrated through simulations and analysis of real data with evident dataset shift. Supplementary materials for this article are available online.",

keywords = "Bayesian, Compositional data, Estimating equations, Machine learning, Quantification",

author = "Jacob Fiksel and Abhirup Datta and Agbessi Amouzou and Scott Zeger",

note = "Funding Information: The authors gratefully acknowledge Bill & Melinda Gates Foundation through the grant number OPP1163221 to Johns Hopkins University for the Countrywide Mortality Surveillance for Action project in Mozambique and Epi/Biostats of Aging Training Grant, Funded by National Institute of Aging T32AG000247 We thank the Editor, the Associate editor, and the reviewers for providing insightful feedback that improved the manuscript. We also thank Dr. Anirban Bhattacharya for helpful discussion about posterior concentration results. Funding Information: The authors gratefully acknowledge Bill & Melinda Gates Foundation through the grant number OPP1163221 to Johns Hopkins University for the Countrywide Mortality Surveillance for Action project in Mozambique and Epi/Biostats of Aging Training Grant, Funded by National Institute of Aging T32AG000247 Acknowledgment Publisher Copyright: {\textcopyright} 2021 The Author(s). Published with license by Taylor & Francis Group, LLC.",

year = "2022",

doi = "10.1080/01621459.2021.1909599",

language = "English (US)",

volume = "117",

pages = "2163--2181",

journal = "Journal of the American Statistical Association",

issn = "0162-1459",

publisher = "Taylor and Francis Ltd.",

number = "540",

}

TY - JOUR

T1 - Generalized Bayes Quantification Learning under Dataset Shift

AU - Fiksel, Jacob

AU - Datta, Abhirup

AU - Amouzou, Agbessi

AU - Zeger, Scott

N1 - Funding Information: The authors gratefully acknowledge Bill & Melinda Gates Foundation through the grant number OPP1163221 to Johns Hopkins University for the Countrywide Mortality Surveillance for Action project in Mozambique and Epi/Biostats of Aging Training Grant, Funded by National Institute of Aging T32AG000247 We thank the Editor, the Associate editor, and the reviewers for providing insightful feedback that improved the manuscript. We also thank Dr. Anirban Bhattacharya for helpful discussion about posterior concentration results. Funding Information: The authors gratefully acknowledge Bill & Melinda Gates Foundation through the grant number OPP1163221 to Johns Hopkins University for the Countrywide Mortality Surveillance for Action project in Mozambique and Epi/Biostats of Aging Training Grant, Funded by National Institute of Aging T32AG000247 Acknowledgment Publisher Copyright: © 2021 The Author(s). Published with license by Taylor & Francis Group, LLC.

PY - 2022

Y1 - 2022

N2 - Quantification learning is the task of prevalence estimation for a test population using predictions from a classifier trained on a different population. Quantification methods assume that the sensitivities and specificities of the classifier are either perfect or transportable from the training to the test population. These assumptions are inappropriate in the presence of dataset shift, when the misclassification rates in the training population are not representative of those for the test population. Quantification under dataset shift has been addressed only for single-class (categorical) predictions and assuming perfect knowledge of the true labels on a small subset of the test population. We propose generalized Bayes quantification learning (GBQL) that uses the entire compositional predictions from probabilistic classifiers and allows for uncertainty in true class labels for the limited labeled test data. Instead of positing a full model, we use a model-free Bayesian estimating equation approach to compositional data using Kullback–Leibler loss-functions based only on a first-moment assumption. The idea will be useful in Bayesian compositional data analysis in general as it is robust to different generating mechanisms for compositional data and allows 0’s and 1’s in the compositional outputs thereby including categorical outputs as a special case. We show how our method yields existing quantification approaches as special cases. Extension to an ensemble GBQL that uses predictions from multiple classifiers yielding inference robust to inclusion of a poor classifier is discussed. We outline a fast and efficient Gibbs sampler using a rounding and coarsening approximation to the loss functions. We establish posterior consistency, asymptotic normality and valid coverage of interval estimates from GBQL, which to our knowledge are the first theoretical results for a quantification approach in the presence of local labeled data. We also establish finite sample posterior concentration rate. Empirical performance of GBQL is demonstrated through simulations and analysis of real data with evident dataset shift. Supplementary materials for this article are available online.

AB - Quantification learning is the task of prevalence estimation for a test population using predictions from a classifier trained on a different population. Quantification methods assume that the sensitivities and specificities of the classifier are either perfect or transportable from the training to the test population. These assumptions are inappropriate in the presence of dataset shift, when the misclassification rates in the training population are not representative of those for the test population. Quantification under dataset shift has been addressed only for single-class (categorical) predictions and assuming perfect knowledge of the true labels on a small subset of the test population. We propose generalized Bayes quantification learning (GBQL) that uses the entire compositional predictions from probabilistic classifiers and allows for uncertainty in true class labels for the limited labeled test data. Instead of positing a full model, we use a model-free Bayesian estimating equation approach to compositional data using Kullback–Leibler loss-functions based only on a first-moment assumption. The idea will be useful in Bayesian compositional data analysis in general as it is robust to different generating mechanisms for compositional data and allows 0’s and 1’s in the compositional outputs thereby including categorical outputs as a special case. We show how our method yields existing quantification approaches as special cases. Extension to an ensemble GBQL that uses predictions from multiple classifiers yielding inference robust to inclusion of a poor classifier is discussed. We outline a fast and efficient Gibbs sampler using a rounding and coarsening approximation to the loss functions. We establish posterior consistency, asymptotic normality and valid coverage of interval estimates from GBQL, which to our knowledge are the first theoretical results for a quantification approach in the presence of local labeled data. We also establish finite sample posterior concentration rate. Empirical performance of GBQL is demonstrated through simulations and analysis of real data with evident dataset shift. Supplementary materials for this article are available online.

KW - Bayesian

KW - Compositional data

KW - Estimating equations

KW - Machine learning

KW - Quantification

UR - http://www.scopus.com/inward/record.url?scp=85104504156&partnerID=8YFLogxK

UR - http://www.scopus.com/inward/citedby.url?scp=85104504156&partnerID=8YFLogxK

U2 - 10.1080/01621459.2021.1909599

DO - 10.1080/01621459.2021.1909599

M3 - Article

AN - SCOPUS:85104504156

SN - 0162-1459

VL - 117

SP - 2163

EP - 2181

JO - Journal of the American Statistical Association

JF - Journal of the American Statistical Association

IS - 540

ER -

Generalized Bayes Quantification Learning under Dataset Shift

Abstract

Keywords

ASJC Scopus subject areas

Access to Document

Other files and links

Fingerprint

Cite this