Separating Voice from Age in COPD Screening
Voice has been proposed as a low-cost screening signal for chronic obstructive pulmonary disease (COPD). COPD is strongly age-associated and voice changes with age, thus such results admit a trivial alternative explanation. We re-evaluate a public sustained-phonation corpus ($1246$ recordings, $68$ participants) under a strictly participant-level protocol. We therefore evaluate on repeatedly drawn age-matched cohorts and report the discrimination achieved by the confounders themselves on those same cohorts. Where raw (unmodelled) age ($0.510$ $[0.469, 0.551]$) and raw gender ($0.479$) are both measured at chance, acoustic models excluding age retain ROC-AUC $0.717$ $[0.552, 0.859]$ and average precision $0.747$ $[0.581, 0.892]$ against a one-to-one baseline of $0.5$, whereas models containing age fall to $0.531$--$0.679$. The separation is reproduced by two further learners with fixed hyperparameters. Two findings have broader methodological implications: models trained with age transfer less effectively to an age-balanced target cohort than otherwise identical models trained without age, and fourteen classical voice-quality and perturbation measures achieve comparable discrimination to a $55$-dimensional combined representation. We conclude that a non-age acoustic signal is present, that confounding by recording conditions cannot be excluded from the released features, and that the evaluation protocol in standard use cannot distinguish these possibilities.