Xiaohua Douglas Zhangstat.AP cs.AI q-bio.QM stat.ML
High-throughput screening (HTS) assays are central to early-stage drug discovery but are often limited by extreme data sparsity, as primary screens typically use only a single replicate per test substance. This sparsity makes conventional machine-learning performance metrics, such as sensitivity, specificity, and area under the receiver operating characteristic curve (AUROC), difficult to estimate empirically because they require adequately sized labeled samples. Here, we introduce a model-based framework that derives these classification metrics from the strictly standardized mean difference (SSMD), a well-established HTS effect-size parameter. Under a Gaussian equal-variance assumption, we derive closed-form relationships linking SSMD to Youden-optimal sensitivity and specificity, and sensitivity at a preset specificity, yielding explicit estimators and exact confidence intervals from the noncentral t-distribution, even under single-replicate designs. Unlike classical statistical power, which approaches 1 as sample size grows regardless of how small the true non-zero difference between group means is, the SSMD-derived sensitivity converges to a finite population value that reflects the true degree of separation between two groups, making it a more meaningful and stable performance measure for hit selection. We demonstrate the utility of this framework in a hepatitis C virus primary siRNA screen comprising approximately 22,000 single-replicate measurements, showing that SSMD, AUROC, and sensitivity-based thresholds yield equivalent and interpretable hit sets. This work bridges classical HTS statistics and machine-learning evaluation theory, providing a statistically principled, reproducible way to estimate classification performance in ultra-low-replication screening workflows.
While machine-learned interatomic potentials (MLIPs) accelerate phonon dispersion calculations, merely identifying dynamical instabilities in computationally predicted materials is insufficient; automated pathways to resolve them are required. We introduce VibroML, an open-source Python toolkit driven by foundational MLIPs that shifts the paradigm from stability verification to automated structural remediation. VibroML employs an energy-guided genetic algorithm that vastly outperforms traditional soft-mode following, efficiently navigating the potential energy surface to uncover diverse, dynamically stable polymorphs. As 0 K harmonic stability does not guarantee macroscopic viability, an automated molecular dynamics workflow evaluates finite-temperature structural retention. VibroML also couples with ProtoCSP, our combinatorial structure prediction engine, to stabilize frustrated crystal topologies via targeted alloying, successfully rescuing functional perovskite networks like Cs$_2$KInI$_6$ and KTaSe$_3$. Demonstrating broader applicability, we mined the Alexandria database -- where ~50% of quaternary and 99.5% of quinary elemental combinations lack any structural entries -- to identify thousands of abandoned, high-symmetry stoichiometries. Deploying ProtoCSP's "cold start" retrieval and VibroML's evolutionary search on a sample, we successfully identified dynamically stable low-symmetry candidates. Through integrated structural remediation, thermal validation, and systematic compositional exploration, VibroML enables a comprehensive deep-screening approach, yielding physically sound structural propositions that far surpass standard high-throughput workflows.