Skip to results
MLSift
← Feed
routineAI Safety, Security & AlignmentSafeguard-Conditioned Uplift2607.13039

Safeguard-Conditioned Uplift: Measuring Utility-Risk Frontiers for Dual-Use Biology Assistants

Dipesh Tharu Mahato

cs.CY cs.AI

Abstract

A refusal rate neither identifies which component intervened nor measures its burden on legitimate users. This paper evaluates safeguards for dual-use biology assistants at the action and answer levels. The framework reconstructs the access path, separates provider refusals from downstream actions, and selects thresholds under an intervention budget. A frozen fresh-generation study satisfies its criterion on Claude Opus 4.5, but both passing configurations share one upstream provider effect; none passes on Gemini 2.5 Flash. The fixed Opus policies retain positive selectivity on 104 previously unused released-label pairs, but both fail a 20\% matched-benign constraint. At the answer level, no joint-scoring verifier qualifies on a response-disjoint 7,200-judgment holdout. A fresh 8,640-judgment factorial experiment finds separate gains from criterion isolation and ordinal representation, with a positive interaction between them; requiring explicit localization lowers aggregate accuracy under a strict no-repair schema. The evidence supports prospective action-level selectivity and identifies verifier interface effects, but not calibrated selective access, verified content removal, or biological-risk reduction.

Topics

Classified with taxonomy v2 on Sat, 5 Sept 2026.

The PDF is 1–3 MB. Open it in your browser's viewer, or load it here.

Open PDF