Skip to results
MLSift
← Feed
Theory & OptimizationTransformer2609.01311

One-Layer Transformer Provably Learns Multiclass One-Nearest Neighbor in Context

Skanda Athreya, Yutong Wang

cs.LG stat.ML

Abstract

We extend recent work establishing an equivalence between one-layer transformers and nearest-neighbor classifiers in the binary setting to the multiclass case. By leveraging the simplex encoding, we show that one-layer transformers with an argmax classification head behave identically to a one-nearest-neighbor classifier in the multiclass setting. This closes a gap left by prior work, whose multiclass result relied on a non-standard rounding-based approach rather than the typical argmax head used in practice.

Topics

Classified with taxonomy v2 on Wed, 2 Sept 2026.

The PDF is 1–3 MB. Open it in your browser's viewer, or load it here.

Open PDF