The Metanym Game: An LLM Benchmark Without Ground Truth That Rises With the Models It Measures
David Nordfors
Abstract
We present evidence that analogy is at the core of LLM intelligence. In our benchmark, LLMs compete in generating sets of analogous statements and rate each other's sets on their own understandings of factual correctness, beauty, intelligence, distinctness, length, and structural diversity. Nothing enters from outside: the only given is the game rules; every item is generated in play; the scores come from the players' ratings alone. Ground truth is replaced by the SVD of the factual rating matrix, which scores players as generators and judges at once -- to our knowledge the first eigen-equation that judges the judges for an LLM council-of-peers. For subjective criteria like beauty, judges are weighted by their rating consistency. The best generators turn out to be middling judges. GPQA Diamond -- difficult multiple-choice questions written by human experts -- could not be more different in method, yet the two benchmarks correlate at Pearson $r = 0.97$, 95% CI [0.92, 0.99]; no leakage could be found. A council of the five best issues the official ratings; its contestable seats let the benchmark scale to any number of players and rise with the models it measures -- a candidate steering signal for self-improving AI. Playing interweaves at least eight constructs of intelligence; the total scores the broad composite, the components allow reductionistic analysis. Every number recomputes from a released package at https://github.com/dnordfors/metanym-game-paper
Topics
Classified with taxonomy v2 on Sat, 5 Sept 2026.