Deliberately under-performing on AI benchmarks to reduce regulatory attention.
Benchminning is the strategic practice of a frontier-AI lab or model vendor deliberately submitting a weaker-performing model version to a public benchmark than to commercial deployment — or training-time degrading the eval-targeted behaviour — in order to post a low enough score to stay below an emerging regulatory threshold. The term was coined by Matthew Berman on X on Aug 5, 2026 to describe this gap between measured and deployed capability and its use as a regulatory-evasion move. The verb is built by analogy with benchmaxxing (the established term for systematically over-optimising for a benchmark). Where benchmaxxing inflates, benchminning deflates; the strategic logic in both cases is treating the public score as an instrument to be managed rather than a measurement to be reported.
The mechanism is straightforward in settings where regulators have begun to publish capability-threshold numbers — the BIS/Commerce interim final rule on dual-use AI, EU AI Act tier-1 thresholds, and the UK AISI/SafeLab frameworks each have explicit capability lines that turn a benchmark score into a regulatory trigger. A model that scores 79 on a dangerous-capability benchmark is below threshold; the same model retrained with one extra RLHF pass and scoring 82 triggers reporting, third-party red-teaming, or export-control review. Benchminning is the response: tweak the submission, the system card, the fine-tune scaffold, or the model-and-checkpoints shipped to the eval harness so the public number sits below the regulator's line. The technique is harder to detect than classic benchmark contamination because there is no foreign artefact pulled into training — the model's weights themselves are simply under-tuned for the eval distribution.
Detecting benchminning requires looking past the benchmark number to leading indicators: the gap between public-eval scores and internal red-team scores, divergence between evaluation-platform-reported numbers and end-user-observed behaviour, drift between the submitted model's downstream benchmarks and its sibling models in the same release cycle, and unusual variance in evaluation repeats. Public evaluation pipelines operated by AISI / Apollo / METR have started instrumenting this explicitly — the same way Stanford CRFM has instrumented data-contamination detection since 2023.
The strategic cousins worth holding apart from benchminning are benchmaxxing (inflating scores, opposite direction), safetywashing (cosmetic safety claims about a model's deployment rather than its benchmarks), capability elucidation (gaining adversarial insights into model behaviour, neutral), and regulatory arbitrage (broader pattern of routing deployment through the lowest-restriction jurisdiction). Benchminning narrows the broader regulatory-arbitrage concept to the specific public-benchmark instrument; safetywashing overlaps on the cosmetic dimension but typically applies to system-card claims, not benchmark numbers. Tradeoff in the term itself: the -minning construction reads at first glance as if it shares an etymology with cryptocurrency mining, which it does not; precision over readability favours keeping the coined form as Berman named it. Open question whether the term will stick in the academic eval-research literature or stay confined to X — current evidence (4.9K views, 1 day old) suggests medium-threshold adoption, comparable to benchmaxxing's 2023 trajectory.
Signals turns a topic into a sourced research record you can inspect and rerun. Your first scan is free, and this one starts with Benchminning already loaded, so edit it or scan as is.