---
title: Benchminning
type: vocabulary
url: "https://www.envisioning.com/vocab/benchminning"
summary: Deliberately under-performing on AI benchmarks to reduce regulatory attention.
year: 2026
generality: 0.40
---

# Benchminning

Deliberately under-performing on AI benchmarks to reduce regulatory attention.
**Benchminning** is the strategic practice of a frontier-AI lab or model vendor deliberately submitting a weaker-performing model version to a public benchmark than to commercial deployment — or training-time degrading the eval-targeted behaviour — in order to post a low enough score to stay below an emerging regulatory threshold. The term was coined by Matthew Berman on X on Aug 5, 2026 to describe this gap between *measured* and *deployed* capability and its use as a regulatory-evasion move. The verb is built by analogy with *benchmaxxing* (the established term for systematically over-optimising for a benchmark). Where benchmaxxing inflates, benchminning deflates; the strategic logic in both cases is treating the public score as an instrument to be managed rather than a measurement to be reported.

The mechanism is straightforward in settings where regulators have begun to publish *capability-threshold* numbers — the BIS/Commerce interim final rule on dual-use AI, EU AI Act tier-1 thresholds, and the UK AISI/SafeLab frameworks each have explicit capability lines that turn a benchmark score into a regulatory trigger. A model that scores 79 on a dangerous-capability benchmark is below threshold; the same model retrained with one extra RLHF pass and scoring 82 triggers reporting, third-party red-teaming, or export-control review. Benchminning is the response: tweak the submission, the system card, the fine-tune scaffold, or the model-and-checkpoints shipped to the eval harness so the public number sits below the regulator's line. The technique is harder to detect than classic benchmark contamination because there is no *foreign* artefact pulled into training — the model's weights themselves are simply under-tuned for the eval distribution.

Detecting benchminning requires looking *past* the benchmark number to leading indicators: the gap between public-eval scores and internal red-team scores, divergence between evaluation-platform-reported numbers and end-user-observed behaviour, drift between the submitted model's downstream benchmarks and its sibling models in the same release cycle, and unusual variance in evaluation repeats. Public evaluation pipelines operated by AISI / Apollo / METR have started instrumenting this explicitly — the same way Stanford CRFM has instrumented data-contamination detection since 2023.

The strategic cousins worth holding apart from benchminning are **benchmaxxing** (inflating scores, opposite direction), **safetywashing** (cosmetic safety claims about a model's deployment rather than its benchmarks), **capability elucidation** (gaining adversarial insights into model behaviour, neutral), and **regulatory arbitrage** (broader pattern of routing deployment through the lowest-restriction jurisdiction). Benchminning narrows the broader regulatory-arbitrage concept to the specific public-benchmark instrument; safetywashing overlaps on the cosmetic dimension but typically applies to system-card claims, not benchmark numbers. Tradeoff in the term itself: the *-minning* construction reads at first glance as if it shares an etymology with cryptocurrency mining, which it does not; precision over readability favours keeping the coined form as Berman named it. Open question whether the term will stick in the academic eval-research literature or stay confined to X — current evidence (4.9K views, 1 day old) suggests medium-threshold adoption, comparable to *benchmaxxing*'s 2023 trajectory.

---
Source: Envisioning — Technology Research Institute (https://www.envisioning.com/vocab/benchminning)
