Benchmark contamination and evaluation leakage in foundation model testing
A completed Data and AI engineering study of benchmark contamination, evaluation leakage, detection evidence, reporting controls, and validity threats.

Project definition
Problem statement
Foundation model benchmarks can become unreliable when test items, task templates, metadata, solution traces, repeated feedback, or evaluation-time retrieval influence model development or scoring.
The engineering problem is to compare what different detectors can observe under realistic access limits, then show how layered controls change residual score inflation and false assurance.
Project objectives
- Define eight distinct contamination and evaluation leakage mechanisms.
- Compare eight detector families across four access regimes.
- Evaluate eight control profiles at four contamination prevalence positions.
- Measure detected leakage, undetected score inflation, audit cost, and false assurance.
- Quantify case-mix uncertainty with fixed-seed bootstrap analysis.
- Produce an evaluation-integrity framework with traceable reporting fields.
Project structure
Project components
Leakage taxonomy
Separates exact overlap, paraphrase, task-template, metadata, solution-trace, repeated tuning, and retrieval pathways.
Detector laboratory
Models corpus, n-gram, semantic, likelihood, membership, differential, canary, and trace-audit evidence.
Access model
Tests which evidence can be obtained under public-only, evaluator, developer, and full-audit access.
Factorial engine
Evaluates all 1,024 combinations with 2,000 fixed-seed synthetic decisions per cell.
Uncertainty analysis
Uses 10,000 bootstrap samples per profile to retain central and interval positions.
Integrity framework
Combines preventative controls, detectors, audit records, and reporting gates without hiding residual uncertainty.
Methodology
Project workflow
- 01Declare
Load the leakage, detector, access, control, and prevalence evidence positions.
- 02Gate by access
Determine which detectors are usable and how sensitive they are under each evidence regime.
- 03Simulate
Generate fixed-seed detection and score-inflation decisions for every factorial cell.
- 04Summarise
Calculate integrity, undetected inflation, audit cost, false assurance, and mechanism-level results.
- 05Test uncertainty
Bootstrap the retained case mix and compare profile sensitivity.
- 06Report
Trace every claim to equations, retained evidence, source annotations, and explicit validity limits.
Demonstration scenario
Run the complete study, compare a minimal public-only check with a layered full-audit profile, inspect which leakage mechanisms remain difficult to observe, then trace one result through detector access, residual inflation, bootstrap uncertainty, and the final evaluation-integrity card.
Engineering
Tools and method
- Study configuration
- A versioned JSON file defines all mechanisms, detectors, profiles, regimes, and evidence positions.
- Numerical analysis
- NumPy and pandas implement the factorial study, seeded decisions, summaries, and bootstrap intervals.
- Figures
- Matplotlib creates twelve labelled mechanism, access, cost, uncertainty, and decision figures.
- Documentation
- The build produces editable Word and fixed PDF report and guide files with contents, figure, and table lists.
- Reproducibility
- Automated tests, repository validation, fixed seeds, pinned dependencies, and Docker reproduce the retained study.
Testing
Evaluation
Evaluation measures
- Detected leakage and residual undetected leakage by mechanism
- Undetected score inflation at four contamination prevalence positions
- Detector sensitivity and specificity across four access regimes
- Integrity score, false-assurance index, and normalized audit cost by control profile
- Ten-thousand-sample bootstrap intervals for each profile
- Lifecycle-stage coverage and layered-profile mechanism outcomes
- Nine automated tests and twelve reproducible result figures
Project boundaries
- All numerical inputs are normalized evidence positions and synthetic study assumptions, not measured contamination rates for named models, vendors, or benchmarks.
- Detector sensitivity and specificity depend on data access, threshold choice, benchmark structure, and the real leakage mechanism.
- A low detected rate does not prove a clean training corpus or an uncontaminated evaluation process.
- The framework supports evaluation design and audit planning but does not certify a foundation model as safe, unbiased, or generally capable.
- No information collected.
Included
- 01Eight benchmark leakage mechanisms and eight detector families
- 02Four evaluation access regimes and eight layered control profiles
- 031,024 factorial study cells and 2,048,000 synthetic decisions
- 0410,000 bootstrap samples for every control profile
- 05Retained CSV and JSON evidence with twelve generated result figures
- 06One attributed NIST literature figure and complete provenance
- 07Nine automated tests and clean-container reproduction
- 08Complete project files, calculations, results, and analysis in a private GitHub repository
- 0973-page project documentation in PDF and editable Word formats
- 1017-page setup and usage guide in PDF and editable Word formats
- 1163 annotated references with a complete source matrix
Project record
No information is collected on this page.
- Permanent project ID
- GP-DA-1O4NTAF
- Catalogued
- 21 Aug 2026
- Completed
- 30 Aug 2026
- Verified
- 30 Aug 2026
- Demonstration
- Included in repository
Handover
After purchase
- 01Payment is confirmed
The project is marked unavailable and cannot be purchased again.
- 02Repository access is granted
The buyer's submitted GitHub account receives access to the private repository.
- 03The purchase record is delivered
The certification sheet is prepared from the reviewed buyer details and sent privately by email.