← Back to project catalogue
GP-DA-1O4NTAFData and AIReady

Benchmark contamination and evaluation leakage in foundation model testing

A completed Data and AI engineering study of benchmark contamination, evaluation leakage, detection evidence, reporting controls, and validity threats.

Benchmark contamination and evaluation leakage in foundation model testing project visual
GP-DA-1O4NTAF · Data and AI
  • Python 3.12
  • NumPy
  • pandas
  • Matplotlib
  • Docker

Project definition

Problem statement

Foundation model benchmarks can become unreliable when test items, task templates, metadata, solution traces, repeated feedback, or evaluation-time retrieval influence model development or scoring.

The engineering problem is to compare what different detectors can observe under realistic access limits, then show how layered controls change residual score inflation and false assurance.

Project objectives

  • Define eight distinct contamination and evaluation leakage mechanisms.
  • Compare eight detector families across four access regimes.
  • Evaluate eight control profiles at four contamination prevalence positions.
  • Measure detected leakage, undetected score inflation, audit cost, and false assurance.
  • Quantify case-mix uncertainty with fixed-seed bootstrap analysis.
  • Produce an evaluation-integrity framework with traceable reporting fields.

Project structure

Project components

01

Leakage taxonomy

Separates exact overlap, paraphrase, task-template, metadata, solution-trace, repeated tuning, and retrieval pathways.

02

Detector laboratory

Models corpus, n-gram, semantic, likelihood, membership, differential, canary, and trace-audit evidence.

03

Access model

Tests which evidence can be obtained under public-only, evaluator, developer, and full-audit access.

04

Factorial engine

Evaluates all 1,024 combinations with 2,000 fixed-seed synthetic decisions per cell.

05

Uncertainty analysis

Uses 10,000 bootstrap samples per profile to retain central and interval positions.

06

Integrity framework

Combines preventative controls, detectors, audit records, and reporting gates without hiding residual uncertainty.

Methodology

Project workflow

  1. 01
    Declare

    Load the leakage, detector, access, control, and prevalence evidence positions.

  2. 02
    Gate by access

    Determine which detectors are usable and how sensitive they are under each evidence regime.

  3. 03
    Simulate

    Generate fixed-seed detection and score-inflation decisions for every factorial cell.

  4. 04
    Summarise

    Calculate integrity, undetected inflation, audit cost, false assurance, and mechanism-level results.

  5. 05
    Test uncertainty

    Bootstrap the retained case mix and compare profile sensitivity.

  6. 06
    Report

    Trace every claim to equations, retained evidence, source annotations, and explicit validity limits.

Demonstration scenario

Run the complete study, compare a minimal public-only check with a layered full-audit profile, inspect which leakage mechanisms remain difficult to observe, then trace one result through detector access, residual inflation, bootstrap uncertainty, and the final evaluation-integrity card.

Engineering

Tools and method

Study configuration
A versioned JSON file defines all mechanisms, detectors, profiles, regimes, and evidence positions.
Numerical analysis
NumPy and pandas implement the factorial study, seeded decisions, summaries, and bootstrap intervals.
Figures
Matplotlib creates twelve labelled mechanism, access, cost, uncertainty, and decision figures.
Documentation
The build produces editable Word and fixed PDF report and guide files with contents, figure, and table lists.
Reproducibility
Automated tests, repository validation, fixed seeds, pinned dependencies, and Docker reproduce the retained study.

Testing

Evaluation

Evaluation measures

  • Detected leakage and residual undetected leakage by mechanism
  • Undetected score inflation at four contamination prevalence positions
  • Detector sensitivity and specificity across four access regimes
  • Integrity score, false-assurance index, and normalized audit cost by control profile
  • Ten-thousand-sample bootstrap intervals for each profile
  • Lifecycle-stage coverage and layered-profile mechanism outcomes
  • Nine automated tests and twelve reproducible result figures

Project boundaries

  • All numerical inputs are normalized evidence positions and synthetic study assumptions, not measured contamination rates for named models, vendors, or benchmarks.
  • Detector sensitivity and specificity depend on data access, threshold choice, benchmark structure, and the real leakage mechanism.
  • A low detected rate does not prove a clean training corpus or an uncontaminated evaluation process.
  • The framework supports evaluation design and audit planning but does not certify a foundation model as safe, unbiased, or generally capable.
  • No information collected.

Included

  1. 01Eight benchmark leakage mechanisms and eight detector families
  2. 02Four evaluation access regimes and eight layered control profiles
  3. 031,024 factorial study cells and 2,048,000 synthetic decisions
  4. 0410,000 bootstrap samples for every control profile
  5. 05Retained CSV and JSON evidence with twelve generated result figures
  6. 06One attributed NIST literature figure and complete provenance
  7. 07Nine automated tests and clean-container reproduction
  8. 08Complete project files, calculations, results, and analysis in a private GitHub repository
  9. 0973-page project documentation in PDF and editable Word formats
  10. 1017-page setup and usage guide in PDF and editable Word formats
  11. 1163 annotated references with a complete source matrix

Project record

No information is collected on this page.

Permanent project ID
GP-DA-1O4NTAF
Catalogued
21 Aug 2026
Completed
30 Aug 2026
Verified
30 Aug 2026
Demonstration
Included in repository

Handover

After purchase

  1. 01
    Payment is confirmed

    The project is marked unavailable and cannot be purchased again.

  2. 02
    Repository access is granted

    The buyer's submitted GitHub account receives access to the private repository.

  3. 03
    The purchase record is delivered

    The certification sheet is prepared from the reviewed buyer details and sent privately by email.