Synthetic data utility, privacy and fairness trade-off assessment
A completed Data and AI engineering study that keeps statistical fidelity, downstream utility, inferential validity, disclosure resistance, formal privacy, subgroup representation and fairness as separate evidence questions.

Project definition
Problem statement
Synthetic data can look realistic while losing important relationships, invalidating ordinary statistical inference, exposing training information, or affecting small groups unevenly.
The engineering problem is to select evidence according to the intended use and prevent an attractive average score from hiding a failed privacy, fairness, subgroup or inference requirement.
Project objectives
- Define ten representative tabular synthesis strategies and twelve evidence dimensions.
- Separate statistical fidelity, predictive utility and inferential validity.
- Evaluate membership, attribute, singling-out and formal privacy evidence.
- Measure subgroup representation, rare-event coverage and fairness stability.
- Compare balanced, privacy-first, fairness-first and inference-first decisions.
- Translate the assessment into validation gates, workflow risks and release documentation.
Project structure
Project components
Strategy register
Declares statistical, GAN, VAE, diffusion, differential-privacy, causal-fairness and hybrid synthesis archetypes.
Factorial assessment
Crosses ten strategies with eight drivers, four contexts and four evidence cases.
Decision gates
Prevents formal privacy, inference, subgroup and fairness weaknesses from being averaged away.
Uncertainty analysis
Runs 20,000 fixed-seed correlated evidence draws for every strategy.
Risk and assurance
Ranks 154 workflow FMEA cells and maps thirteen release requirements to evidence positions.
Evidence package
Retains every CSV, JSON, figure, source annotation and document build input.
Methodology
Project workflow
- 01Declare the purpose
Identify supported tasks, estimands, protected units, groups and prohibited uses before selecting metrics.
- 02Load the evidence positions
Review the transparent strategy, driver, context, evidence-case and scenario registers.
- 03Run the factorial study
Generate all 1,280 deterministic decision cells and hard-gate outcomes.
- 04Test uncertainty
Compare percentile ranges, scenario dependence and dimension sensitivity.
- 05Prioritise validation
Use FMEA and the assurance crosswalk to identify missing evidence.
- 06Record the decision
Document supported uses, failed gates, residual risks, version hashes and the accountable reviewer.
Demonstration scenario
Run the complete study, compare a high-fidelity non-private generator with marginal-based differential privacy and a fairness-aware strategy, inspect which hard gates fail under an informed attack or severe imbalance, then trace the recommendation through uncertainty, FMEA and the release decision card.
Engineering
Tools and method
- Study configuration
- Versioned Python data structures declare every strategy, dimension, weight, context and validation requirement.
- Numerical analysis
- NumPy and pandas implement deterministic cases, uncertainty, sensitivity, FMEA and assurance tables.
- Figures
- Matplotlib creates twelve labelled strategy, driver, context, uncertainty and risk figures.
- Documentation
- The build produces editable Word and fixed PDF report and guide files with contents, figure and table lists.
- Reproducibility
- Automated tests, pinned dependencies, repository validation and Docker reproduce the retained study.
Testing
Evaluation
Evaluation measures
- Statistical fidelity, task utility and inferential validity positions
- Subgroup fidelity, rare-event coverage and fairness stability
- Membership and attribute disclosure resistance
- Formal privacy, reproducibility, computing effort and governance clarity
- Gate-pass rates across drivers, contexts and evidence cases
- Balanced, privacy-first, fairness-first and inference-first scenarios
- Twenty-thousand-draw uncertainty intervals and one-at-a-time sensitivity
- Ten automated tests and twelve reproducible result figures
Project boundaries
- All numerical inputs are literature-informed normalized evidence positions, not measured performance from a named generator, dataset, population or protected group.
- The project does not establish a differential-privacy guarantee or prove resistance to an untested adversary.
- Predictive utility does not establish valid statistical inference, and population fidelity does not establish subgroup fairness.
- A real release requires authorized data, independent testing, explicit privacy parameters, attack evaluation, group analysis and accountable approval.
- No information collected.
Included
- 01Ten statistical, adversarial, variational, diffusion, private and fairness-aware strategy archetypes
- 02Twelve evidence dimensions and eight non-compensable trade-off drivers
- 031,280 deterministic assessment cases
- 0420,000 fixed-seed uncertainty draws per strategy
- 05Four application contexts and four decision scenarios
- 06Workflow FMEA, assurance crosswalk and twelve labelled result figures
- 07One attributed NIST literature figure with complete provenance
- 08Ten automated tests, dependency audit and clean Docker reproduction
- 09Complete project files, calculations, results, and analysis in a private GitHub repository
- 1093-page project documentation in PDF and editable Word formats
- 1116-page setup and usage guide in PDF and editable Word formats
- 1250 annotated references with a complete source matrix
Project record
No information is collected on this page.
- Permanent project ID
- GP-DA-0YUA6AY
- Catalogued
- 21 Aug 2026
- Completed
- 03 Sept 2026
- Verified
- 03 Sept 2026
- Demonstration
- Included in repository
Handover
After purchase
- 01Payment is confirmed
The project is marked unavailable and cannot be purchased again.
- 02Repository access is granted
The buyer's submitted GitHub account receives access to the private repository.
- 03The purchase record is delivered
The certification sheet is prepared from the reviewed buyer details and sent privately by email.