Antimicrobial-resistance genomic surveillance
A completed provenance-first study of antimicrobial-resistance marker detection, assembly quality, genomic relatedness, and metadata-aware surveillance evidence.

Project definition
Problem statement
A detected resistance gene can be technically correct while the conclusion about phenotype, lineage, transmission, or prevalence is wrong.
The engineering problem is to preserve sequence, database, threshold, QC, relatedness, and metadata evidence as separate and reproducible layers.
Project objectives
- Extract a traceable eight-marker panel from the current AMRFinderPlus database.
- Generate 24 deterministic reduced assemblies with declared truth and planted lineages.
- Compare standard, high-coverage, and high-identity detector settings.
- Apply assembly QC and metadata eligibility before controlled aggregation.
- Compare exact canonical 21-mer relatedness with AMR-family profile similarity.
- Retain raw evidence, metrics, figures, tests, provenance, and claim boundaries.
Project structure
Project components
Reference panel
Stores symbols, families, accessions, lengths, database version, container digest, and sequence hashes.
Assembly generator
Builds full, partial, divergent, and reverse marker fixtures in eight planted lineages.
AMRFinderPlus runner
Runs the pinned detector under three identity and coverage settings and preserves raw output.
Evidence analysis
Calculates call recovery, assembly QC, genome similarity, profile similarity, clusters, and denominators.
Verification
Checks code, retained results, container execution, documents, references, figures, and claim limits.
Methodology
Project workflow
- 01Prepare references
Load the recorded eight-marker panel and its database provenance.
- 02Generate assemblies
Create 24 deterministic teaching assemblies and configuration-specific truth records.
- 03Run marker detection
Execute AMRFinderPlus under standard, high-coverage, and high-identity settings.
- 04Analyse evidence
Measure call recovery, QC, relatedness, AMR profiles, metadata completeness, and denominators.
- 05Review result
Inspect retained CSV, JSON, figures, raw calls, dashboard, tests, and report.
Demonstration scenario
The standard configuration recovers all 36 expected sample-family calls. Raising identity to 98 percent removes one partial vanA call measured at 97.64 percent. One planted high-N assembly is excluded by QC. The relatedness comparison reaches 99.28 percent pair accuracy while still showing that two unrelated lineages can share an AMR profile.
Engineering
Tools and method
- Tools
- The project uses Python, Biopython, NCBI AMRFinderPlus 4.2.7, Pandas, Matplotlib, Jupyter for subject analysis, simulation, and results.
- Sequence handling
- Python and Biopython validate FASTA data and create deterministic marker fixtures.
- AMR evidence
- NCBI AMRFinderPlus 4.2.7 and database 2026-08-07.1 provide versioned marker calls.
- Analysis
- Pandas and exact canonical 21-mer sets produce QC, threshold, relatedness, and metadata tables.
- Figures
- Matplotlib generates nine labelled figures directly from retained result files.
- Reproducibility
- Docker, exact dependency pins, checksums, tests, and a static offline dashboard preserve the experiment.
Testing
Evaluation
Evaluation measures
- True positive, false positive, false negative, precision, recall, and F1 by threshold configuration
- Exact AMR-family profiles for every controlled sample
- Total length, contigs, N50, GC, ambiguous bases, and QC eligibility
- Pairwise planted-lineage classification from canonical 21-mer similarity
- Unrelated-lineage profile matches and within-lineage profile changes
- Automated test coverage, dependency audit, document validation, and container execution
Project boundaries
- The 24 assemblies are controlled reduced teaching fixtures, not sequenced organisms.
- Detected markers do not establish clinical resistance or replace phenotypic susceptibility testing.
- The planted k-mer result does not establish transmission, outbreak membership, or a species threshold.
- Source and period proportions are controlled fixtures and are not prevalence.
- Real sequence or health data require supervisor approval, governance, organism-specific methods, and new validation.
Included
- 01Complete AMR genomic-surveillance Python workflow
- 02Twenty-four controlled reduced teaching assemblies from eight planted lineages
- 03Raw NCBI AMRFinderPlus results under three threshold configurations
- 04Assembly QC, AMR-profile, genomic-relatedness, and metadata result tables
- 05Nine generated result figures and three attributed literature figures
- 06Offline evidence dashboard
- 07Seventy-one automated tests with 99.31 percent statement coverage
- 08Complete project files, calculations, results, and analysis material in a private GitHub repository
- 0976-page project documentation in PDF and editable Word formats
- 1017-page setup and usage guide in PDF and editable Word formats
- 11Fifty annotated references
Project record
No information is collected on this page.
- Permanent project ID
- GP-BT-1EKR0B9
- Catalogued
- 21 Aug 2026
- Completed
- 25 Aug 2026
- Verified
- 25 Aug 2026
- Demonstration
- Included in repository
Handover
After purchase
- 01Payment is confirmed
The project is marked unavailable and cannot be purchased again.
- 02Repository access is granted
The buyer's submitted GitHub account receives access to the private repository.
- 03The purchase record is delivered
The certification sheet is prepared from the reviewed buyer details and sent privately by email.