← Back to project catalogue
GP-BT-1EKR0B9BiotechnologyReady

Antimicrobial-resistance genomic surveillance

A completed provenance-first study of antimicrobial-resistance marker detection, assembly quality, genomic relatedness, and metadata-aware surveillance evidence.

Antimicrobial-resistance genomic surveillance project visual
GP-BT-1EKR0B9 · Biotechnology
  • Python
  • Biopython
  • NCBI AMRFinderPlus 4.2.7
  • Pandas
  • Matplotlib
  • Jupyter

Project definition

Problem statement

A detected resistance gene can be technically correct while the conclusion about phenotype, lineage, transmission, or prevalence is wrong.

The engineering problem is to preserve sequence, database, threshold, QC, relatedness, and metadata evidence as separate and reproducible layers.

Project objectives

  • Extract a traceable eight-marker panel from the current AMRFinderPlus database.
  • Generate 24 deterministic reduced assemblies with declared truth and planted lineages.
  • Compare standard, high-coverage, and high-identity detector settings.
  • Apply assembly QC and metadata eligibility before controlled aggregation.
  • Compare exact canonical 21-mer relatedness with AMR-family profile similarity.
  • Retain raw evidence, metrics, figures, tests, provenance, and claim boundaries.

Project structure

Project components

01

Reference panel

Stores symbols, families, accessions, lengths, database version, container digest, and sequence hashes.

02

Assembly generator

Builds full, partial, divergent, and reverse marker fixtures in eight planted lineages.

03

AMRFinderPlus runner

Runs the pinned detector under three identity and coverage settings and preserves raw output.

04

Evidence analysis

Calculates call recovery, assembly QC, genome similarity, profile similarity, clusters, and denominators.

05

Verification

Checks code, retained results, container execution, documents, references, figures, and claim limits.

Methodology

Project workflow

  1. 01
    Prepare references

    Load the recorded eight-marker panel and its database provenance.

  2. 02
    Generate assemblies

    Create 24 deterministic teaching assemblies and configuration-specific truth records.

  3. 03
    Run marker detection

    Execute AMRFinderPlus under standard, high-coverage, and high-identity settings.

  4. 04
    Analyse evidence

    Measure call recovery, QC, relatedness, AMR profiles, metadata completeness, and denominators.

  5. 05
    Review result

    Inspect retained CSV, JSON, figures, raw calls, dashboard, tests, and report.

Demonstration scenario

The standard configuration recovers all 36 expected sample-family calls. Raising identity to 98 percent removes one partial vanA call measured at 97.64 percent. One planted high-N assembly is excluded by QC. The relatedness comparison reaches 99.28 percent pair accuracy while still showing that two unrelated lineages can share an AMR profile.

Engineering

Tools and method

Tools
The project uses Python, Biopython, NCBI AMRFinderPlus 4.2.7, Pandas, Matplotlib, Jupyter for subject analysis, simulation, and results.
Sequence handling
Python and Biopython validate FASTA data and create deterministic marker fixtures.
AMR evidence
NCBI AMRFinderPlus 4.2.7 and database 2026-08-07.1 provide versioned marker calls.
Analysis
Pandas and exact canonical 21-mer sets produce QC, threshold, relatedness, and metadata tables.
Figures
Matplotlib generates nine labelled figures directly from retained result files.
Reproducibility
Docker, exact dependency pins, checksums, tests, and a static offline dashboard preserve the experiment.

Testing

Evaluation

Evaluation measures

  • True positive, false positive, false negative, precision, recall, and F1 by threshold configuration
  • Exact AMR-family profiles for every controlled sample
  • Total length, contigs, N50, GC, ambiguous bases, and QC eligibility
  • Pairwise planted-lineage classification from canonical 21-mer similarity
  • Unrelated-lineage profile matches and within-lineage profile changes
  • Automated test coverage, dependency audit, document validation, and container execution

Project boundaries

  • The 24 assemblies are controlled reduced teaching fixtures, not sequenced organisms.
  • Detected markers do not establish clinical resistance or replace phenotypic susceptibility testing.
  • The planted k-mer result does not establish transmission, outbreak membership, or a species threshold.
  • Source and period proportions are controlled fixtures and are not prevalence.
  • Real sequence or health data require supervisor approval, governance, organism-specific methods, and new validation.

Included

  1. 01Complete AMR genomic-surveillance Python workflow
  2. 02Twenty-four controlled reduced teaching assemblies from eight planted lineages
  3. 03Raw NCBI AMRFinderPlus results under three threshold configurations
  4. 04Assembly QC, AMR-profile, genomic-relatedness, and metadata result tables
  5. 05Nine generated result figures and three attributed literature figures
  6. 06Offline evidence dashboard
  7. 07Seventy-one automated tests with 99.31 percent statement coverage
  8. 08Complete project files, calculations, results, and analysis material in a private GitHub repository
  9. 0976-page project documentation in PDF and editable Word formats
  10. 1017-page setup and usage guide in PDF and editable Word formats
  11. 11Fifty annotated references

Project record

No information is collected on this page.

Permanent project ID
GP-BT-1EKR0B9
Catalogued
21 Aug 2026
Completed
25 Aug 2026
Verified
25 Aug 2026
Demonstration
Included in repository

Handover

After purchase

  1. 01
    Payment is confirmed

    The project is marked unavailable and cannot be purchased again.

  2. 02
    Repository access is granted

    The buyer's submitted GitHub account receives access to the private repository.

  3. 03
    The purchase record is delivered

    The certification sheet is prepared from the reviewed buyer details and sent privately by email.