← Back to project catalogue
GP-BT-0OOKW8MBiotechnologyReady

Single-Cell Trajectory Robustness Study

A completed single-cell biology study measuring how quality control, normalisation, batch correction, neighborhood size, and root-cell choice alter diffusion-pseudotime conclusions.

Single-Cell Trajectory Robustness Study project visual
GP-BT-0OOKW8M · Biotechnology
  • Python
  • Scanpy 1.12.3
  • AnnData
  • NumPy
  • Pandas
  • SciPy
  • Matplotlib
  • Jupyter

Project definition

Problem statement

A single-cell pseudotime plot can look convincing even when another reasonable filtering, normalisation, correction, graph, or root choice produces a different ordering.

The biotechnology problem is to measure that sensitivity and separate stable lineage evidence from conclusions that depend on the selected pipeline.

Project objectives

  • Validate and retain the released Paul15 mouse myeloid progenitor count matrix.
  • Compare three quality-control profiles, two normalisations, two batch-correction choices, and two neighborhood sizes.
  • Infer diffusion pseudotime from three biologically declared root cells.
  • Measure rank agreement, endpoint overlap, graph topology, batch mixing, cluster purity, and root sensitivity.
  • Check lineage direction with seven erythroid and myeloid marker genes.
  • Retain every result, figure, test, source, and interpretation boundary.

Project structure

Project components

01

Data validation

Checks the Paul15 file checksum, dimensions, integer counts, cell labels, and required marker genes.

02

Technical perturbation

Creates balanced synthetic batches and applies a controlled count-depth reduction to one batch.

03

Trajectory pipeline

Runs quality control, normalisation, feature selection, optional ComBat, PCA, neighbor graphs, diffusion maps, and DPT.

04

Robustness study

Builds 24 graph pipelines and 72 root-specific trajectories under one deterministic configuration.

05

Evidence and verification

Retains complete CSV and JSON results, figures, references, tests, and document checks.

Methodology

Project workflow

  1. 01
    Load released counts

    Validate the retained Paul15 matrix and published cell-state labels.

  2. 02
    Apply controlled perturbations

    Select the quality-control, normalisation, correction, and neighborhood combination.

  3. 03
    Infer trajectories

    Calculate diffusion pseudotime independently from the published, MEP, and GMP roots.

  4. 04
    Compare robustness

    Measure rank, endpoint, topology, batch, cluster, root, and marker evidence.

  5. 05
    Interpret the biology

    Identify conclusions that remain stable and state the limits of snapshot trajectory inference.

Demonstration scenario

The complete study finds a median reference agreement of 0.904690, but the weakest preprocessing and root combination falls to 0.315010. All log1p trajectories remain above 0.931501. ComBat raises mean balanced batch mixing from 0.823602 to 0.881327, while separate purity and topology evidence prevents mixing alone from being treated as success.

Engineering

Tools and method

Tools
The project uses Python, Scanpy 1.12.3, AnnData, NumPy, Pandas, SciPy, Matplotlib, Jupyter for subject analysis, simulation, and results.
Single-cell data
AnnData preserves the released count matrix, cell labels, gene names, and study metadata.
Scientific workflow
Scanpy implements normalisation, ComBat, PCA, neighbor graphs, diffusion maps, and diffusion pseudotime.
Statistics
NumPy, Pandas, and SciPy calculate rank agreement, endpoint overlap, graph summaries, and marker trends.
Evidence
CSV, JSON, PNG, SVG, PDF, and Word files retain the complete experiment and report.
Verification
Automated tests cover configuration, data integrity, processing branches, metrics, commands, and the complete study.

Testing

Evaluation

Evaluation measures

  • Reference Spearman agreement across all seventy-two trajectories
  • Early and late decile Jaccard overlap
  • Top intercluster-edge topology overlap
  • Balanced batch mixing and published-cluster purity
  • Minimum, mean, and maximum pairwise root agreement
  • Gata1, Klf1, Hba-a2, Mpo, Elane, Irf8, and Cebpa marker directions

Project boundaries

  • The study uses one released mouse dataset and one diffusion-pseudotime method.
  • Pseudotime is a computational order, not measured developmental time or direct lineage ancestry.
  • The synthetic batch represents controlled count thinning and is not a biological condition.
  • Marker correlations are descriptive and do not prove causal gene regulation.
  • Experimental validation, human generalisation, RNA velocity, fate probabilities, and clinical use are outside scope.

Included

  1. 01Complete single-cell trajectory analysis source code
  2. 02Validated Paul15 matrix with 2,730 cells and 3,451 genes
  3. 03Twenty-four graph pipelines and seventy-two trajectories
  4. 04Complete cell-level pseudotime, marker, endpoint, topology, and root evidence
  5. 05Fourteen generated result figures and three attributed biological images
  6. 06Thirty-four automated tests with 100 percent statement and branch coverage
  7. 07Complete project files, calculations, results, and analysis material in a private GitHub repository
  8. 0883-page project documentation in PDF and editable Word formats
  9. 098-page setup and usage guide in PDF and editable Word formats
  10. 10Fifty annotated references

Project record

No information is collected on this page.

Permanent project ID
GP-BT-0OOKW8M
Catalogued
21 Aug 2026
Completed
26 Aug 2026
Verified
26 Aug 2026
Demonstration
Included in repository

Handover

After purchase

  1. 01
    Payment is confirmed

    The project is marked unavailable and cannot be purchased again.

  2. 02
    Repository access is granted

    The buyer's submitted GitHub account receives access to the private repository.

  3. 03
    The purchase record is delivered

    The certification sheet is prepared from the reviewed buyer details and sent privately by email.