Critical-mineral compositional anomaly study
A completed geology study using compositional data analysis, robust regional and local backgrounds, field duplicates, and uncertainty to screen multielement stream-sediment anomalies.

Software compatibility
The completed workflow uses the July 2026 USGS Alaska sediment reanalysis release, its retained checksums, and the pinned Python packages. Replacing the source data requires a new provenance, quality, and sensitivity review.
Project definition
Problem statement
Regional stream-sediment chemistry reflects bedrock, weathering, drainage transport, grain size, analytical limits, and possible mineral-system signals at the same time.
The geology problem is to identify unusual multielement compositions without treating raw concentration thresholds or statistical rank as proof of mineralisation.
Project objectives
- Prepare a documented Yukon-Tanana study population from the July 2026 USGS release.
- Represent twenty-five selected elements with centred and isometric log-ratio coordinates.
- Combine a robust regional model with a 35-neighbour local-background model.
- Evaluate matched field duplicates, censored results, spatial autocorrelation, uncertainty, and model sensitivity.
- Retain traceable candidate evidence, tables, figures, tests, documentation, and source provenance.
Project structure
Project components
Source preparation
Validates checksums, selects the regional population, aligns field duplicates, and decodes reporting-limit flags.
Compositional model
Calculates closure, CLR, ILR, Aitchison distance, variation matrices, and interpretable elemental balances.
Anomaly model
Combines robust global Mahalanobis distance with departure from a 35-neighbour local background.
Validation
Measures duplicate agreement, uncertainty stability, model-choice sensitivity, and Moran spatial autocorrelation.
Evidence builder
Retains complete scores, candidate sheets, CSV and JSON outputs, sixteen figures, references, and documentation.
Methodology
Project workflow
- 01Verify the source
The workflow checks the retained USGS files, row inventory, analytical method, coordinates, identifiers, and element columns.
- 02Build compositions
Reporting-limit flags are handled under declared assumptions before CLR and ILR transformations.
- 03Fit both backgrounds
A robust regional covariance model and local neighbour medians produce comparable empirical percentiles.
- 04Rank candidates
Global and local evidence are combined under one fixed top-two-percent screening rule.
- 05Test stability
Duplicates, 250 uncertainty runs, seven sensitivity cases, and 999 spatial permutations test the retained evidence.
Demonstration scenario
The completed study screens 1,549 regular samples and identifies 31 candidates under the fixed top-two-percent rule. Thirty remain stable at the declared 80 percent uncertainty condition. Field-duplicate status agrees for 98.2 percent of pairs, and the weak positive Moran I is separated from the permutation distribution.
Engineering
Tools and method
- Tools
- The project uses Python 3.14, NumPy, Pandas, SciPy, scikit-learn, pyrolite, Matplotlib for subject analysis, simulation, and results.
- Compositional layer
- SciPy-based Helmert coordinates implement CLR and ILR transforms independently checked against pyrolite 0.3.7.
- Robust model
- scikit-learn minimum covariance determinant and nearest-neighbour structures provide regional and local evidence.
- Uncertainty layer
- Seeded simulations vary censored values and duplicate-derived analytical precision, while named scenarios vary model settings.
- Verification
- Automated tests, dependency audit, Docker, retained-result contracts, document checks, and provenance checks verify the delivery.
Testing
Evaluation
Evaluation measures
- Candidate rank, global percentile, local percentile, and combined score
- Conditional candidate stability across 250 uncertainty runs
- Candidate-set overlap and full score-rank correlation across seven sensitivity cases
- Matched field-duplicate score and candidate-status agreement
- Aitchison distance and element-level duplicate precision
- Moran I with 999 seeded spatial permutations
Project boundaries
- Candidate rank is screening evidence and does not prove mineralisation, grade, tonnage, continuity, ownership, or economic value.
- The local model uses sample-point neighbours rather than delineated upstream catchments or geological domains.
- Uncertainty covers declared censoring and duplicate-derived variation, not every sampling, geological, spatial, or laboratory uncertainty.
- Field follow-up requires current geological review, permissions, repeat sampling, mineralogy, independent assays, and qualified geologists.
Included
- 01Complete Python source code
- 02Retained public-domain USGS source data with checksums and provenance
- 03Complete scores for 1,549 regular samples and 167 matched field duplicates
- 04Thirty-one candidate evidence sheets with concentrations, rankings, uncertainty, and interpretation
- 05Two hundred and fifty uncertainty runs and seven model sensitivity cases
- 06Sixteen project figures and two public-domain literature maps
- 07133-page project documentation in PDF and editable Word formats
- 0817-page setup and usage guide in PDF and editable Word formats
- 09Fifty-two annotated references
- 1053 automated tests with 98.85 percent source coverage
Project record
No information is collected on this page.
- Permanent project ID
- GP-GE-1EA9NOY
- Catalogued
- 21 Aug 2026
- Completed
- 25 Aug 2026
- Verified
- 25 Aug 2026
- Demonstration
- Included in repository
Handover
After purchase
- 01Payment is confirmed
The project is marked unavailable and cannot be purchased again.
- 02Repository access is granted
The buyer's submitted GitHub account receives access to the private repository.
- 03The purchase record is delivered
The certification sheet is prepared from the reviewed buyer details and sent privately by email.