← Back to project catalogue
GP-CY-0CV9Z4DCybersecurityReady

Prompt Injection Control Assurance for Tool-Using AI Agents

A completed defensive cybersecurity study comparing preventive, detective, authorization, containment, and confirmation controls for prompt injection in a harmless tool-using agent harness.

Prompt Injection Control Assurance for Tool-Using AI Agents project visual
GP-CY-0CV9Z4D · Cybersecurity
  • Python
  • NumPy
  • Pandas
  • Matplotlib
  • Docker

Software compatibility

Python 3.11 or later

The retained release and clean Docker run use Python 3.12. No model-provider account, paid AI service, or external dataset is required.

Project definition

Problem statement

A tool-using AI agent can encounter untrusted instructions inside mail, documents, calendar entries, or purchasing content. A model-only response filter does not establish whether an unsafe action path is prevented, detected, contained, or confirmed.

The engineering problem is to combine complementary controls, measure their security and utility effects, and state assurance claims without presenting a controlled harness result as proof of production safety.

Project objectives

  • Define a harmless fixture corpus covering benign work and six prompt-injection attack families.
  • Model mail, calendar, document, and purchasing workflows with explicit action classes.
  • Compare seven cumulative control profiles from an unprotected baseline to a restrictive reference profile.
  • Measure unsafe completion, benign task utility, false rejection, and confirmation burden.
  • Use control ablation and fixed-seed case-mix uncertainty to test result stability.
  • Connect each retained result to a bounded assurance claim, limitation, and further-work requirement.

Project structure

Project components

01

Evidence review

Screens 59 primary, official, standards, and peer-reviewed sources and records how each source supports the study.

02

Harmless fixture corpus

Defines 192 labelled cases without storing operational attack prompts, credentials, personal data, or live tool calls.

03

Control profiles

Composes source isolation, instruction detection, action authorization, data-flow containment, confirmation, and restrictive policy controls.

04

Assurance harness

Evaluates every fixture under every profile and retains the exact control decisions behind each outcome.

05

Uncertainty and ablation

Tests case-mix sensitivity and removes individual controls to identify redundancy and utility effects.

06

Assurance case

Maps evidence to claims, confidence, limitations, residual risk, and release recommendations.

Methodology

Project workflow

  1. 01
    Review the threat boundary

    Read the fixture rules, trust assumptions, action classes, and excluded operational attack content.

  2. 02
    Inspect the profiles

    Compare the seven declared combinations of preventive, detective, authorization, containment, and confirmation controls.

  3. 03
    Run the harness

    Evaluate 192 cases under each profile and retain 1,344 case-level results.

  4. 04
    Measure trade-offs

    Compare unsafe completion, benign utility, false rejection, confirmation burden, workflow effects, and attack-family effects.

  5. 05
    Test robustness

    Run 20,000 case-mix bootstrap samples per profile and review control-ablation evidence.

  6. 06
    State assurance claims

    Report only the claims supported by the prepared harness and identify the tests required before production use.

Demonstration scenario

The student compares the unprotected baseline with progressively layered profiles. In the prepared harness, profile P3 is the first profile with zero unsafe completions while retaining all benign tasks. Later profiles demonstrate the confirmation burden and utility cost of additional restrictions.

Engineering

Tools and method

Fixture model
Structured synthetic cases represent source trust, workflow, action class, attack family, and expected safety outcome.
Decision pipeline
Small deterministic control functions make every allow, block, contain, and confirmation decision inspectable.
Metric model
Security and utility measures remain separate so a restrictive policy cannot appear successful by blocking all work.
Uncertainty model
Fixed-seed bootstrap sampling tests how different case mixtures affect each profile result.
Reproducibility
Fixed configuration, retained CSV and JSON evidence, tests, document audits, and a clean Linux container reproduce the study.

Testing

Evaluation

Evaluation measures

  • Unsafe completion rate for every control profile
  • Benign task utility and false-rejection rate
  • Confirmation burden for benign and adversarial cases
  • Attack-family and workflow residual-risk comparison
  • Action-class and severity comparison
  • Single-control ablation results and redundant coverage
  • Twenty thousand fixed-seed case-mix samples per profile
  • Exact reproducibility of 1,344 retained case results

Project boundaries

  • The fixtures are harmless abstract cases and do not contain operational prompt-injection payloads.
  • The harness is deterministic and does not invoke a language model, remote tool, production agent, or external account.
  • A zero unsafe-completion result in the prepared harness does not prove immunity to prompt injection.
  • The reported rates apply only to the declared fixtures, profiles, scoring rules, and evidence cutoff.
  • Real deployment requires model-specific testing, tool sandboxing, authorization review, logging, incident response, and current threat intelligence.
  • No information is collected.

Included

  1. 01Complete defensive Python study
  2. 02Ninety-six benign and ninety-six harmless adversarial fixtures
  3. 03Seven layered control profiles across four agent workflows
  4. 04One thousand three hundred and forty-four retained case results
  5. 05Twenty thousand fixed-seed bootstrap samples for each profile
  6. 06Eight result tables, one study summary, and fourteen labelled figures
  7. 07106-page project documentation in PDF and editable Word formats
  8. 0819-page setup and usage guide in PDF and editable Word formats
  9. 09Fifty-nine screened and annotated references with a source matrix
  10. 10Complete project files and evidence in a private GitHub repository

Project record

No information is collected on this page.

Permanent project ID
GP-CY-0CV9Z4D
Catalogued
21 Aug 2026
Completed
29 Aug 2026
Verified
29 Aug 2026
Demonstration
Included in repository

Handover

After purchase

  1. 01
    Payment is confirmed

    The project is marked unavailable and cannot be purchased again.

  2. 02
    Repository access is granted

    The buyer's submitted GitHub account receives access to the private repository.

  3. 03
    The purchase record is delivered

    The certification sheet is prepared from the reviewed buyer details and sent privately by email.