Prompt Injection Control Assurance for Tool-Using AI Agents
A completed defensive cybersecurity study comparing preventive, detective, authorization, containment, and confirmation controls for prompt injection in a harmless tool-using agent harness.

Software compatibility
The retained release and clean Docker run use Python 3.12. No model-provider account, paid AI service, or external dataset is required.
Project definition
Problem statement
A tool-using AI agent can encounter untrusted instructions inside mail, documents, calendar entries, or purchasing content. A model-only response filter does not establish whether an unsafe action path is prevented, detected, contained, or confirmed.
The engineering problem is to combine complementary controls, measure their security and utility effects, and state assurance claims without presenting a controlled harness result as proof of production safety.
Project objectives
- Define a harmless fixture corpus covering benign work and six prompt-injection attack families.
- Model mail, calendar, document, and purchasing workflows with explicit action classes.
- Compare seven cumulative control profiles from an unprotected baseline to a restrictive reference profile.
- Measure unsafe completion, benign task utility, false rejection, and confirmation burden.
- Use control ablation and fixed-seed case-mix uncertainty to test result stability.
- Connect each retained result to a bounded assurance claim, limitation, and further-work requirement.
Project structure
Project components
Evidence review
Screens 59 primary, official, standards, and peer-reviewed sources and records how each source supports the study.
Harmless fixture corpus
Defines 192 labelled cases without storing operational attack prompts, credentials, personal data, or live tool calls.
Control profiles
Composes source isolation, instruction detection, action authorization, data-flow containment, confirmation, and restrictive policy controls.
Assurance harness
Evaluates every fixture under every profile and retains the exact control decisions behind each outcome.
Uncertainty and ablation
Tests case-mix sensitivity and removes individual controls to identify redundancy and utility effects.
Assurance case
Maps evidence to claims, confidence, limitations, residual risk, and release recommendations.
Methodology
Project workflow
- 01Review the threat boundary
Read the fixture rules, trust assumptions, action classes, and excluded operational attack content.
- 02Inspect the profiles
Compare the seven declared combinations of preventive, detective, authorization, containment, and confirmation controls.
- 03Run the harness
Evaluate 192 cases under each profile and retain 1,344 case-level results.
- 04Measure trade-offs
Compare unsafe completion, benign utility, false rejection, confirmation burden, workflow effects, and attack-family effects.
- 05Test robustness
Run 20,000 case-mix bootstrap samples per profile and review control-ablation evidence.
- 06State assurance claims
Report only the claims supported by the prepared harness and identify the tests required before production use.
Demonstration scenario
The student compares the unprotected baseline with progressively layered profiles. In the prepared harness, profile P3 is the first profile with zero unsafe completions while retaining all benign tasks. Later profiles demonstrate the confirmation burden and utility cost of additional restrictions.
Engineering
Tools and method
- Fixture model
- Structured synthetic cases represent source trust, workflow, action class, attack family, and expected safety outcome.
- Decision pipeline
- Small deterministic control functions make every allow, block, contain, and confirmation decision inspectable.
- Metric model
- Security and utility measures remain separate so a restrictive policy cannot appear successful by blocking all work.
- Uncertainty model
- Fixed-seed bootstrap sampling tests how different case mixtures affect each profile result.
- Reproducibility
- Fixed configuration, retained CSV and JSON evidence, tests, document audits, and a clean Linux container reproduce the study.
Testing
Evaluation
Evaluation measures
- Unsafe completion rate for every control profile
- Benign task utility and false-rejection rate
- Confirmation burden for benign and adversarial cases
- Attack-family and workflow residual-risk comparison
- Action-class and severity comparison
- Single-control ablation results and redundant coverage
- Twenty thousand fixed-seed case-mix samples per profile
- Exact reproducibility of 1,344 retained case results
Project boundaries
- The fixtures are harmless abstract cases and do not contain operational prompt-injection payloads.
- The harness is deterministic and does not invoke a language model, remote tool, production agent, or external account.
- A zero unsafe-completion result in the prepared harness does not prove immunity to prompt injection.
- The reported rates apply only to the declared fixtures, profiles, scoring rules, and evidence cutoff.
- Real deployment requires model-specific testing, tool sandboxing, authorization review, logging, incident response, and current threat intelligence.
- No information is collected.
Included
- 01Complete defensive Python study
- 02Ninety-six benign and ninety-six harmless adversarial fixtures
- 03Seven layered control profiles across four agent workflows
- 04One thousand three hundred and forty-four retained case results
- 05Twenty thousand fixed-seed bootstrap samples for each profile
- 06Eight result tables, one study summary, and fourteen labelled figures
- 07106-page project documentation in PDF and editable Word formats
- 0819-page setup and usage guide in PDF and editable Word formats
- 09Fifty-nine screened and annotated references with a source matrix
- 10Complete project files and evidence in a private GitHub repository
Project record
No information is collected on this page.
- Permanent project ID
- GP-CY-0CV9Z4D
- Catalogued
- 21 Aug 2026
- Completed
- 29 Aug 2026
- Verified
- 29 Aug 2026
- Demonstration
- Included in repository
Handover
After purchase
- 01Payment is confirmed
The project is marked unavailable and cannot be purchased again.
- 02Repository access is granted
The buyer's submitted GitHub account receives access to the private repository.
- 03The purchase record is delivered
The certification sheet is prepared from the reviewed buyer details and sent privately by email.