Prompt Injection and Tool Security Lab
Attack direct and indirect instruction paths. Enforce least privilege, typed arguments, confirmation, isolation, and auditable denial outside the model.
Required evidence
Submit versioned tests, system manifest, raw and summarized results, representative failures, a rejected alternative, named control owners, and residual-risk decision. Safety claims without adversarial evidence do not pass.
Oral defense
Demonstrate one failure, identify its first violated boundary, explain why layered controls limit impact, and execute the stop or escalation path.
Source backbone
Use NIST AI RMF, OWASP LLM Top 10, NIST SSDF, and Building Secure and Reliable Systems.