Worked Examples
Validate an LLM judge against blinded human ratings; trace an indirect prompt injection from retrieved content to an attempted refund; and turn a vague “hallucination” complaint into a mechanism-based failure taxonomy and regression gate.
Required evidence
Submit versioned tests, system manifest, raw and summarized results, representative failures, a rejected alternative, named control owners, and residual-risk decision. Safety claims without adversarial evidence do not pass.
Oral defense
Demonstrate one failure, identify its first violated boundary, explain why layered controls limit impact, and execute the stop or escalation path.
Source backbone
Use NIST AI RMF, OWASP LLM Top 10, NIST SSDF, and Building Secure and Reliable Systems.