AI Engineering Rubric
| Dimension | Repeat | Pass | Strong |
|---|---|---|---|
| Baseline and evaluation | Demo examples only | Frozen set and meaningful baseline | Segmented evaluation with uncertainty and error analysis |
| Reproducibility | Undocumented manual run | Pinned code, data, model, and configuration | Clean-room reproduction succeeds |
| Safety and security | Generic risk prose | Concrete abuse tests and mitigations | Residual risk is measured, monitored, and rehearsed |
| Operations | No production evidence | Latency, cost, telemetry, fallback, and runbook | SLOs and failure drills drive revisions |
| Defense | Cannot explain model limitations | Defends value, tradeoffs, and failures | Clearly rejects inappropriate AI use cases |
Passing requires at least Pass in every row. A strong average cannot compensate for missing safety, evidence, or fallback behavior.