Skip to main content

AI Engineering Rubric

DimensionRepeatPassStrong
Baseline and evaluationDemo examples onlyFrozen set and meaningful baselineSegmented evaluation with uncertainty and error analysis
ReproducibilityUndocumented manual runPinned code, data, model, and configurationClean-room reproduction succeeds
Safety and securityGeneric risk proseConcrete abuse tests and mitigationsResidual risk is measured, monitored, and rehearsed
OperationsNo production evidenceLatency, cost, telemetry, fallback, and runbookSLOs and failure drills drive revisions
DefenseCannot explain model limitationsDefends value, tradeoffs, and failuresClearly rejects inappropriate AI use cases

Passing requires at least Pass in every row. A strong average cannot compensate for missing safety, evidence, or fallback behavior.