Skip to main content

Serving Load and Cost Lab

Deploy a bounded endpoint and measure throughput, tail latency, queueing, memory, and cost per successful outcome under realistic bursts.

Required evidence

Submit reproducible deployment configuration, correlated telemetry, load and failure results, user-impact and unit-cost calculations, one rejected alternative, and the executed response. A dashboard screenshot without raw query and version context does not pass.

Oral defense

Trace one request across boundaries, explain a tail-latency or quality regression, demonstrate containment and rollback, and defend the residual operational risk.

Source backbone

Use Building Secure and Reliable Systems, Google SRE books, OpenTelemetry specifications, and NIST AI RMF.