Skip to main content
MODEL-DEPLOYMENT6 MIN READ

Set an SLO for a Prediction Endpoint

Create a model-serving SLO that reflects user impact and operational response.

A recommendation endpoint is officially "up" but mobile users see empty or late shelves during traffic spikes. Start with user-visible behavior, then define SLIs and SLOs that trigger release and recovery decisions. The common trap is copying a generic uptime SLO from the platform and missing model-specific failure modes like fallback spikes, feature missingness, and long-tail latency. Before 99.9 percent endpoint uptime. No target for P95/P99 inference latency, feature missingness, fallback rate, or scored-list completeness. After 99 percent of mobile requests return scored recommendations or approved fallback within 250 ms, fallback rate below 4 percent over 30 minutes, feature missingness…

Read the full lesson

Sign up free — one personalized lesson every day, matched to your role and goals.

Already have an account? Sign in

← Back to library
Contact us