Set an SLO for a Prediction Endpoint
Create a model-serving SLO that reflects user impact and operational response.
A recommendation endpoint is officially "up" but mobile users see empty or late shelves during traffic spikes. Start with user-visible behavior, then define SLIs and SLOs that trigger release and recovery decisions. The common trap is copying a generic uptime SLO from the platform and missing model-specific failure modes like fallback spikes, feature missingness, and long-tail latency. Before 99.9 percent endpoint uptime. No target for P95/P99 inference latency, feature missingness, fallback rate, or scored-list completeness. After 99 percent of mobile requests return scored recommendations or approved fallback within 250 ms, fallback rate below 4 percent over 30 minutes, feature missingness…
Sign up free — one personalized lesson every day, matched to your role and goals.
Already have an account? Sign in