Interpret input, cached-input, output, and batch price magnitudes to choose a practical first optimization move.
Read the pricing magnitudes, then choose the first move for a workflow that repeats context and produces long nightly summaries. Example gpt-5.4-mini short-context prices per 1M tokens Standard output Completion tokens in the live lane $4.50 Standard input Fresh prompt tokens in the live lane $0.75 Cached input Repeated prefix when cache hits $0.075 Batch output Completion tokens in the asynchronous lane $2.25 Output length can dominate the bill. Repeated prefixes are cheapest when they actually cache. Batch trades immediacy for lower unit cost. For a nightly summary workflow with repeated policy context and long answers, which first move best…
Sign up free — one personalized lesson every day, matched to your role and goals.
Already have an account? Sign in