Estimate the cost of running an AI workload, including caching, retries and human review. Compare the total with the number of results you can actually use. This calculator works with unexpired starter rates or your custom rates while KAPI-backed current pricing is paused.
How the estimate works
Calls = attempted tasks × (base calls per task + extra retry calls per 100 tasks ÷ 100). The calculator prices ordinary input, cache reads, cache writes and output separately. A cache miss uses the cache-write rate. Human review is charged once per attempted task.
Use measured cacheable tokens and hit rates. Caching is off by default because eligibility depends on your request. Astra requires a cacheable prefix of at least 1,024 visible tokens. OpenAI explains the caching requirements.
Accepted results = attempted tasks × accepted-result rate. At 0%, the calculator still shows spending, but cost per accepted task is undefined. A low token price does not prove that a model completes your workload well.
Where the prices come from
The table above is generated from the same loaded price catalog used by the calculator. It shows each source-check date and expiry, standard rates per million tokens, documented limits and official model-page links. Tiered models list standard rates through 272K input tokens and the published long-context tier above 272K; the GPT-5.4 long-context cached-input price is not stated in the source and cache-hit estimates in that tier are blocked.
KAPI-backed current pricing is paused; this calculator uses the separate starter catalog or your custom rates. Expired or malformed provider prices are excluded. Custom rates are never labeled as provider-verified, and the retry button reloads the local price catalog without changing verification dates.
Inspect KAPI’s August 24 historical snapshot and method. For generation by clip duration, use the AI video cost calculator.
