Run locally or pay per job?
Compare cash and operating costs for the same volume of accepted work. Change the assumptions to match your hardware and model. These starting values are illustrative.
JavaScript is required for interactive estimates.
Compare the same job, then compare cost
A small local model and a hosted model may produce different quality. Enter separate attempt counts and review times for each. Passing a memory-fit check does not establish speed or task accuracy. This is a capacity and cost scenario, not a benchmark.
Local cost includes new hardware cash spending when selected, active and idle power, human review, and maintenance. Already-owned hardware has no new purchase cost in this view; depreciation, financing, resale value and opportunity cost are excluded. Available hours assume one sequential worker. Local total is withheld if the work exceeds capacity.
Hosted cost applies uncached and cached input rates, output rates, attempts, review, and your extra-fee budget. Include any billed reasoning tokens in output. Cache-write/storage fees, long-context multipliers and tool calls must be included in your chosen effective rates or extras. Price increases over the horizon are not forecast.
Reuse your measurements from Local Lab and verify rates in Model Economics. The optional preset preserves a record from the existing KAPI dataset and its original date; it does not update that dataset.
Inspect a measured local-model test
See how three exact Qwen3 files handled four short JSON tasks on an M4 Pro, with observed timing, sampled process memory and downloadable outputs. The small task set and differing quantizations limit what the results establish.