AI Infra Interviews logo

Inference economics workbench

Seven free calculators for API budgets, fleet cost, peak capacity, API-versus-fleet break-even, cache reuse, human review and hardware ownership. Change the assumptions and save your scenario.

7 calculators · no account needed

Change the assumptions, inspect the formula and save your scenario. The starting numbers are teaching examples. Calculations stay in your browser.

API budget

Follow a month of business tasks through model calls and billable attempts. Token categories are disjoint averages per attempt.

Your calculation

Billable attempts
1,080,000 attempts
Model charges
11,880 USD
Other attempt-based charges
1,080 USD
Other monthly cost
2,000 USD
Total cost
14,960 USD
Accepted business tasks
900,000 tasks
Cost per accepted task
0.016622 USD / task
Compare the amounts
Model charges11,880 USD
Other attempt-based charges1,080 USD
Other monthly cost2,000 USD

All bars start at zero. Values come from the inputs above.

How the result is calculated

Attempts = tasks × calls per task × attempts per call. Spend = attempts × (sum of tokens × category price / 1,000,000 + tool cost) + fixed cost.

Default inputs are hypothetical. Include reasoning in billable output where your provider charges for it. Put cache storage and other monthly charges in fixed cost. Acceptance is measured per original business task.

Use the result to make a decision

A useful budget connects a business task to the calls it creates, the capacity those calls need and the outcomes the product accepts. Use the same workload and quality target when comparing options. Each calculator shows its formula and the conditions that could change the decision.

Questions people ask

Are the starting values current vendor prices?▲

The calculator defaults are hypothetical teaching inputs. Replace them with your invoice or quote. The linked playbook includes a dated register of selected public prices and explains the billing units.

Why measure cost per accepted task?▼
Does the capacity calculator predict GPU performance?▼
Are my inputs uploaded?▼