New: the Throttle Cost Index →

Find and prove inference savingsfor agent workloads on your own GPUs.

Throttle measures your agent workload's $ per million tokens with a 95% interval, keeps only the config changes that beat the noise, and proves the savings. Qwen2.5-72B on one MI300X: $1.67 per million tokens. On two: $2.14. Measured, not guessed.

Start a Throttle Pilot → Book a demo →
Watch the 75-second film

Every config change is a new bill. Throttle checks it.

  1. Measure

    throttle check prints your $ per million tokens, with a 95% interval.

  2. Change one thing

    A flag, the model, the engine, the quantization.

  3. Get a verdict

    Only when the change beats your measured run-to-run noise.

    CHEAPERMORE EXPENSIVENO WINNER

Real runs. Real dollars.

GPU count · AMD MI300X

One GPU beat two.

$1.67vs$2.14
per M output tokens · 1 GPU vs 2

Two GPUs ran 1.56× faster but cost 28% more per token.

Qwen2.5-72B BF16, vLLM 0.23 ROCm, 32 concurrent requests. 4 + 3 checks, Oct 1 2026.

one vLLM flag · AMD MI300X

One flag tripled the bill.

$0.227→$0.700
per M output tokens · +208%

--max-num-seqs 8 at 32 concurrent requests.

Qwen2.5-7B, vLLM 0.23 ROCm. MORE EXPENSIVE in 4 of 4 checks.

one vLLM flag · NVIDIA A100

One flag cut it by two-thirds.

$0.746→$0.234
per M output tokens · −68.6%

max_num_seqs 1 → 8.

Qwen2.5-0.5B, vLLM 0.16, A100 80GB. Counterbalanced golden run.

GPU $/hr are list prices (assumed). Tokens and time are measured.Every measured result → Cost Index

Real terminal runs of throttle-pro 0.4.2 on a MacBook (GPU rate assumed at $1.50/hr), and the A100 result from the saved golden run.

Throttle Pilot. We find savings on your agent workload and prove them.

throttle pilot · a small group of teams

Free for 3 months

We’re working with a small group of teams running agents on their own GPUs. For 3 months, free, we set Throttle up on your stack, find the config changes that lower your cost per token, and prove each one. We’d love to help you with your results.

After the pilot, pricing is based on the savings we verify together, at the conservative end of the 95% interval.

Apply for the pilot →

For agent workloads on GPUs you own or rent. How the pilot works

how it works

Baseline, test, prove

  1. Baseline. We measure your agent workload’s $ per million tokens with a 95% interval, on your own GPUs.
  2. Test. We try config changes and keep only the winners that beat the noise with unchanged outputs.
  3. Prove. Each change comes with a verified-savings statement (throttle savings), counted at the conservative end of the interval.

See it on your stack. Book a demo.

demo · agent workloads on your own GPUs

Walk through Throttle with the founder

Email us your model, your GPUs and what your agents do. We’ll reply with a time and show you how Throttle measures your cost per token and which changes are worth testing.

Book a demo →

Or write to hello@throttle-pro.com