Pricing

One public index. No negotiation games.

Our price is a fixed share of the public RunPod on-demand rate for your server type, re-based quarterly. Check it yourself before you talk to us.

The rate card

What that works out to per server

Training Engine

50% of the RunPod on-demand rate

Training and fine-tuning throughput.

Taliesin

30% of the RunPod on-demand rate

Long-context inference and exact recall.

Merlin

15% of the RunPod on-demand rate

Repeated-context and agent pipelines.

Start with one module if you like — you can move to the bundle at any time.

Server Reference rate Your monthly cost Typical saving
H100 SXM (80 GB) $2.99/hr $1,746 ~60%
H200 SXM (141 GB) $4.39/hr $2,564 ~60%
B200 SXM (180 GB) $5.89/hr $3,440 ~60%
A100 SXM (80 GB) $1.49/hr $870 ~60%
L40S (48 GB) $0.99/hr $578 ~60%
L4 (24 GB) $0.39/hr $228 ~60%
RTX 4090 (24 GB) $0.69/hr $403 ~60%

The rules, in plain sight

  • Minimum 2 servers.
  • Volume: 5–19 servers −10%, 20–49 −15%, 50+ −20%.
  • Annual prepayment: −5%.
  • Pilot: up to 3 months, 2 servers, 50% off.
  • No lock-in: cancel per server, monthly, after the first year.

The things people push back on, answered straight

Doesn't vLLM already do this?

No. Every other cache remembers by approximating — it keeps a fuzzy, compressed version to go fast and gives up when the memory gets big. We remember exactly, bit for bit, provably identical, no matter how much you give it.

Why can't a bigger company just add this?

This is the hard part everyone else avoids. It is not a feature bolted onto an approximate cache — being exact is a different architecture, and it closes the shortcuts the incumbents built their performance on.

Does it work on today's hardware?

Yes. It works today, on today's hardware, and cuts real bills for real companies. No new law and no future breakthrough required.

Which of your workloads is bleeding budget?

We can most likely cut it. The fastest way to see exactly what we are is to prove it on your own numbers.