Pricing
One public index. No negotiation games.
Our price is a fixed share of the public RunPod on-demand rate for your server type, re-based quarterly. Check it yourself before you talk to us.
The rate card
What that works out to per server
Full Stack
Everything, bundled. The modules bought separately add up to 95%, so the bundle is always the better deal.
Training Engine
Training and fine-tuning throughput.
Taliesin
Long-context inference and exact recall.
Merlin
Repeated-context and agent pipelines.
Start with one module if you like — you can move to the bundle at any time.
| Server | Reference rate | Your monthly cost | Typical saving |
|---|---|---|---|
| H100 SXM (80 GB) | $2.99/hr | $1,746 | ~60% |
| H200 SXM (141 GB) | $4.39/hr | $2,564 | ~60% |
| B200 SXM (180 GB) | $5.89/hr | $3,440 | ~60% |
| A100 SXM (80 GB) | $1.49/hr | $870 | ~60% |
| L40S (48 GB) | $0.99/hr | $578 | ~60% |
| L4 (24 GB) | $0.39/hr | $228 | ~60% |
| RTX 4090 (24 GB) | $0.69/hr | $403 | ~60% |
The rules, in plain sight
- Minimum 2 servers.
- Volume: 5–19 servers −10%, 20–49 −15%, 50+ −20%.
- Annual prepayment: −5%.
- Pilot: up to 3 months, 2 servers, 50% off.
- No lock-in: cancel per server, monthly, after the first year.
The things people push back on, answered straight
Doesn't vLLM already do this?
No. Every other cache remembers by approximating — it keeps a fuzzy, compressed version to go fast and gives up when the memory gets big. We remember exactly, bit for bit, provably identical, no matter how much you give it.
Why can't a bigger company just add this?
This is the hard part everyone else avoids. It is not a feature bolted onto an approximate cache — being exact is a different architecture, and it closes the shortcuts the incumbents built their performance on.
Does it work on today's hardware?
Yes. It works today, on today's hardware, and cuts real bills for real companies. No new law and no future breakthrough required.
Which of your workloads is bleeding budget?
We can most likely cut it. The fastest way to see exactly what we are is to prove it on your own numbers.