Pricing
One public index. No negotiation games.
Our price is a fixed share of the public RunPod on-demand rate for your server type, re-based quarterly. Check it yourself before you talk to us.
The rate card
What that works out to per server
Galahad Full Stack
Merlin and Taliesin together — the verified-knowledge flywheel. Work your model has already done becomes a permanent asset that is reused at zero generation tokens, byte-exact, every time.
Merlin
Deduplication only. Strips redundant context before it reaches your model, so every billed token does new work. Runs against any provider, including pure API setups with no GPU of your own.
Taliesin performs the graft and is licensed as part of Full Stack: it acts on what Merlin finds, so it is not sold on its own.
| Server | Reference rate | Your monthly cost | Typical saving |
|---|---|---|---|
| H100 SXM (80 GB) | $2.99/hr | $1,746 | ~60% |
| H200 SXM (141 GB) | $4.39/hr | $2,564 | ~60% |
| B200 SXM (180 GB) | $5.89/hr | $3,440 | ~60% |
| A100 SXM (80 GB) | $1.49/hr | $870 | ~60% |
| L40S (48 GB) | $0.99/hr | $578 | ~60% |
| L4 (24 GB) | $0.39/hr | $228 | ~60% |
| RTX 4090 (24 GB) | $0.69/hr | $403 | ~60% |
The rules, in plain sight
- Minimum 2 servers.
- Volume: 5–19 servers −10%, 20–49 −15%, 50+ −20%.
- Annual prepayment: −5%.
- Pilot: up to 3 months, 2 servers, 50% off.
- No lock-in: cancel per server, monthly, after the first year.
The things people push back on, answered straight
Doesn't vLLM already do this?
No. Every other cache remembers by approximating — it keeps a fuzzy, compressed version to go fast and gives up when the memory gets big. We remember exactly, bit for bit, provably identical, no matter how much you give it.
Why can't a bigger company just add this?
This is the hard part everyone else avoids. It is not a feature bolted onto an approximate cache — being exact is a different architecture, and it closes the shortcuts the incumbents built their performance on.
Does it work on today's hardware?
Yes. It works today, on today's hardware, and cuts real bills for real companies. No new law and no future breakthrough required.
Which of your workloads is bleeding budget?
We can most likely cut it. The fastest way to see exactly what we are is to prove it on your own numbers.