Pricing
One public index.
No negotiation games.
Our price is a fixed share of the public RunPod on-demand rate for your server type, re-based quarterly. Check it yourself before you talk to us.
The rate card
Your price, per server.
| Server | Reference rate | Your monthly cost | Typical saving |
|---|---|---|---|
| H100 SXM (80 GB) | $2.99/hr | $1,746 | ~60% |
| H200 SXM (141 GB) | $4.39/hr | $2,564 | ~60% |
| B200 SXM (180 GB) | $5.89/hr | $3,440 | ~60% |
| A100 SXM (80 GB) | $1.49/hr | $870 | ~60% |
| L40S (48 GB) | $0.99/hr | $578 | ~60% |
| L4 (24 GB) | $0.39/hr | $228 | ~60% |
| RTX 4090 (24 GB) | $0.69/hr | $403 | ~60% |
How we start our partnerships
Three steps, quickly settled
| Feature | What happens | What it costs you |
|---|---|---|
| Scoping call | A quick check on your fleet and workload under 30 minutes, no obligation. | Nothing. If Galahad will not pay off on your numbers, we say so and part as friends. |
| Measured pilot | We prove it on your own servers, with clear success criteria set upfront. | Nothing. You decide if it worked. |
| Rollout | Scale up server by server, at your own pace. | Per-server licences: daily, weekly, monthly or yearly, your choice. |
The things people push back on, answered straight
Doesn't vLLM already do this?
No. Every other cache remembers by approximating — it keeps a fuzzy, compressed version to go fast and gives up when the memory gets big. We remember exactly, bit for bit, provably identical, no matter how much you give it.
Why can't a bigger company just add this?
This is the hard part everyone else avoids. It is not a feature bolted onto an approximate cache — being exact is a different architecture, and it closes the shortcuts the incumbents built their performance on.
Does it work on today's hardware?
Yes. It works today, on today's hardware, and cuts real bills for real companies. No new law and no future breakthrough required.