Product

Galahad: the harness that makes your model cheaper and smarter at once

A per-server software layer that sits between your workload and your silicon. It keeps a byte-exact memory of work your model has already proven, and grafts it back into new requests with zero tokens recomputed. The model stays frozen and stays yours — any model can run inside the harness. Remove it and everything runs as before. No lock-in, no data leaving your infrastructure.

The harness

Two engines, one flywheel

Merlin decides what can be reused and proves it is safe. Taliesin performs the graft. Together they turn work your model has already done into a permanent, verified asset.

Merlin

the memory

alone: 15% of the reference rate

Hashes every incoming prompt and asks one question: have we done this before? A live in-VRAM registry answers in under two microseconds; a durable on-disk ledger survives restarts and moves byte-identical between machines. A lookup is only a hit when two independent hashes agree, so it never returns the wrong memory. Merlin also runs standalone as content-aware deduplication — stripping redundant context before it reaches any model, including pure API setups.

1.4 µs to select · exact addressing makes zero errors where approximate retrieval picks wrong 94.3% of the time

Taliesin

the graft

included in Full Stack

Takes the state Merlin found and splices it into a live request with zero tokens recomputed. The restored logits are byte-for-byte identical to computing fresh — SHA-256 equal, zero KL divergence, verified across two model scales and two GPU targets. Blocks are merged with position re-binding and recompute-fused, so the model can reason across them rather than merely read them.

bit-exact restore · 100% argmax agreement over 50 samples · 6–23 ms per full reuse at 36 mWh

Boundaries

What Galahad is not

Not a cloud

You keep renting your servers wherever you like. We only license the software layer on top.

Not a model

The harness is model-agnostic — any model can run inside it, chosen for your needs. Whatever you run stays 100% yours, and we never see your data or weights.

Not a leap of faith

Every claim is measured, SHA-256 verifiable, and re-measured on your own workload during the pilot.

How buying works

Three steps, and an honest exit at each one

Feature What happens What it costs you
Free scoping call We look at your fleet, utilisation and workload. Nothing. If Galahad will not pay off on your numbers, we say so and part as friends.
Measured pilot Up to 3 months, up to 2 servers, with a measurement protocol agreed before we start. 50% of list price. The exit criterion is yours.
Rollout Per-server monthly licences on the servers you choose. Add or remove servers per calendar month; cancel per server after the first year.

See what it does to your bill

Five questions, sixty seconds, an honest answer.