The verified memory layer for AI

Pay once for the answer. Reuse it forever.

Your model recomputes the same work on every request. Galahad stores what it has already proven and grafts it back byte-exact — so repeat work costs zero generation tokens, and the answer is identical every time.

every number measured · sha-256 verified · works with the model you already run

galahad verify --exact
galahad reuse --model gemma-4-12b --ctx 64kmerlin  lookup ................. 1.4 µs  HITtaliesin graft ................. 0 tokens recomputedsha-256  a3f9c1..8e2d  expectedsha-256  a3f9c1..8e2d  actualbit-exact match — 0 divergent bytescost -60% · answer identical

The harness

Two engines, one flywheel

Merlin decides what can be reused and proves it is safe. Taliesin performs the graft. Together they turn work your model has already done into a permanent, verified asset.

Merlin

the memory

alone: 15% of the reference rate

Hashes every incoming prompt and asks one question: have we done this before? A live in-VRAM registry answers in microseconds; a durable on-disk ledger survives restarts and moves byte-identical between machines. A lookup is only a hit when two independent hashes agree, so it never returns the wrong memory. Merlin also runs standalone as deduplication, against any provider including pure API setups.

1.4 µs to select · exact addressing makes zero errors where approximate retrieval picks wrong 94.3% of the time

Taliesin

the graft

included in Full Stack

Takes the state Merlin found and splices it into a live request with zero tokens recomputed. The restored logits are byte-for-byte identical to computing fresh — SHA-256 equal, zero KL divergence. Blocks are merged with position re-binding and recompute-fused, so the model reasons across them rather than merely reading them.

bit-exact restore · 100% argmax agreement over 50 samples · 6–23 ms per reuse at 36 mWh

Radical honesty

If it won't pay off, we tell you.

Free scoping call

Fleet, utilisation, workload fit. Wrong fit? We say so and part as friends.

Measured pilot

Up to 3 months, 2 servers, 50% off — with a measurement protocol agreed up front.

No lock-in

Your data never leaves your infrastructure. Cancel per server, monthly, after the first year.

Your GPU bill, measurably smaller.

−60% Typical cut to your monthly server bill, on the hardware you already rent. agreed with an enterprise customer
0 Generation tokens on work already proven. The answer returns byte-identical. 180/180 across nine problem families
8.8× Faster 64K-context loads, bit-exact. Verified across four model vendors. sha-256 · 25/25 byte-equal
−82% Cost per repeated query in RAG and agent loops. Break-even from the second query. measured · 5.5× faster

Stop building bigger. Build the layer below.

Per-server licensing · indexed to public RunPod rates · pilots that measure before you commit