Blog

We pushed 3,400 labeled pairs through our live API — here's where verification actually pays off

production testfine-tuningbenchmarks

Most of what we've published so far comes from the research behind CacheVerifier: offline experiments, run with research code. This time we tested the product itself. We created fresh accounts on the live service, pushed 3,400 labeled cache candidates through the same API a customer would call, fine-tuned through the same endpoint, and measured what came back. Then we deleted the accounts.

The short version: verification pays off clearly in one place, and that place is narrower than "always". This post covers where it is, where it isn't, and what we fixed along the way.

The setup

  • Data. Two public benchmarks, 1,700 gray-zone candidates each (cosine similarity 0.80–0.97), every one labeled correct or wrong for its query:
    • LMArena: long, varied chat prompts. 17% of candidates are correct reuses.
    • AmazonHelp: short customer-support tweets. 2% are correct, which is typical of support traffic.
  • Protocol. Candidates stay in time order. The first 1,200 go in as feedback and train the fine-tune; the last 500 are a holdout that neither the stock model nor the fine-tune ever saw. Every number below is on that holdout.
  • Same path as a customer. POST /v1/verify/batch, POST /v1/feedback/batch, POST /v1/finetune/jobs. No special access.

What we fixed first

The first full run surfaced two problems, and we fixed both before taking the numbers below.

  • Fine-tuning was barely training. The trainer never set a learning-rate warmup, so the library default kicked in. That default is 10,000 warmup steps, and a typical fine-tune is only a few hundred steps. The model spent its whole run warming up and never got past about 2% of its learning rate, so it shifted scores without ever learning to reorder them. Warmup is now 10% of the run, and fine-tunes train for 3 epochs.
  • Heavy concurrent load could error. Under several simultaneous batch requests, some failed. That's fixed: under the same load every request now succeeds, and throughput is higher than before.

Fine-tuning on LMArena

Here's the 500-pair holdout (27% correct), before and after fine-tuning.

Stock model Fine-tuned Fine-tuned + similarity
AUC 0.903 0.930 —
Correct answers reused (recall) 74% 84% 94%
Wrong answers let through (false-accept rate) 11.0% 12.9% 10.4%
Reused answers that were correct (precision) 71% 71% 77%

The last column adds the cache's own similarity score to the verdict: if you send similarity_score, the fine-tune fits a small rule that combines it with the verifier score. It only keeps that rule when it beats the verifier alone on held-out data, which it did here. Recall rises by 20 points while false accepts fall.

The comparison that matters: just raising the similarity threshold

The fair question isn't "fine-tuned versus not fine-tuned." It's "verifier versus what you'd do without one", and what you'd do without one is raise your similarity threshold. So we ran the same holdout through a similarity threshold alone, and for both methods we picked the best threshold with hindsight on the same 500 pairs. That's generous to both sides, and especially to the threshold, which has nothing else to tune.

On LMArena's 500-pair holdout Fine-tuned verifier Similarity threshold only
Recall at the same false-accept rate (12.9%) 84% 97%
Recall when 90% of reused answers must be correct 44% 7%

The two rows tell different stories:

  • At moderate risk, a good similarity threshold is hard to beat. If you can tolerate letting through about 1 in 8 of the wrong candidates (on this data, 26–29% of reused answers end up wrong at that point), raising the threshold gets you there, and here it even does better.
  • At strict precision, it falls apart. On this data, right and wrong answers are mixed together at the high-similarity end, so cutting on similarity can't separate them: when 90% of reused answers must be correct, a similarity threshold can only reuse 7% of the reusable answers. The fine-tuned verifier reuses about 6× as many. (Inside the 0.90–0.97 band the stock model is close to a coin flip too, at AUC 0.53–0.55; fine-tuning is what moves it.)

That's where verification pays off: the regime where a wrong cached answer is expensive enough that you need most reuses to be right. It's also the regime CacheVerifier's risk certification is built for.

When there's nothing to reuse

AmazonHelp shows the other side. Only 10 of the 500 holdout candidates are correct reuses.

  • The stock model let 85 wrong answers through (17% of the wrong ones) and caught 2 of the 10 correct ones.
  • After fine-tuning, the verifier passed our automatic checks and activated. It let through about 15–20 wrong answers and at most 1 correct one.

That's much safer, and it saves almost nothing, because there were almost no correct answers to save. A similarity threshold did no better. If your gray zone looks like this, verification mostly keeps bad answers out, and the cost savings you'd want from a cache won't show up. A Health Check on a sample of your own traffic tells you which case you're in before you commit to anything.

Latency

Server-side model time per /v1/verify call was 18 ms at the median for short texts and 32 ms for long ones. That matches what we already publish for the full server round trip. Batching helps a lot: from our test client, where network time dominates, sending candidates in batches raised throughput more than tenfold over one request at a time.

Caveats

  • Public benchmarks, not your traffic. "Correct" here means the benchmark's own equivalence labels, which aren't the same as "this answer is right for this user". LMArena's labels may also be friendlier to similarity than real traffic is.
  • One run. We'll repeat this as the product changes.
  • Hindsight thresholds. Both best-recall comparisons pick the threshold on the holdout itself. In production you pick it on past data, so expect somewhat lower numbers for both methods.
  • Risk certification needs stable traffic. A risk-certified threshold held its target on traffic from the same period, but ran 1.4–1.8× over on later traffic that had drifted. We cover that in our post on the guarantee. For traffic that keeps changing, you can now opt in to an online adaptive threshold that keeps adjusting from randomized audits.

What to take from this

If wrong cached answers are cheap for you, tune your similarity threshold and you may not need a verifier. If they're expensive, and you need most reused answers to be right, a fine-tuned verifier reuses several times more answers than a threshold can at the same precision. Either way, measure it on your own traffic first: the free Health Check runs the same stock-versus-fine-tuned comparison on a sample of it.