CacheVerifier
ProductGuidesResearchBlogPricingDocs
Log inSign up

Blog

Data from CacheVerifier's own research, and notes from running the service in production — what the experiments actually showed, what broke, and what we changed.

September 11, 2026Fine-tuning on real feedback didn't protect against adversarial inputs — it made things slightly worseAn off-the-shelf verifier false-accepts 84% of deliberately adversarial query pairs. Fine-tuning on natural gray-zone feedback doesn't fix that — it pushes the number to 87.6%. What actually worked, and the honest caveat about how well the test taxonomy predicts real errors.Read September 11, 2026Our own risk guarantee overshot its target — here's what we found when we checkedA Conformal Risk Control threshold is supposed to hold a false-reuse-risk bound exactly. Under the realistic chronological test we actually run it against, it overshot on one dataset and looked like it overshot worse on another — until we found that second one was our own measurement artifact.Read September 11, 2026The dataset where fine-tuning turned actively harmful — and how we now catch itOn real multi-year Comcast support traffic, a fine-tuned verifier's gray-zone positive rate drifted 5x and the model went from helping to actively hurting. Here's the failure, and the three-detector monitor we built after finding it.Read
© 2026 CacheVerifier. Hosted semantic cache verification.
ProductGuidesResearchBlogPricingDocsFAQPrivacyTermsPython SDK