Financial services · Evaluation and assurance, 12 weeks
Putting a measured error rate behind claims-evidence retrieval at a specialist insurer
An evaluation harness, adversarial suite and assurance pack retrofitted to a retrieval system already in production, so the firm could state its error rate rather than describe its intentions.
- Sector
- Financial services
- Scale
- UK specialist insurer, ~1,200 staff
- Duration
- 12 weeks

Context
A specialist insurer had a retrieval system in limited production over claims correspondence, medical reports and policy wordings. It worked, in the sense that claims handlers used it and liked it.
It could not be extended to a wider set of claim types, because risk and compliance asked a question the delivery team could not answer: how often is it wrong, and in what way.
The problem
The system had been evaluated during build by the people who built it, on examples they selected. There was no held-out set, no regression suite, and no measurement of what happened to quality when a prompt or a model changed which by then had happened several times.
The specific worry was not fluency but silent degradation. A model update had already produced a noticeable change in answer style, and nobody could say whether it had also changed accuracy, because there was nothing to compare against.
Under the Consumer Duty the firm also needed to show that customer-affecting decisions supported by the system were based on evidence that was actually present in the file, and that the system's limitations were understood by the people relying on it.
Why earlier approaches failed
Spot-checking by handlers caught obvious errors and nothing systematic. It could not detect a four-point drop in recall on one document family, which is exactly the failure that matters.
A vendor-supplied accuracy figure was available and unusable: it was measured on a public benchmark bearing no relationship to medical reports written by consultants in a specialty, which is where this system's difficulty actually lives.
The pipeline we built
No changes to retrieval were made until it could be measured. Measurement then determined which changes were worth making.
01Question set construction
380 questions drawn from twelve months of query logs and from handler interviews, stratified across claim type, document family and question type. Labelled by two senior handlers with a third adjudicating disagreements; inter-annotator agreement measured and reported alongside every subsequent score.
02Retrieval baseline
Recall@10, MRR and answer-supported rate measured per document family. Aggregate performance was acceptable; medical reports were substantially worse than the mean, which the aggregate had been hiding.
03Generation scoring
Faithfulness, citation correctness and appropriate abstention scored with a rubric and an LLM judge. The judge was calibrated against human scores on a 120-item sample and its agreement rate published with every result, so it is treated as an instrument with known precision rather than as an oracle.
04Adversarial suite
Prompt injection through document content, including a real submitted PDF containing instruction-like text, attempts to elicit unsupported medical conclusions, and probes for cross-claim leakage.
05Permission and segregation testing
Synthetic users exercising claim-level and role-level boundaries on every release, including inference attacks that ask about a claim indirectly rather than naming it.
06Targeted remediation
Only then were changes made. Medical-report chunking moved from fixed windows to section-aware splitting on report structure, and a document-family filter was added ahead of ranking. Both were adopted because they moved the measured number on the family that was failing.
07CI regression
The full suite runs on every prompt, model or pipeline change, with thresholds that fail the build. A model version change is now a routine gated release rather than an unmonitored event.
08Production monitoring
Sampled live scoring, drift alerting on retrieval quality, and a handler feedback control whose corrections flow into the labelled set. The set has grown every month since handover.
What shipped
- A labelled evaluation set owned by the insurer, versioned alongside the codebase.
- A regression suite wired into CI with failing thresholds, plus the adversarial and permission suites.
- An assurance pack for the risk committee: method, measured error rates by failure mode, known limitations and mitigations.
- Production monitoring on retrieval quality, abstention rate and cost per answered question.
Outcomes
Each figure below carries the method behind it and the baseline it is measured against, which is the form a result has to take before it means anything.
- Answer-supported rate, medical reports
- Before and after section-aware chunking and a document-family pre-filter, on the held-out set. Aggregate performance across all families moved far less, from 88.1% to 94.7% which is why the aggregate had been hiding the problem.
- 72.8% → 93.1%
- Regressions caught before release
- Changes blocked by CI thresholds in the first six months. Two would have cost more than five points of recall on medical reports; one was a prompt edit, one a model version bump.
- 7 blocked
- Claim types approved for rollout
- Following the risk committee's review of the assurance pack. The fourth was deferred pending a data-residency question unrelated to retrieval quality.
- 3 of 4 requested
- Judge-to-human agreement
- Agreement rate of the calibrated LLM judge against two senior handlers on a 120-item stratified sample. Faithfulness judgements on medical causation fell below the threshold and stayed human-scored.
- 88.6%
“We did not need the system to be perfect. We needed to be able to say how imperfect it was, in writing, to a committee. Once the error rate existed as a number with a method behind it, the conversation about extending it took twenty minutes instead of two quarters.”
What next
The evaluation harness is being extended to cover a second retrieval system in the underwriting function, reusing the scoring pipeline and replacing only the question set.
Services involved
More engagements

Water & utilities
Making thirty years of asset documentation answerable at a UK water utility
A retrieval pipeline over 380,000 asset documents: drawings, condition reports, permits and handover packs, with permission inheritance preserved from the source document management system.
UK regional water utility, ~4,500 staff · 16 weeks
Read the engagement
Central government
Assembling regulatory evidence from casework correspondence in central government
An extraction and evidence-assembly pipeline over eleven years of casework correspondence, built to a standard where every extracted fact traces to a source passage and a pipeline version.
UK central government agency, ~2,000 staff · 20 weeks
Read the engagement
Professional services
A two-year AI roadmap grounded in what the document estate could actually support
Fourteen candidate use cases screened against real corpora and against whether anyone could define a correct answer. Six survived; the sequence was chosen so the first delivery paid for the second.
Mid-market UK professional services firm, ~600 staff · 7 weeks
Read the engagement
Get in touch
Talk to us.
A first conversation runs about forty-five minutes and covers three things: what your estate actually looks like, whether anyone can define a correct answer or an agreed number, and whether your permission model resolves per user. Any one of them can rule the work out, and we would rather tell you in week one.
- Prefer email
- hello@vectisflow.com
- Response time
- One working day, from a person who has read it.