Skip to main content
VectisFlow
VectisFlow
All case studies

Financial services · Evaluation and assurance, 12 weeks

Putting a measured error rate behind claims-evidence retrieval at a specialist insurer

An evaluation harness, adversarial suite and assurance pack retrofitted to a retrieval system already in production, so the firm could state its error rate rather than describe its intentions.

Sector
Financial services
Scale
UK specialist insurer, ~1,200 staff
Duration
12 weeks
Glass and steel facade of a City office building reflecting a grey sky

Context

A specialist insurer had a retrieval system in limited production over claims correspondence, medical reports and policy wordings. It worked, in the sense that claims handlers used it and liked it.

It could not be extended to a wider set of claim types, because risk and compliance asked a question the delivery team could not answer: how often is it wrong, and in what way.

The problem

The system had been evaluated during build by the people who built it, on examples they selected. There was no held-out set, no regression suite, and no measurement of what happened to quality when a prompt or a model changed which by then had happened several times.

The specific worry was not fluency but silent degradation. A model update had already produced a noticeable change in answer style, and nobody could say whether it had also changed accuracy, because there was nothing to compare against.

Under the Consumer Duty the firm also needed to show that customer-affecting decisions supported by the system were based on evidence that was actually present in the file, and that the system's limitations were understood by the people relying on it.

Why earlier approaches failed

Spot-checking by handlers caught obvious errors and nothing systematic. It could not detect a four-point drop in recall on one document family, which is exactly the failure that matters.

A vendor-supplied accuracy figure was available and unusable: it was measured on a public benchmark bearing no relationship to medical reports written by consultants in a specialty, which is where this system's difficulty actually lives.

The pipeline we built

No changes to retrieval were made until it could be measured. Measurement then determined which changes were worth making.

  1. 01Question set construction

    380 questions drawn from twelve months of query logs and from handler interviews, stratified across claim type, document family and question type. Labelled by two senior handlers with a third adjudicating disagreements; inter-annotator agreement measured and reported alongside every subsequent score.

  2. 02Retrieval baseline

    Recall@10, MRR and answer-supported rate measured per document family. Aggregate performance was acceptable; medical reports were substantially worse than the mean, which the aggregate had been hiding.

  3. 03Generation scoring

    Faithfulness, citation correctness and appropriate abstention scored with a rubric and an LLM judge. The judge was calibrated against human scores on a 120-item sample and its agreement rate published with every result, so it is treated as an instrument with known precision rather than as an oracle.

  4. 04Adversarial suite

    Prompt injection through document content, including a real submitted PDF containing instruction-like text, attempts to elicit unsupported medical conclusions, and probes for cross-claim leakage.

  5. 05Permission and segregation testing

    Synthetic users exercising claim-level and role-level boundaries on every release, including inference attacks that ask about a claim indirectly rather than naming it.

  6. 06Targeted remediation

    Only then were changes made. Medical-report chunking moved from fixed windows to section-aware splitting on report structure, and a document-family filter was added ahead of ranking. Both were adopted because they moved the measured number on the family that was failing.

  7. 07CI regression

    The full suite runs on every prompt, model or pipeline change, with thresholds that fail the build. A model version change is now a routine gated release rather than an unmonitored event.

  8. 08Production monitoring

    Sampled live scoring, drift alerting on retrieval quality, and a handler feedback control whose corrections flow into the labelled set. The set has grown every month since handover.

What shipped

  • A labelled evaluation set owned by the insurer, versioned alongside the codebase.
  • A regression suite wired into CI with failing thresholds, plus the adversarial and permission suites.
  • An assurance pack for the risk committee: method, measured error rates by failure mode, known limitations and mitigations.
  • Production monitoring on retrieval quality, abstention rate and cost per answered question.

Outcomes

Each figure below carries the method behind it and the baseline it is measured against, which is the form a result has to take before it means anything.

Answer-supported rate, medical reports
Before and after section-aware chunking and a document-family pre-filter, on the held-out set. Aggregate performance across all families moved far less, from 88.1% to 94.7% which is why the aggregate had been hiding the problem.
72.8% → 93.1%
Regressions caught before release
Changes blocked by CI thresholds in the first six months. Two would have cost more than five points of recall on medical reports; one was a prompt edit, one a model version bump.
7 blocked
Claim types approved for rollout
Following the risk committee's review of the assurance pack. The fourth was deferred pending a data-residency question unrelated to retrieval quality.
3 of 4 requested
Judge-to-human agreement
Agreement rate of the calibrated LLM judge against two senior handlers on a 120-item stratified sample. Faithfulness judgements on medical causation fell below the threshold and stayed human-scored.
88.6%
We did not need the system to be perfect. We needed to be able to say how imperfect it was, in writing, to a committee. Once the error rate existed as a number with a method behind it, the conversation about extending it took twenty minutes instead of two quarters.
Director of Claims Operations, UK specialist insurer

What next

The evaluation harness is being extended to cover a second retrieval system in the underwriting function, reusing the scoring pipeline and replacing only the question set.

More engagements

  • Water treatment infrastructure at dusk, concrete channels running into the distance

    Water & utilities

    Making thirty years of asset documentation answerable at a UK water utility

    A retrieval pipeline over 380,000 asset documents: drawings, condition reports, permits and handover packs, with permission inheritance preserved from the source document management system.

    UK regional water utility, ~4,500 staff · 16 weeks

    Read the engagement
  • Repeating stone facade of a government building, shot from below against overcast sky

    Central government

    Assembling regulatory evidence from casework correspondence in central government

    An extraction and evidence-assembly pipeline over eleven years of casework correspondence, built to a standard where every extracted fact traces to a source passage and a pipeline version.

    UK central government agency, ~2,000 staff · 20 weeks

    Read the engagement
  • Empty modern meeting room with a long table and floor-to-ceiling windows

    Professional services

    A two-year AI roadmap grounded in what the document estate could actually support

    Fourteen candidate use cases screened against real corpora and against whether anyone could define a correct answer. Six survived; the sequence was chosen so the first delivery paid for the second.

    Mid-market UK professional services firm, ~600 staff · 7 weeks

    Read the engagement

Get in touch

Talk to us.

A first conversation runs about forty-five minutes and covers three things: what your estate actually looks like, whether anyone can define a correct answer or an agreed number, and whether your permission model resolves per user. Any one of them can rule the work out, and we would rather tell you in week one.

Response time
One working day, from a person who has read it.