Skip to main content
VectisFlow
VectisFlow
All case studies

Professional services · Data and AI strategy, 7 weeks

A two-year AI roadmap grounded in what the document estate could actually support

Fourteen candidate use cases screened against real corpora and against whether anyone could define a correct answer. Six survived; the sequence was chosen so the first delivery paid for the second.

Sector
Professional services
Scale
Mid-market UK professional services firm, ~600 staff
Duration
7 weeks
Empty modern meeting room with a long table and floor-to-ceiling windows

Context

A mid-market professional services firm had a board commitment to AI, a budget attached to it, and fourteen candidate use cases collected from practice leads.

What it did not have was any assessment of whether the underlying data could support them. The estimates in the list had been produced in a workshop, without anyone opening a document.

The problem

The firm's competitive knowledge sits in engagement files: reports, working papers, correspondence and precedent documents, held across a document management system, several practice-specific shares and, for older material, an archive nobody had opened in years.

Several candidates assumed a level of structure that did not exist. Others assumed access that partners would not grant. At least one assumed a definition of correctness that three interviewees defined three different ways.

The firm's real risk was not choosing the wrong first project. It was spending a year discovering, one project at a time, that its estate was not ready, and losing the board's appetite in the process.

Why earlier approaches failed

A previous strategy engagement had produced a value-versus-effort matrix. Effort had been estimated from the use-case description rather than from the data, and the two projects started on its recommendation both stalled on document quality within a quarter.

An internal proof of concept had worked well on a curated set of eighty documents. That set had been assembled by the person running the proof of concept, which made it a demonstration of the concept rather than a test of the estate.

The pipeline we built

Seven weeks, run jointly with the firm's technology and risk leads. The screening work is deliberately front-loaded: the point is to disqualify early and cheaply.

  1. 01Opportunity capture

    Interviews with fee earners as well as practice leads. Two of the strongest candidates came from associates describing work they considered unremarkable, and had not appeared on the original list at all.

  2. 02Corpus sampling

    Stratified samples from each repository, assessed for text quality, structural consistency, revision hygiene and duplication. The archive was 64% scanned with no reliable text layer, which removed two candidates outright and rescoped a third.

  3. 03Permission modelling

    Mapping who can see what, and whether that model can be resolved per user at query time. Ethical walls in one practice could not be enforced by the document system's own API, which made a firm-wide precedent search undeployable until that is remediated. Better found in week three than in month nine.

  4. 04Evaluability screening

    For each surviving candidate: what does correct look like, who adjudicates it, and can we assemble two to three hundred labelled examples? One high-enthusiasm candidate failed here: three senior people gave three incompatible definitions of a good output, and no system can be improved against a target that has not been agreed.

  5. 05Reference architecture

    A target-state design sized to the firm's estate: ingestion, a shared retrieval layer, a shared evaluation harness, and per-use-case extraction schemas on top. The shared layers are what make the second delivery cheaper than the first.

  6. 06Sequencing and costing

    Six surviving candidates ordered so that early work builds reusable capability. Build and run costs modelled per candidate, with the cost driver named, for two of them it was OCR volume, not inference.

What shipped

  • A two-year sequenced roadmap with dependencies, decision points and explicit drop conditions per item.
  • Feasibility findings for all fourteen candidates, including the eight recommended against and why.
  • A reference architecture for a shared retrieval and evaluation layer.
  • A costed and scoped first delivery, ready to start within three weeks of sign-off.

Outcomes

Each figure below carries the method behind it and the baseline it is measured against, which is the form a result has to take before it means anything.

Candidates recommended against
Disqualified on corpus quality, permission model or evaluability before any build budget was committed.
8 of 14
Build budget released
Allocated to the two candidates the firm had been closest to starting, both disqualified on corpus quality during week three.
£410,000
Time from sign-off to first delivery start
Enabled by scoping and costing the first delivery during the roadmap rather than after it.
19 days
Reusable layers identified
Shared ingestion, retrieval and evaluation layers, each serving four of the six surviving candidates, which is what makes the second delivery cheaper than the first.
3 layers, 4 of 6 candidates
The uncomfortable part was being told that eight of our fourteen ideas were not viable, and the useful part was being told in week four rather than after we had funded two of them. The finding about our ethical walls alone justified the engagement: we would have built something we could not have switched on.
Chief Operating Officer, mid-market UK professional services firm

What next

The firm has begun the first delivery, precedent retrieval within a single practice where the permission model resolves cleanly, with the shared retrieval and evaluation layers built to serve the following three.

More engagements

  • Water treatment infrastructure at dusk, concrete channels running into the distance

    Water & utilities

    Making thirty years of asset documentation answerable at a UK water utility

    A retrieval pipeline over 380,000 asset documents: drawings, condition reports, permits and handover packs, with permission inheritance preserved from the source document management system.

    UK regional water utility, ~4,500 staff · 16 weeks

    Read the engagement
  • Repeating stone facade of a government building, shot from below against overcast sky

    Central government

    Assembling regulatory evidence from casework correspondence in central government

    An extraction and evidence-assembly pipeline over eleven years of casework correspondence, built to a standard where every extracted fact traces to a source passage and a pipeline version.

    UK central government agency, ~2,000 staff · 20 weeks

    Read the engagement
  • Glass and steel facade of a City office building reflecting a grey sky

    Financial services

    Putting a measured error rate behind claims-evidence retrieval at a specialist insurer

    An evaluation harness, adversarial suite and assurance pack retrofitted to a retrieval system already in production, so the firm could state its error rate rather than describe its intentions.

    UK specialist insurer, ~1,200 staff · 12 weeks

    Read the engagement

Get in touch

Talk to us.

A first conversation runs about forty-five minutes and covers three things: what your estate actually looks like, whether anyone can define a correct answer or an agreed number, and whether your permission model resolves per user. Any one of them can rule the work out, and we would rather tell you in week one.

Response time
One working day, from a person who has read it.