# Document intelligence at scale

> An NLP pipeline grading millions of documents a month, replacing five provider systems with one query.

- Role: Architect & technical lead
- Sector: Risk intelligence
- Period: 2014 – 2023
- Engagement: Multi-year build & evolution

A security and risk consultancy monitors people and entities on behalf of its clients, checking for sanctions, fraud allegations, political exposure, adverse media. The traditional way to do that is a point-in-time report: thorough, expert, and stale the day it's delivered.

The starting point was five separate third-party systems, each with its own query language, its own data structures and its own result format. Analysts ran the same search five times and reconciled the answers by hand. That was slow and easy to get wrong, but the real cost was subtler. Departments had each built their own workarounds, so the same underlying data could support different conclusions depending on who went looking.

The brief was to make monitoring continuous. Search many sources, including sanctions and compliance databases, court records and news media, over and over, automatically, and surface only the findings that deserve a human's attention. At the volumes involved, no team of analysts could read everything. The system had to do the reading, so the analysts could do the judging.

## Working out what to build

We focused on the analysts before we touched the pipeline. Structured interviews with staff across departments and seniority levels covered how they ran an investigation and what made them trust a finding, including the tools they had built for themselves to get around the systems they had. The senior analysts had developed workarounds that worked, so we kept them. The newer ones had needs the project could actually meet.

Design sprints then put analysts, department leads and technical stakeholders in the same room, which is where the useful tensions surfaced. What an individual analyst wants is not always what their department wants, and neither is always what the business needs. Better to have that argument in a workshop than to discover it in production.

Alongside that we audited the vendor ecosystem: what each provider could actually do, where they overlapped, where they disagreed, and where they had nothing to offer. That audit shaped the interface and the architecture at the same time, which was the point. Designing them together, rather than one and then the other, is what made a single query model possible at all. None of it was model work. In a system like this, the models are rarely the part that decides whether it succeeds.

## What the work involved

The hard problem was that every source is different. Each database and media feed has its own API, its own formats, its own quirks. The architecture answered that with a microservice per source, each one responsible for exactly one thing: turning that source's output into a single, normalised document structure. Once everything speaks the same shape, everything downstream gets simpler. Keeping each client small and isolated was deliberate. Providers change their APIs without warning, and when one does, the blast radius is a single service rather than the platform.

The same assumption runs into the interface. Providers are slow, they rate-limit, and sometimes they simply don't answer. Rather than hide that behind a spinner, the UI reports what each provider is doing while a query runs and what has come back so far, so an analyst can see that four sources of five have returned and decide whether the fifth is worth waiting for. Treating provider uncertainty as a normal operating condition rather than an error state is a large part of why analysts trusted what came back.

From there, documents flow through a streaming pipeline of queues and workers. Natural-language models score each document against risk categories such as bribery, fraud and political exposure, and when a score crosses a threshold the item is flagged by priority into an analyst's review queue, where it can be assessed and, where warranted, reported onward.

On top of the pipeline sits the tool analysts actually touch: a UI for creating monitoring searches, choosing sources, and expressing the query once, in a single unified structure, rather than once per source. The pipeline does the fan-out.

The system uses Go microservices and AWS SQS with a React front end. It processed over 20 million data points a month by the time I left. Because the architecture is built for streaming, more data means scaling the resources that already exist rather than redesigning the system.

> The pipeline's job was never to replace the analysts. It was to make sure the next document they opened was worth opening.

## Where it stands

The effect I didn't predict was who benefited most. Not the senior analysts, who already knew which source to reach for and what its answers were worth. It was the newer ones, who got that judgement from the tool instead of spending five years accumulating it. Good tooling turned out to be a training lever as much as a productivity one.

This was production document intelligence years before the current AI wave: purpose-built models, scoring real documents, with real consequences riding on the output, and a human decision at the end of every flag. The architecture it forced on us, normalise at the edge, stream everything, keep humans on the judgement calls, is the same architecture I bring to document and AI pipeline work today.

## Focus areas

- Analyst research & design sprints
- NLP classification & scoring
- Per-source ingestion microservices
- Streaming queues (AWS SQS)
- Unified query model
- Analyst review & escalation workflow
- AWS infrastructure & CI/CD

Canonical page: https://blizzard.consulting/work/document-intelligence-at-scale/
