The situation
A healthcare provider had a data incident. The most important question was who was affected. The answer sat in around 85,000 scanned faxes, PDFs, screenshots and exported files, with no index and no consistent format. Until each file had been read, the provider could not tell which patients needed to be notified. A notification that reached the wrong people, or missed the right ones, would have been a second incident.
Details in this case study have been altered to protect client confidentiality. The core facts, method and outcomes are accurate.
What we found
Standard e-discovery tools were built for email and office documents. Most of this material was scanned paper: fax images and PDFs with no text layer, along with screenshots and files from legacy systems. The material had to be read as images. Doing that by hand across 85,000 files would have taken months and produced errors nobody could measure.
How we responded
We built the review pipeline for the job rather than forcing the files through a tool that did not fit them.
- Separated the file types. Scanned images, PDFs, screenshots and legacy exports each got their own handling, so nothing was pushed through a step that did not suit it.
- Turned pictures into data. Each scanned page was converted to an image and read by an AI model. The model returned the names, dates of birth and identifiers it found, as structured data.
- Ran two models on every file. A second, independent model read the same page. Where the two agreed the result was accepted. Where they disagreed the file went to an exceptions list for a person to check.
- Matched against the patient record. Every extracted identity was matched to the provider’s patient records system so the client could see, per patient, which documents contained their information.
- Kept the work in Australia. All processing ran on infrastructure in Sydney. No document left the country.
The client received a main report listing the affected patients and the documents behind each one. A separate exceptions report listed the files that needed a human decision.
The outcome
The provider had the list it needed to notify the people who were affected, and only them, on the basis of a documented method rather than an estimate. The review of the document pool took days. The pipeline and its exception handling gave the client a defensible record of what was read, what was matched and what was set aside.
Lessons for similar organisations
- Scanned data is the hardest to search. Faxes and scanned forms hold the most sensitive information and no search tool reads them. Know where they are before you need to.
- Use two independent readers. Cross-checking two models gives you a measured exception rate instead of a guess.
- Ask where the processing happens. Health information should not leave Australia to be read. Ask any provider to say, in writing, where the work runs.
