DOCUMENT INTELLIGENCE / OCTOBER 2026

Why I Built Our Own KYC Flow: OCR, MRZ Parsing and the Business Decision Behind It

The trigger was not a technical failure, and we were not deciding what regulation required. The regulated banking and compliance programme defined the KYC requirements we had to satisfy. The mismatch was that our verification provider's hosted flow enforced an additional check outside that required path and did not let us remove it. Once that extra step started adding friction, the product decision was to own the journey while preserving every requirement imposed by the regulated programme.

Passport, phone and headphones arranged on a desk
Identity documents increasingly sit at the boundary between physical evidence and software. Photo by Alex Robert on Unsplash.

I did not set out to build a proprietary KYC flow. I set out to remove a piece of onboarding friction that should not have existed in the first place. The OCR, document parsing, machine-readable-zone work and guided capture came later. The first decision was simply that our customer journey should implement the KYC requirements set by our regulated banking and compliance programme, rather than inherit additional requirements from a third-party workflow that were not part of that programme.

A note on disclosure.

This is a product and engineering field note. I have intentionally left out the names of the financial product, verification provider, banking provider, private infrastructure, internal routes, operational thresholds and customer information. The useful part of the story is the decision process and the document-recognition engineering, not the private integration details.

The problem was business, not code

Our first approach was sensible: use a specialist identity-verification platform for the KYC journey and integrate its result into the rest of the financial application. In theory, that meant less custom work and a faster path to production. The provider positioned configurable verification journeys as one of the important parts of its product, so I expected us to be able to shape the flow around the requirements of our own banking relationship.

That assumption eventually broke down. A particular verification requirement was enforced in the provider-managed journey even though it was not required by the banking provider for the product path we were implementing. More importantly, we could not disable that step for our use case. The configuration interface gave us control over several parts of the flow, but not enough control over the part that mattered.

We did not decide the regulatory KYC requirements. Those requirements came from the regulated banking and compliance programme we were integrating with, and our job was to implement them faithfully. The change was not an attempt to lower, reinterpret or bypass that standard. It was a decision to stop carrying an extra vendor-imposed workflow requirement that sat outside the programme we were required to satisfy.

I do not think that makes the provider wrong. A verification company has to design defaults and risk controls for many customers, markets and compliance programmes. But our responsibility was different. We were responsible for the experience our own customers would go through while still meeting the requirements set by the regulated programme. If a customer had to complete an additional step only because a vendor's workflow could not be adapted, that friction still belonged to us.

That was the moment the question changed. It was no longer, “How do I integrate this provider correctly?” It became, “How do we implement the verification journey required by our regulated banking and compliance programme without forcing customers through an extra vendor-specific step that the programme itself does not require?”

The decision came before the implementation

This distinction matters because it would be easy to retell the story as an engineering exercise: I wanted to build OCR, so we replaced a hosted KYC flow. That is backwards.

The directive was a business one: make onboarding easy for customers while satisfying every KYC requirement defined by the regulated banking and compliance programme.

Technology had to implement that directive.

That gave me a better decision rule than “build versus buy.” We could still buy specialist capabilities. We could still ask an external identity service to validate an identifier or return an authoritative result. We could still rely on our banking partner for the financial capabilities and compliance decisions it owned. What we did not need to outsource was the entire sequence the customer saw.

External services can verify facts. The application should own the journey. That separation let us keep specialist verification where it was useful while removing an additional workflow requirement that was not part of the regulated programme. The regulated provider still made the final acceptance decision for the banking product.

Owning the flow did not mean reinventing identity infrastructure

“Build our own KYC” can sound much more dramatic than it is. I was not trying to recreate government identity databases, build a biometric research lab, define regulatory policy or claim that our software could independently establish legal identity from first principles.

The proprietary part was the orchestration and product layer: how we present and collect the information and evidence required by the programme, how a document is captured, how we extract structured fields, how we recover from unreadable evidence, how we present progress, and when an authoritative external verification needs to be called.

That separation was important because it kept the system honest. OCR can read text. An MRZ parser can validate a machine-readable travel-document format. A geocoding service can help normalize an address. None of those things, by themselves, mean “this person is verified.” They are components in a larger evidence and decision process, and they do not replace the acceptance decision made by the regulated financial provider.

The architecture therefore became less like a hosted widget and more like a document-intelligence pipeline:

CaptureUpload or guided camera capture with a clear document frame.
PrepareCorrect perspective and improve the image for machine reading.
RecognizeRun OCR and select a parser based on the document class.
ValidateCheck MRZ structure, labels, dates and other document-specific constraints.
ContinueUse structured evidence in the broader verification journey.
Passport resting on a computer keyboard
A passport is unusually useful for automated extraction because the machine-readable zone is deliberately standardized. Photo by Oxana Melis on Unsplash.

OCR became the centre of the work

Once we owned the screens, the next obvious question was whether we could stop asking users to type information that was already printed on the document sitting in front of them.

That is where optical character recognition became more than a convenience feature. If the system could reliably extract document fields, then the user could spend less time copying names, document numbers, dates and addresses into forms. The application could prefill what it recognized, ask the user to confirm where appropriate, and reserve manual input for the places where automation was not confident enough.

My initial OCR path worked well enough to prove the idea, but it also exposed an important lesson: general OCR and document understanding are not the same problem.

An OCR engine gives you text and confidence. It does not inherently know that one line is a surname, another is an issuing authority, another is a location, and two dense lines at the bottom of a passport are a machine-readable data structure with checksums. If you treat every document as a page of unrelated text, you end up writing heuristics that work until the first unfamiliar layout arrives.

So I stopped thinking in terms of “read the image” and started thinking in terms of “recognize the document, then use the grammar of that document.”

Why the machine-readable zone changed the passport pipeline

Passports gave me the clearest example of that idea. The lower part of a passport biodata page contains a machine-readable zone, or MRZ. ICAO Doc 9303 defines the formats used by machine-readable travel documents, including the constrained character set, field positions and check digits that make the data machine-verifiable.

That is much stronger than asking OCR to guess which uppercase line “looks like a name.”

The passport pipeline became deliberately MRZ-first:

01

Find the machine-readable region

Use the document geometry to focus on the lower code zone instead of running the same extraction logic over the whole page.

02

Prepare it for OCR

Perspective correction, grayscale conversion, contrast adjustment, thresholding and deskewing make the characters easier to separate from the background.

03

Constrain recognition

The MRZ alphabet is intentionally narrow: uppercase letters, digits and the filler character. Constraining recognition reduces the space of plausible OCR mistakes.

04

Parse the format

Once the lines are recognized, parse them according to the applicable TD1, TD2 or TD3 structure rather than treating the output as prose.

05

Validate check digits

A parsed field is not accepted just because it looks plausible. The MRZ contains check digits that let the software detect many recognition errors before using the extracted values.

After a valid parse, the software can obtain structured fields such as the holder's names, document number, nationality, date of birth and expiry information from the document's own standardized machine-readable representation.

That changed the quality of the system because it converted OCR from a free-form text problem into a constrained recognition-and-validation problem.

The failure that killed my generic name heuristic

One of the most useful failures came from a passport that OCR could read, but not correctly enough for the original parser.

The MRZ was not successfully parsed. The fallback logic then looked elsewhere on the page for an uppercase line that might represent the holder's name. It selected part of an authority or location line instead. The resulting value contained “LAGOS” and was confidently wrong.

The downstream identity comparison rejected it, which was the correct final behaviour, but that was not good enough for me. A safe rejection can still expose a bad extraction design. The system had reached the right outcome for the wrong reason.

I removed the idea that a passport could fall back to an arbitrary uppercase line. The rule became much narrower: a passport name comes from a valid MRZ, or from explicit printed surname and given-name labels when those labels can be read reliably. If neither path works, the honest result is that the document needs another capture or review.

That single bug changed the extraction architecture beyond passports. It made me much less interested in clever generic heuristics and much more interested in document-specific evidence.

Laptop displaying source code in Lagos
Most of the work was not “AI magic”; it was deterministic image preparation, parsing, validation and careful fallback design. Photo by Desola Lanre-Ologun on Unsplash.

Guided capture improved the input before I tried to improve the OCR

There is only so much software can recover from a bad photograph. A passport cropped through the MRZ, a card covered in glare or a document photographed from a severe angle can defeat a good OCR model before recognition even starts.

So the next part of the system was a reusable guided document scanner. Instead of a plain camera button, the camera experience could detect the document boundary, show the user a framing guide, measure common quality problems and correct the selected frame before upload.

The useful checks were mundane but effective: is the whole document inside the frame, is it large enough, is the image badly blurred, is there severe glare, is exposure usable, is the page excessively skewed, and has the device been stable long enough to capture a clean frame?

OpenCV.js provided the computer-vision primitives, and a lightweight document-scanning layer helped with contour detection and perspective correction. The important design decision was not the library choice. It was keeping capture separate from verification. The camera experience exists to produce a better image; it does not get to declare the person verified.

That also made the experience more forgiving. Upload remained useful for people who already had a clean scan or PDF. If OCR failed because the image quality was poor, the application could offer guided camera capture as a recovery path instead of returning a dead-end “document unreadable” message.

Different documents need different parsers

The passport work made it obvious that one extractor should not pretend every identity document has the same grammar.

For MRZ-equipped documents, the machine-readable zone should be the primary structured source. For national identity cards, licences and similar cards without a usable MRZ, the parser can use OCR words with bounding boxes, recognize labels such as surname, given names, date of birth, document number or expiry, and then inspect the text spatially beside or below those labels.

Proof-of-address documents are different again. The useful fields are more likely to be the issuer, document category, holder or customer name, address block and document date. The parser therefore needs to understand the semantics of a statement or bill rather than force it through an identity-card template.

PDFs also deserved their own route. If a PDF already contains extractable text, there is no reason to rasterize it immediately. If it is a scanned PDF, then rendering the relevant page for OCR makes sense. Only after those paths fail should the user be asked to capture the physical document again.

The pattern is simple: shared OCR engine, document-aware interpretation. That is more reliable than multiplying unrelated OCR stacks, and much safer than using one giant heuristic parser for every possible document.

Address intelligence removed another form users should not have to fight

Document recognition was only one source of friction. Address entry is another place where onboarding forms routinely make users do unnecessary work.

I added an address-search layer backed by a geocoding API so the user could begin typing an address and select a normalized candidate instead of manually reproducing every component. The service could return structured locality, region and country information alongside the formatted address. Where a user explicitly permitted location-based assistance, reverse lookup could also help propose a nearby address for confirmation.

The important word there is confirmation. Geocoding is useful for search and normalization; it is not proof that a person lives at a location. The address result still belongs inside the wider evidence flow, especially when proof-of-address documentation is required.

From a product perspective, though, this was a meaningful improvement. Every field you can safely infer, extract or search is one less field the customer has to type. KYC often feels difficult because teams treat friction as evidence of seriousness. I think the better standard is the opposite: verification should be rigorous where it needs to be and nearly invisible everywhere else.

Retries are part of the product, not an error screen

Owning the flow also meant owning failure states.

“Document unreadable” is technically accurate and practically useless. The system knows more than that. It can often distinguish a document that is too blurry from one that is cropped, too dark, affected by glare, too small, missing a visible MRZ, or simply the wrong document type.

I turned those differences into structured retry reasons that the interface could translate into useful instructions. If the image is blurry, ask for a steadier capture. If the MRZ is cut off, tell the user to include the code lines at the bottom of the passport. If the document type is wrong, do not send the user into a camera loop; ask for the correct evidence. If a valid document is ambiguous, route it to the appropriate review path rather than inventing certainty.

This is one of the least glamorous parts of KYC engineering and one of the most important. A verification journey is not defined by the happy path. It is defined by how quickly a legitimate customer can recover when the happy path fails.

What I took away from building it

The most important lesson was not about OCR. It was about product ownership.

A third-party workflow can be technically excellent and still be wrong for your product if it forces requirements that are not part of the regulated banking and compliance programme you actually have to satisfy. That is not permission to invent your own KYC standard. We did not define the regulatory bar; we implemented the bar set for the programme and removed an additional vendor-imposed step that sat outside it.

The second lesson was that document intelligence gets better when it becomes less generic. Passports have an MRZ grammar. Cards have labels and spatial relationships. Statements have issuers, dates and address blocks. A parser that respects those structures can be simpler and more reliable than a supposedly universal heuristic.

The third lesson was that better input usually beats more sophisticated recognition. Framing, perspective correction, blur detection and clear retry guidance can improve real-world OCR before you reach for a larger model.

And the final lesson was that automation should remove typing, not remove judgment. OCR can extract. MRZ checksums can validate structure. A geocoding API can normalize. A camera can improve capture. But the application still needs a clear authority model for what those signals mean, and the regulated provider still made the final acceptance decision for the banking product.

What started as a decision to remove one unnecessary onboarding requirement ended up becoming a much better identity-document system. That is the part I find interesting: the technology was not the reason for the change. The technology became interesting because the business decision forced us to build the experience we actually wanted customers to have without changing the regulated KYC requirements we were obligated to satisfy.

Sources and further reading

  1. ICAO Doc 9303, Part 3: Specifications Common to all Machine Readable Travel Documents
  2. Tesseract.js: OCR for the browser and Node.js
  3. OpenCV.js: contours and document-shape primitives
  4. jscanify: browser document detection and perspective correction
  5. cheminfo/mrz: TD1, TD2 and TD3 MRZ parsing
  6. MDN: MediaDevices.getUserMedia()