Beyond Text: Base64 Receives U.S. Patent for Multi-Modal Document Classification

A document’s words can tell you a lot. Its layout, visual features, and barcode can tell you more.

We are pleased to announce that Base64.ai has been awarded U.S. Patent No. 12,743,902, “Multi-Modal Document Type Classification Systems and Methods.” The patent recognizes an approach to identifying document types using more than extracted text alone.

Why text alone is not always enough

Before an organization can extract the right information from a document, it needs to know what kind of document it received. That sounds straightforward until a workflow encounters similar forms, unfamiliar layouts, poor scans, or documents with little readable text. Consider an identification card. A name, date of birth, and address may suggest one document type, but those words do not tell the whole story. A portrait, signature, overall shape, or barcode can provide additional context. Looking at these signals together helps a system determine whether the document matches the type it appears to be.

How multi-modal classification works

The patented approach describes an AI system that can analyze several kinds of evidence in a document, including:

  • Text: Words and fields captured through optical character recognition (OCR).

  • Visual features: Elements such as portraits, signatures, markings, and their positions on the page.

  • Document shape: The overall form of the document compared with what is expected for a candidate type.

  • Barcodes: Encoded information that can help identify or check a document type when present.

For example, text might suggest that an uploaded image is a driver’s license. The system can then compare its visual features and shape with what it expects from that document type. If the signals do not align, the workflow can flag the inconsistency for review instead of relying on the text alone. The patent also describes using barcode data as another route to identify and check a document.

Why this matters for document workflows

Classification is the first decision in many document processes. It determines which fields to extract, which rules to apply, and where a document should go next. A stronger classification step can give downstream systems better context and help teams focus attention on uncertain cases. This is especially relevant when organizations receive a mix of IDs, forms, receipts, invoices, and other documents through the same channel. Each may require different handling, even when some of the text looks similar.

The patent reflects our team’s work on a simple idea: documents communicate through more than words. By considering text alongside visual and encoded signals, document AI can make a more informed first decision and support workflows built around the document actually received.

Explore Base64.ai to learn more about intelligent document processing.

U.S. Patent No. 12,743,902, “Multi-Modal Document Type Classification Systems and Methods,” issued September 22, 2026. View the USPTO patent record.