Schema Builder

Tell Docxster exactly what to pull from a document, and how to trust it


Schema Builder is where you define what data Docxster should extract from a document type: the fields, the tables, and the rules each one has to follow. Draw a box on a sample document, tell it what that field is, and every document of that type gets extracted the same way from then on. If you've been hand-checking extracted data because there was no way to enforce what "correct" even looked like, this closes that gap.

How it works


A schema starts with fields, string, number, boolean, or date, and tables with their own columns underneath. You link each one to a document by drawing a box directly on a sample file. That works across multi-page documents too, and zoom and pan won't throw off the box coordinates once you've set them.


A schema can hold more than one sample instance, and each one is tracked separately for whether it's been tested. Run Extraction lets you check the real output against a real document before anything goes near production. Toggle a field or table off and it's skipped during extraction instead of deleted, and toggling one as required means a reviewer downstream can't submit with that value missing or blank.


Two controls sit directly on a field:

  1. Use Constant Value – turn this on for a field and the model stops looking for it in the document. It just returns whatever fixed value you typed in, every time, which works well for something like a cost center code or a department tag that was never going to be in the document anyway.

  2. Validation – attach a regex, either your own or a built-in preset (numbers only, all caps, a length range, an email format). If a reviewer enters something that doesn't match, they get a red border and an inline "doesn't match required format" message, with Approve blocked server-side until it's fixed. An empty value skips the regex check and falls back to whatever the Required toggle says instead.


If only a couple of fields on a document came back wrong, Selective Re-Run lets you rerun just those fields, or just that schema, instead of sending the whole document through extraction again. That saves time and the AI usage a full rerun would burn on fields that were already fine. Long documents get their own fix too: anything six pages or more used to be prone to timeouts and fields silently coming back as "No data found," and a chunking change now splits longer documents smarter at extraction time, without changing how shorter documents are processed.


Once you're working with more than a handful of document types, folders keep them organized. Create one, rename it, move a schema in or out, or delete it, and anything you haven't sorted sits in an "Uncategorized" bucket pinned at the top. Publishing a schema is what makes it available to flows, and every publish is versioned. The underlying active state for a field or table persists correctly through import and export too, so a disabled label doesn't quietly reactivate itself when a schema moves between environments.


Where this is useful

Freight and logistics operations

A BOL number, a PRO number, or a rate confirmation code has to follow an exact format, and a regex validation rule catches a malformed one before it moves downstream, instead of someone catching it after the shipment's already in motion. Mark fields like weight or ship date as required so a reviewer can't approve a load manifest that's missing one.

Customs brokers

An HS code or an entry number needs to follow an exact format, and regex validation catches a malformed one before it ever leaves review. A constant field can carry a fixed processing code that has nothing to do with what's actually printed on the commercial invoice or entry summary, so you're not hunting for a value that was never going to be there.

Lending and mortgage

A loan file has fields that can't go through blank: income figures, borrower names, dates. Mark those required on a loan application or income verification sheet, and a reviewer physically can't approve a file that's missing one, instead of catching it two steps later.

Manufacturing and procurement

Purchase orders and spec sheets often run long, and the chunking fix means a ten-page document doesn't come back with half its fields marked "No data found" the way it used to. Mark critical fields, like part numbers or quantities, as required so a bad extraction gets caught before it reaches a buyer.

On This Page

No headings found

Turn documents into decisions.

See how Docxster gets you from inbox to insight in minutes, not days. Bring your toughest workflow. We'll show you what it looks like solved.