PolyForm

How it works
← Open the app

ENGINEERING DOCUMENTATION

How PolyForm works

A buyer's onboarding form goes in. The same file comes back — filled from one company master record, every answer traced to the fact behind it, in the form's own language. This page is the full picture: the pipeline, how each format is handled, the reasoning behind the design, and how it's deployed.

Overview

What PolyForm is

PolyForm takes an incoming supplier-onboarding form — a PDF, Excel sheet or Word document, usually in German — and fills it automatically from a single structured company record (“master data”). It returns the original file, completed, with a review panel that shows where every answer came from.

The problem

A supplier receives hundreds of these questionnaires. Every buyer uses a different layout; most are in a language the supplier's staff may not read fluently; all are answered today by hand, an hour or more per form, retyping the same company facts over and over. The information is always the same — only the form changes.

The core idea

The hard part is never the data; it's understanding each form's layout and filling it without breaking it. So PolyForm separates the two: a format-agnostic brain that reasons about meaning once, and thin per-format adapters that only know how to read fields out of a file and write answers back in. Add a format → add an adapter; the brain never changes.

You don't read German? You don't have to. Every answer shows the master fact behind it — you verify by fact, not by language.

The pipeline

Seven stages

A form flows through a fixed sequence. The first four stages read the form into a neutral model, stage five is the brain, and the last two write and then verify. Stages 1, 3 and 4 are PDF-specific; the adapters for Office replace them with their own read step.

1
Synthesize PDF only
A flat PDF with no form fields? PolyForm reads the printed layout and creates a virtual widget layer first — so even a scanned-style form becomes fillable.
2
Extract read
Every widget on every page is captured with a stable internal id — the raw anatomy of the form, independent of the publisher's unreliable field names.
3
Label read
Geometry attaches meaning: which printed question does each box actually belong to? Column-aware so dense tables (Straße | PLZ | Telefon) bind correctly.
4
Vision rescue read
Where the layout garbles a label, a vision model reads that region of the page like a human would — but only for the labels geometry couldn't resolve.
5
Map the brain
One batched, schema-enforced LLM call maps the company master record onto the fields. It reasons, it answers in the form's language, and it never invents.
6
Fill write
The real file is filled in place. Only values and appearances change — page count, widgets, fonts, formulas and checkboxes stay exactly as they were.
7
Verify no LLM
Deterministic code traces every written answer back to the master fact it came from, before a human ever sees it. This is the trust signal, not model confidence.
Geometry decides where. Vision decides what. The LLM decides meaning. Deterministic code decides trust.

The format-agnostic brain

Stage five (mapping) and stage seven (verify) read only four things from each field: its id, its kind (text / yes-no / multi-select), its label, and a small extra payload. Nothing format-specific. That's what makes a single brain serve PDF, Excel and Word identically — the null-literal scrub, the unfounded-yes/no backstop, dependency blanking and the grounding cross-check all run unchanged regardless of the source file.

Formats & adapters

Each format is an adapter with two ends — read (file → field model) and fill (answers → file) — wrapped around the shared brain.

PDF

The original target. Native AcroForm PDFs are read directly; flat PDFs with no form fields get a synthesized widget layer from their printed layout. Fill changes only the widget value and appearance — never the page structure. Checkboxes, radio yes/no groups and multi-page chapter logic are all handled.

Excel (.xlsx)

A label is a cell; its answer is the first empty cell to the right of the label's (merge-aware) span. Full-width banners and prose are treated as headings, not fields. The fill writes only the answer cells — embedded logos and images survive the round-trip; the brain leaves anything it can't ground blank.

Word (.docx) + content controls

Word has two field sources, and a real form often mixes them:

  • Plain table cells — a label cell with the answer cell beside it.
  • Content controls (the dominant pattern in real German forms): text controls for free text, and checkbox controls for questions. Two checkboxes in a row under a Ja / Nein header become a single yes-no field; the correct box is ticked by setting its checked state and swapping its glyph to that box's own checked symbol.
Non-destruction is a hard rule. The naïve way to write a Word cell wipes its XML — which would destroy any content controls or checkboxes inside it. PolyForm never does that: it edits only the control's own elements, and a region it doesn't understand (a multi-select group, a dropdown) is left perfectly intact rather than mangled.

On-screen preview

Office files have no native page images, so the filled file is rendered to a PDF by a headless LibreOffice and shown page-by-page in the same viewer as a PDF. The preview is strictly additive: if rendering ever fails, the app falls back to “download to view” — it can never block the fill. Conversions are serialized and memory-capped so they can't affect anything else on the host.

Design principles — the “why”

Right-or-flagged

The product's whole promise is trust. So the contract is absolute: fill what can be proven correct, flag what's uncertain, and never silently drop or invent a value. A wrong answer on a legal supplier form is worse than a blank — a blank is a five-second human fix; a confident wrong answer is a liability the customer only finds later.

Grounding is the trust signal

Model confidence is weak — a model can be confidently wrong. So the review decision is driven by grounding instead: deterministic code traces each answer back to a real master fact. Grounded answers are trusted even when their label was geometrically ambiguous; a filled value that can't be traced is always sent to review. Trust is earned by provenance, not by a probability.

The file is sacred

The filled file is the customer's legal artifact. For PDF the guarantee is absolute — page count, widget set and rectangles are byte-identical to the source; only values and appearances change. For Office it's best-effort (the libraries rewrite the file on save), but we still touch only the answer slots and never the structure. The review panel, by contrast, is our surface — free to redesign.

Deterministic gates

Every behaviour that matters is locked by a no-LLM gate (G1–G24) that runs before and after every change: null-literal scrubbing, dependency blanking, radio binding, the Office adapters, the render-verify gate and the non-destruction guarantee. They're deterministic so a regression fails loudly and instantly, instead of being discovered in a demo.

Trust & verification

The mechanisms that make the output defensible — the part that wins contracts.

Grounding cross-check
Every answer is traced to the master fact it came from. The banner states it plainly — e.g. 38/38 grounded · 0 invented. A filled value that can't be traced is always flagged, however confident the model claimed to be.
Unfounded-answer backstop
A printed “Ja” needs a truthy fact; a “Nein” an explicitly false one. No fact — no answer, decided deterministically after the model, not by it.
Dead-section blanking
When a controlling question is “no”, its dependent section stays completely blank. No data leaks into chapters that don't apply.
Residue stripping
A previous submitter's leftovers are wiped and reported — including a warning when a foreign company name is printed into the form's own prose.
Honest gaps
If the master data can't answer a question, the field stays blank and is surfaced for a human. A blank is honest; a guess is a liability.
Language discipline
Answers come out in the form's own language. An English “no” never leaks onto a German form — the engine enforces it.
Render-verify gate
After filling, the blank and filled pages are rendered and compared pixel-by-pixel: every checkbox must look exactly as intended on the printed page. A form that fails is held for human review — it never ships wrong.

Deployment

Topology

Two deployables. The frontend is a Next.js app on Vercel (auto-deployed on every push to main). The API is a FastAPI service in Docker on a Linux droplet, reached over a private network through a Caddy reverse proxy that terminates TLS — the API has no public port of its own. The mapping stage calls the OpenAI API; the key lives only in an .env on the server, never in git. The API image also bundles headless LibreOffice for the Office preview.

Browser
→
VercelNext.js UI
→
CaddyTLS + routing
→
FastAPIengine · Docker
→
OpenAImapping

Deploy flow

  1. Work on a branch; the gate suite must be green locally.
  2. Merge to main → Vercel rebuilds the frontend automatically.
  3. On the droplet: pull, docker compose build, then up -d — the API image is pinned to the checked-out commit.
  4. Smoke-test the live API (a real upload → filled download) before calling it done.

Operational safety

  • The container has a hard memory limit, so a LibreOffice spike is contained by its cgroup and can't starve the host.
  • The compose project is explicitly named and never torn down with “remove orphans”, so neighbouring services on the box are untouched.
  • CORS is locked to the production origin; uploads are size-capped and type-checked; corrupt or password-protected files return a clean error instead of crashing.

Proof, not promises

241golden answers, hand-authored across real forms
100 · 100 · 95.2% stage-4 accuracy on SOMA · WEBER · ViscoTec — zero mapping errors
G1–G24deterministic gates run before and after every change
3 → 0live field-test rounds, 11 findings driven to zero defects

Accuracy is measured per pipeline stage against hand-authored golden truth, so every miss is attributed to the stage that caused it — that's what makes the system improvable instead of lucky.

Direction — what's next

The business requirement is simple to state: a new form should just work — the customer should never have to call us because a layout we'd never seen broke the tool. This section is honest about where that stands and what we're building next. (Confirmed and locked with the team, July 2026.)

The product promise

From the customer's side there are only three possible experiences with a filled form:

  • It comes back filled, correctly. The goal — and the common case.
  • It comes back filled except a few fields marked “please check these.” Acceptable: thirty seconds of review instead of an hour of typing.
  • It comes back silently wrong and gets sent to a buyer. The one experience that kills trust — and the one the render-verify gate + held-for-review posture is built to make impossible.
The promise is 100% safe, most-of-it automated — and the automated share grows every month. Never “100% magic”.

Why not “100%”?

Because the input space doesn't allow it — for anyone. Some files are corrupt or photographed at an angle; some questions are business decisions (“do you accept our payment terms?”) that a tool should not answer alone; some answers simply aren't in the master data yet. That's why the industry's serious players all ship confidence + human review rather than claiming perfection. What is achievable is 100% trustworthiness: every output is either right or clearly flagged, and the original file is never damaged. That bar is already built and live.

The locked roadmap

Today, structure is understood by geometric rules — proven on the pilot corpus, but each genuinely new layout convention needs engine work (mechanism-level, never per-form). The next architecture inverts this: a vision model reads the rendered page and proposes the structure — fields, pairings, meanings, with coordinates — and deterministic code grounds each proposal to a real widget, fills the original file, and verifies the rendered result. Vision proposes; geometry grounds; the gate judges. The steps, in order:

  1. Extend the render-verify gate to text fields — the safety floor must cover everything before anything new stands on it.
  2. Vision-structure spike — scored side-by-side against the geometric engine on the hardest known forms, with hard pass/fail criteria. If it doesn't clearly win, it doesn't ship.
  3. Review workspace — flagged fields shown on the page image with confidence and provenance, so a human clears a held form in seconds.
  4. Self-correcting fill loop — on a failed verification the engine retries the specific miss before ever involving a human.

The fill itself stays deterministic file surgery, and the gate stays the final judge — intelligence only ever proposes; it never gets to write the file unchecked.

Tech stack

EnginePython · pikepdf · pdfplumber · pypdfium2
Officeopenpyxl · python-docx · headless LibreOffice
IntelligenceOpenAI structured outputs (schema-enforced)
APIFastAPI · streaming progress (NDJSON)
FrontendNext.js · Vercel
InfraDocker · Caddy · DigitalOcean droplet

Built as a working pilot. Multi-tenant accounts, persistence and SSO are deliberate post-pilot work — the engine and its trust model came first.