Invoice example
null rather than be guessed.
Claim intake
Extract facts such as claim reference, policy reference, incident date, incident type, and claimed amount. Keep extraction separate from coverage, eligibility, or payment decisions. Reviewers handle those decisions.Prepare examples
Pair each document with a reviewed expected output. Include difficult layouts, absent fields, multiple totals, and ambiguous dates. Group related documents when splitting data so pages from the same case do not appear in both training and evaluation. Scans require OCR or a model that accepts images. Do not assume a text-only model can read PDFs directly. Record the parsing or OCR version so errors can be traced to the source stage.Measure the result
Check schema validity, per-field accuracy, numeric/date normalization, missing-value behavior, and evidence alignment. Review failures individually; an average score can hide costly errors in totals or policy references.The data workspace supports typed fields, label review, and frozen dataset versions. Document parsing/OCR, training, and automated evaluation execution are still being integrated. The current insurance chat demo is a support-response model, not an extraction engine.

