> ## Documentation Index
> Fetch the complete documentation index at: https://docs.r3al.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Document extraction

> Define reliable structured outputs for financial and insurance documents.

Start with a narrow document family and an explicit output contract. Invoice extraction and claim intake are good pilot tasks because their fields can be checked against reviewed examples.

## Invoice example

```json theme={null}
{
  "invoice_number": "INV-2041",
  "issuer": "Cedar Office Supplies",
  "invoice_date": "2026-08-12",
  "total": 247.50,
  "currency": "EUR"
}
```

This is a fictional example. Decide whether dates must follow ISO 8601, how currencies are represented, and whether amounts include tax. A missing currency should remain `null` rather than be guessed.

## Claim intake

Extract facts such as claim reference, policy reference, incident date, incident type, and claimed amount. Keep extraction separate from coverage, eligibility, or payment decisions. Reviewers handle those decisions.

## Prepare examples

Pair each document with a reviewed expected output. Include difficult layouts, absent fields, multiple totals, and ambiguous dates. Group related documents when splitting data so pages from the same case do not appear in both training and evaluation.

Scans require OCR or a model that accepts images. Do not assume a text-only model can read PDFs directly. Record the parsing or OCR version so errors can be traced to the source stage.

## Measure the result

Check schema validity, per-field accuracy, numeric/date normalization, missing-value behavior, and evidence alignment. Review failures individually; an average score can hide costly errors in totals or policy references.

<Note>
  The data workspace supports typed fields, label review, and frozen dataset versions. Document parsing/OCR, training, and automated evaluation execution are still being integrated. The current insurance chat demo is a support-response model, not an extraction engine.
</Note>


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.