Reproducible sources
Public starter definitions pin Hugging Face repository revisions and file paths. Review their dataset cards and license terms before using or redistributing data. The synthetic examples contain fictional information and are labeled separately from the public collections.From a starter to company data
Public receipts are not a substitute for representative company invoices or claim files. Build a reviewed evaluation set from the target task before deciding whether specialization helps. Keep evaluation documents out of training. The data workspace versions imported records, fields, labels, and split assignments. The small built-in samples are for exploring setup, not evidence of production accuracy.Pilot preparation
The Platform repository includesscripts/prepare_pilot_data.py. It downloads two pinned Parquet files, verifies SHA-256 checksums, saves the publisher cards, and writes JSONL plus a provenance manifest. It does not execute code from a dataset repository.
For invoices, unsupported source labels are quarantined rather than used to teach invented dates or currencies. Dates are preserved as written. The remaining 269 records are split by vendor: 213 training, 28 validation, and 28 test.
For insurance, only four first-notice fields with exact text evidence enter the labeled set: claim reference, policy reference, claimant, and loss date. The 392 accepted notices are split by claim: 314 training, 39 validation, and 39 test. The 6,606-document corpus remains separately labeled unlabeled. All documents from a claim share its split, and structured policy/rule records are never added to model input. There is no claim approval target.
These filters are intentionally conservative; they can exclude valid alternate date formats. Evidence matching is an automated check, not a human audit of label meaning. Synthetic templates can inflate performance, so these sets exercise the pipeline but cannot establish customer-level accuracy.
Financial invoice and insurance starter imports require the prepared, checksum-verified artifacts to be installed in the platform environment. If those files are unavailable, import reports an error instead of manufacturing a dataset.
Import your own records
For a project with managed preparation/training, open Data and import UTF-8 TXT or JSONL. The browser accepts files totaling less than 4.5 MB and at most 500 records. PDF, image/OCR, and third-party connector import are not available yet. Each JSONL line uses this shape:group to keep related documents together. The platform assigns train, validation, and test splits by group. The browser importer reads id, group, input, and expected; do not use it to preserve externally assigned split labels.

