Skip to main content
Use Starter datasets to inspect provenance, format, license, and limitations. Selecting a starter in setup records the choice. Import an available starter from the project’s Data section to create a reviewable dataset. Neither action starts training.

Reproducible sources

Public starter definitions pin Hugging Face repository revisions and file paths. Review their dataset cards and license terms before using or redistributing data. The synthetic examples contain fictional information and are labeled separately from the public collections.

From a starter to company data

Public receipts are not a substitute for representative company invoices or claim files. Build a reviewed evaluation set from the target task before deciding whether specialization helps. Keep evaluation documents out of training. The data workspace versions imported records, fields, labels, and split assignments. The small built-in samples are for exploring setup, not evidence of production accuracy.

Pilot preparation

The Platform repository includes scripts/prepare_pilot_data.py. It downloads two pinned Parquet files, verifies SHA-256 checksums, saves the publisher cards, and writes JSONL plus a provenance manifest. It does not execute code from a dataset repository. For invoices, unsupported source labels are quarantined rather than used to teach invented dates or currencies. Dates are preserved as written. The remaining 269 records are split by vendor: 213 training, 28 validation, and 28 test. For insurance, only four first-notice fields with exact text evidence enter the labeled set: claim reference, policy reference, claimant, and loss date. The 392 accepted notices are split by claim: 314 training, 39 validation, and 39 test. The 6,606-document corpus remains separately labeled unlabeled. All documents from a claim share its split, and structured policy/rule records are never added to model input. There is no claim approval target. These filters are intentionally conservative; they can exclude valid alternate date formats. Evidence matching is an automated check, not a human audit of label meaning. Synthetic templates can inflate performance, so these sets exercise the pipeline but cannot establish customer-level accuracy. Financial invoice and insurance starter imports require the prepared, checksum-verified artifacts to be installed in the platform environment. If those files are unavailable, import reports an error instead of manufacturing a dataset.

Import your own records

For a project with managed preparation/training, open Data and import UTF-8 TXT or JSONL. The browser accepts files totaling less than 4.5 MB and at most 500 records. PDF, image/OCR, and third-party connector import are not available yet. Each JSONL line uses this shape:
Use group to keep related documents together. The platform assigns train, validation, and test splits by group. The browser importer reads id, group, input, and expected; do not use it to preserve externally assigned split labels.

Review and freeze

Review the source text, expected JSON, and field types. Save corrected outputs and confirm labels. Resolve validation issues before freezing. A frozen dataset must have reviewed labels, valid field values, and nonempty train, validation, and test splits. It becomes immutable; fork it to continue editing. The platform does not automatically send this dataset to the onboarding agent or start fine-tuning. Company imports are blocked when the project requires private preparation.