> ## Documentation Index
> Fetch the complete documentation index at: https://docs.r3al.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# Evaluate your model

> Compare outputs against examples that were not used for training.

Use the standalone chat to inspect individual responses and the inference API to build repeatable evaluations. Automated experiment execution and a comparison dashboard are not yet connected to the platform.

## Keep evaluation data separate

Prepare representative examples with reviewed expected outputs. Keep all documents from the same customer case or source group in the same split. Freeze a dataset version before running comparisons and avoid using test examples to revise the training recipe.

## Choose metrics for the task

For extraction, measure valid JSON, per-field correctness, missing-value handling, evidence, and unsupported claims. For support responses, review factual correctness, policy adherence, completeness, and placeholder links. Also record latency and failures under the conditions you intend to deploy.

Character or word similarity to a reference can help compare a fixed benchmark. It does not by itself establish factual accuracy or safe behavior on confidential customer documents.

## Reproduce a result

Keep the model revision, registered fingerprint, tokenizer, chat template, prompts, generation settings, dataset version, and evaluation code with each result. Compare a base model and specialized model under the same conditions.

The insurance demo demonstrates inference from a previously fine-tuned model. Its earlier benchmark is external to the platform; the console does not present it as a platform-executed evaluation.


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.