Run and improve
Evaluate your model
Compare outputs against examples that were not used for training.
Use the standalone chat to inspect individual responses and the inference API to build repeatable evaluations. Automated experiment execution and a comparison dashboard are not yet connected to the platform.

