> ## Documentation Index
> Fetch the complete documentation index at: https://docs.r3al.ai/llms.txt
> Use this file to discover all available pages before exploring further.

# On-demand compute

> Run a configured model for a demo without leaving a GPU on all day.

The managed pilot uses RunPod Serverless behind the platform. R3AL configures the endpoint and model connection; you control its demo session from the project.

## Session lifecycle

1. **Off:** the demo has no active lease and worker limits are zero.
2. **Preparing:** a worker starts, obtains the checkpoint, verifies files, and loads the model.
3. **Ready:** the expected model is loaded and can answer requests.
4. **Stopping:** the platform requests shutdown and waits for provider confirmation.
5. **Off:** provider shutdown is confirmed.

Choose a duration before **Start demo**. The UI offers 15, 30, and 60 minutes. The automatic deadline runs on the server, so it does not depend on leaving your browser open. Use **End demo** to finish early.

## Loading time and cost

Cold starts can take 5–10 minutes. GPU capacity, download speed, and checkpoint size affect this time. Compute is billed while warming and serving; loading is part of the selected duration. Serverless improves scheduling flexibility but does not guarantee immediate GPU capacity.

Stopping compute preserves the model artifacts. Artifact storage may still incur charges. The setup budget is separate from the demo's duration limit and is not a hard spending cap.

## Customer infrastructure

Project setup records customer-owned preparation/training and serving requirements independently. Self-service Linux/NVIDIA/Docker runner installation is not available yet. A customer runtime requires a separately configured and verified connection.

A saved location preference does not by itself establish a private processing boundary. Read [Privacy](/platform/privacy).


This documentation is built and hosted on [Mintlify](https://mintlify.com), a developer documentation platform.