Use this guide for source development and local checks. Contribution policy lives in Contributing; deployment branch management and publishing are consolidated in Space deployment.
Read the repository's AGENTS.md and the guidance in the component you change.
Use Python 3.11 and uv for evaluator/, and Node.js 22.22.2/npm for viewer/.
Install dependencies in the checkout where you work:
# From evaluator/
uv sync --locked# From viewer/
npm ciInspect git status --short and git worktree list before switching branches.
Worktrees share Git history, but dependencies, ignored data, and caches are local
to each checkout. Preserve the tracked viewer/data symlink to ../evaluator/data.
Use explicit result-directory arguments to read measurements stored elsewhere.
Ordinary source contributions require no deployment access or deployment worktree.
| Change | Guide |
|---|---|
| Model integration | Adapter contract and testing |
| Evaluation behavior | Evaluation runbook and scoring specification |
| UI or local result comparison | Viewer setup |
| Display loading and recovery | Prepared display data |
| Result publication | Dataset PR workflow |
| Deployment branch or hosted viewer | Space deployment |
To develop against an existing results folder, run from viewer/:
npm run prepare-display -- --results-dir /absolute/path/to/results
npm run devThe wrapper prints a Tailscale IPv4 or localhost URL. See the viewer guide for synchronization, combining sources, Docker, and restart behavior. The viewer reads saved metrics; it does not evaluate models or download Hub updates itself.
Keep downloaded datasets, model weights, measurements, credentials, and generated reports out of source commits. Standard ignored locations include:
| Content | Location |
|---|---|
| Evaluation inputs | evaluator/data/datasets/ |
| Original local runs | evaluator/data/results/<run-id>/ |
| Verified Hub snapshots | evaluator/data/hub-results and evaluator/data/.hub-results* |
| Recorded dataset revisions for validation | evaluator/data/result-datasets/ |
| Audits and scratch reports | evaluator/audits/, tmp/, output/, evaluator/output/ |
| Prepared display JSON and reports | viewer/display/ |
Keep reports and auxiliary JSON outside result roots. Custom output destinations
must be outside the checkout or have their own ignore rules. Keep benchmark and
category definitions and shared synthetic fixtures tracked; do not ignore all of
evaluator/data/ or all JSON/XZ files. Configuration names belong in
.env.sample; real credentials belong in ignored local files.
From viewer/:
npm test
npm run typecheck
npm run buildThe build includes Storybook, served under /storybook/ by the production viewer.
Use synthetic stories to review visual changes at desktop and mobile sizes,
including comparison, model details and loading states.
From evaluator/:
uv sync --locked
uv run toxPublic CI must work without private data, credentials, model APIs or GPU access. Use the dedicated dataset integration tests when authorized data is available; do not make ordinary unit tests depend on it.
Describe the problem, resulting behavior, and validation in the source PR. Keep measurements in a separate results Dataset PR and link related code changes. Before a source release, follow RELEASING.md, including review of the release file list and Git history. Ignore rules do not remove previously committed artifacts.