Evaluation workspace · client-side

Test AI models on your own work.

Choose candidates, use one real prompt, paste their answers, and score them without seeing the model names. Labs does not call model APIs.

Loading…

1A

Choose the work to test

Start with one real case, use a starter suite, or import up to 100 cases.

Choose fixed cases for coding, research, writing, spreadsheets, support, video, Arabic translation, RAG, or tool use. Every prompt and rubric is inspectable, editable, and replaceable.
Suite actionLoading a suite replaces the current evaluation.

CSV headers: title,prompt,criteria,category,language,weight. JSON may be an array or an object with a cases array. Import files must be 5 MB or smaller.

1B

Choose target models

Use 2–5 models for a model comparison, or 1–5 fixed target models for a prompt A/B experiment. Names and current price fields come from the verified benchr model index.

0 of 5 selected
This creates the comparison fields. It does not call the models or use an API key.
Local prompt experiment history

The active prompt experiment is restored temporarily in this tab with sessionStorage and is removed by Reset. It is not synced or uploaded. Saving history is a separate, explicit action: it keeps prompt versions, ratings, evidence, and the result on this device, while excluding raw model outputs.

When to use it

Use Labs after you have a shortlist and before committing production traffic. Cases sampled from your actual workload are usually more decision-relevant than one broad leaderboard.

Cost and latency

Cost uses the current input/output rates in models.json. Exact token counts improve it; otherwise the tool marks a rough character-based estimate. Latency appears only when you enter a measured value for that output; the model index does not supply a fallback speed estimate.

Evaluation limits

A small suite can reveal clear failures, not guarantee production quality. Review edge cases, repeat runs where generation varies, and use multiple qualified evaluators for high-impact decisions.

Open test material

The prompts and rubrics are published under CC BY 4.0. Inspect or reuse the core suites and the use-case packs. They contain no hidden model outputs or private benchmark scores.

Frequently asked questions

Does benchr send my prompts or outputs to a server?

No. The workspace does not upload them. It temporarily keeps an active prompt experiment in this tab's sessionStorage so a refresh can recover it; Reset removes that snapshot. Persistent prompt history is saved only after your confirmation and excludes raw model outputs. The page otherwise downloads only benchr's public model and evaluation-suite data files.

Does the lab call the selected AI models?

No. Select models for labels and pricing context, then paste or import outputs generated in your own provider accounts. The lab never asks for an API key.

Are the starter-suite scores an official benchmark?

No. The included coding, reasoning, RAG, and tool-use cases are original synthetic evaluation prompts for decision support, not a scientific or provider-certified benchmark. Use representative production cases and multiple qualified reviewers for important decisions.

What is included in a Labs share link?

Only after you confirm and click Share, the current cases, outputs, scores, and setup are encoded in the URL fragment. Anyone who receives that URL can read the embedded content, so do not share confidential or personal data.