When to use it
Use Labs after you have a shortlist and before committing production traffic. Cases sampled from your actual workload are usually more decision-relevant than one broad leaderboard.
Bring your own cases and outputs, hide the model names while people score them, then compare human preference with cost and latency context. No account or API key required.
Start from a benchr suite, add a case, or import up to 100 cases from CSV or JSON.
CSV headers: title,prompt,criteria,category,language,weight. JSON may be an array or an object with a cases array. Import files must be 5 MB or smaller.
Model labels and current price fields come from the shared benchr model index.
Generate answers in your own provider accounts. Add measured latency and token counts when available.
Model identities use a stable randomized order per case and remain hidden until you reveal results.
The ranking sorts by human score, then estimated cost. Missing data stays missing rather than becoming zero.
Loading…
Use Labs after you have a shortlist and before committing production traffic. Cases sampled from your actual workload are usually more decision-relevant than one broad leaderboard.
Cost uses the current input/output rates in models.json. Exact token counts improve it; otherwise the tool marks a rough character-based estimate. Latency is measured only when you enter it. Index latency is labelled editorial.
A small suite can reveal clear failures, not guarantee production quality. Review edge cases, repeat runs where generation varies, and use multiple qualified evaluators for high-impact decisions.
No. This evaluation workspace runs in your browser and does not upload or store your prompts, model outputs, scores, or notes. It only downloads benchr's public model and evaluation-suite data files.
No. Select models for labels and pricing context, then paste or import outputs generated in your own provider accounts. The lab never asks for an API key.
No. The included cases are original synthetic evaluation prompts for decision support, not a scientific or provider-certified benchmark. Use your own representative cases and multiple qualified reviewers for important decisions.
Only after you confirm and click Share, the current cases, outputs, scores, and setup are encoded in the URL fragment. Anyone who receives that URL can read the embedded content, so do not share confidential or personal data.