When to use it
Use Labs after you have a shortlist and before committing production traffic. Cases sampled from your actual workload are usually more decision-relevant than one broad leaderboard.
Choose candidates, use one real prompt, paste their answers, and score them without seeing the model names. Labs does not call model APIs.
Loading…
Start with one real case, use a starter suite, or import up to 100 cases.
CSV headers: title,prompt,criteria,category,language,weight. JSON may be an array or an object with a cases array. Import files must be 5 MB or smaller.
Use 2–5 models for a model comparison, or 1–5 fixed target models for a prompt A/B experiment. Names and current price fields come from the verified benchr model index.
Generate answers in your provider accounts, then bring them here. Latency and token counts are optional.
Model names stay hidden while you score the same criteria for every answer.
Compare human score, estimated cost, and entered latency; then export the result or create an explicit share link.
Recommendations use only paired rubric scores, evaluator evidence, and the structured A/B difference.
The active prompt experiment is restored temporarily in this tab with sessionStorage and is removed by Reset. It is not synced or uploaded. Saving history is a separate, explicit action: it keeps prompt versions, ratings, evidence, and the result on this device, while excluding raw model outputs.
Use Labs after you have a shortlist and before committing production traffic. Cases sampled from your actual workload are usually more decision-relevant than one broad leaderboard.
Cost uses the current input/output rates in models.json. Exact token counts improve it; otherwise the tool marks a rough character-based estimate. Latency appears only when you enter a measured value for that output; the model index does not supply a fallback speed estimate.
A small suite can reveal clear failures, not guarantee production quality. Review edge cases, repeat runs where generation varies, and use multiple qualified evaluators for high-impact decisions.
The prompts and rubrics are published under CC BY 4.0. Inspect or reuse the core suites and the use-case packs. They contain no hidden model outputs or private benchmark scores.
No. The workspace does not upload them. It temporarily keeps an active prompt experiment in this tab's sessionStorage so a refresh can recover it; Reset removes that snapshot. Persistent prompt history is saved only after your confirmation and excludes raw model outputs. The page otherwise downloads only benchr's public model and evaluation-suite data files.
No. Select models for labels and pricing context, then paste or import outputs generated in your own provider accounts. The lab never asks for an API key.
No. The included coding, reasoning, RAG, and tool-use cases are original synthetic evaluation prompts for decision support, not a scientific or provider-certified benchmark. Use representative production cases and multiple qualified reviewers for important decisions.
Only after you confirm and click Share, the current cases, outputs, scores, and setup are encoded in the URL fragment. Anyone who receives that URL can read the embedded content, so do not share confidential or personal data.