Evaluation workspace · client-side

Test AI models on the work you actually do.

Bring your own cases and outputs, hide the model names while people score them, then compare human preference with cost and latency context. No account or API key required.

1

Build the test set

Start from a benchr suite, add a case, or import up to 100 cases from CSV or JSON.

The Arabic/Gulf suite contains original synthetic prompts for MSA, Gulf dialect, code-switching, customer service, and RTL formatting.
Suite actionLoading a suite replaces the current evaluation.

CSV headers: title,prompt,criteria,category,language,weight. JSON may be an array or an object with a cases array. Import files must be 5 MB or smaller.

2

Choose 2–5 models

Model labels and current price fields come from the shared benchr model index.

0 of 5 selected

Loading…

When to use it

Use Labs after you have a shortlist and before committing production traffic. Cases sampled from your actual workload are usually more decision-relevant than one broad leaderboard.

Cost and latency

Cost uses the current input/output rates in models.json. Exact token counts improve it; otherwise the tool marks a rough character-based estimate. Latency is measured only when you enter it. Index latency is labelled editorial.

Evaluation limits

A small suite can reveal clear failures, not guarantee production quality. Review edge cases, repeat runs where generation varies, and use multiple qualified evaluators for high-impact decisions.

Frequently asked questions

Does benchr send my prompts or outputs to a server?

No. This evaluation workspace runs in your browser and does not upload or store your prompts, model outputs, scores, or notes. It only downloads benchr's public model and evaluation-suite data files.

Does the lab call the selected AI models?

No. Select models for labels and pricing context, then paste or import outputs generated in your own provider accounts. The lab never asks for an API key.

Are the Arabic and Gulf scores an official benchmark?

No. The included cases are original synthetic evaluation prompts for decision support, not a scientific or provider-certified benchmark. Use your own representative cases and multiple qualified reviewers for important decisions.

What is included in a Labs share link?

Only after you confirm and click Share, the current cases, outputs, scores, and setup are encoded in the URL fragment. Anyone who receives that URL can read the embedded content, so do not share confidential or personal data.