One screenshot in, a working page out
Hand a model a picture of a screen and get markup and styles back that reproduce the layout, spacing and copy closely enough to edit.
Open the full record →
Each record is one concrete thing a provider documents a model or agent can do — with the steps, the limits, the surfaces it runs on, and the official page benchr read it from. None of it has been tested by benchr yet.
These records describe what providers officially document. benchr has not run these capabilities itself, so nothing here is a test result. Each record shows the vendor page it was read from and the date.
Ledger updated: September 3, 2026
Hand a model a picture of a screen and get markup and styles back that reproduce the layout, spacing and copy closely enough to edit.
Open the full record →
The model gets a live browser: it navigates, reads the accessibility tree, clicks by element reference, fills forms and manages tabs, in a loop, until the task is done.
Open the full record →
Screenshot, move, click, type. The model works a native application the way a person does, for software that never got an API.
Open the full record →
Supply a schema and the model's response adheres to it - no missing required key, no invented enum value, no defensive parser.
Open the full record →
Instead of describing an analysis, the model executes it in a sandbox, reads the actual output, and corrects itself before answering.
Open the full record →
Not OCR. The model reads the text and looks at each page, so a figure that only exists as a chart is still answerable.
Open the full record →
The model decides when to search, the API runs the searches, and the final answer carries source URLs, titles and the exact quoted span.
Open the full record →
One open protocol between an AI application and your systems, so a connector you write once works in every client that speaks it.
Open the full record →
Mark the stable part of a prompt as cacheable and every later request reads it at 0.1x the input price instead of paying full price again.
Open the full record →
Submit many requests as one asynchronous batch, poll for completion, and pay 50% less than the same work sent one request at a time.
Open the full record →
An agent that reads the codebase itself, edits across many files, runs the commands, reads the failures, and opens the pull request.
Open the full record →
A folder with a SKILL.md, optional reference files and scripts. The model reads the name and description always, the instructions only when your request matches.
Open the full record →
Speech to text is the easy half. The useful half is a transcript where each line is attributed to a speaker.
Open the full record →
Not transcribe-then-answer-then-synthesise. One session on gpt-realtime-2.1 where audio goes in and audio comes back, with the turn-taking handled for you.
Open the full record →
The failure everyone remembers from early image models - garbled lettering - is documented as a supported capability now, for infographics, menus, diagrams and marketing assets.
Open the full record →
Text or an image goes in, a short video comes out - and on one of the two documented models, with native audio.
Open the full record →
A whole codebase, a full deposition, an hour of video. The window is real; the retrieval behaviour inside it is the part people get wrong.
Open the full record →
Open-weight models on your own hardware, where the constraint is memory arithmetic rather than a price per token.
Open the full record →
Start the work in the terminal, move it to the cloud, check it from a phone, and have it run again next Tuesday without you.
Open the full record →
Closed to new users. The documentation states the fine-tuning platform is being wound down and is no longer accessible to new users, with existing users able to create jobs temporarily.
Open the full record →
Nothing matches that yet
The ledger holds 20 documented capabilities, so a narrow filter empties quickly. That is the honest state, not a search failure.