Reference database
Explore local model records
Filter source-linked initial records by size, license, modality, native context, and documented runtime route. Every factual row leads back to a provider source.
Open the local model explorer →Local AI reference · updated 15 August 2026
Benchr connects provider facts—parameters, license, context, and published memory figures—to a visible local-deployment method. It is a planning reference, not a benchmark lab or a download directory.
Use the Explorer when you need to inspect a record. Use the planner when you know the memory available and need an honest shortlist with trade-offs.
Reference database
Filter source-linked initial records by size, license, modality, native context, and documented runtime route. Every factual row leads back to a provider source.
Open the local model explorer →Interactive planner
Choose NVIDIA VRAM, Apple unified memory, or a CPU-only memory path, then see what has comfortable headroom, what is tight, and what is not recommended at your context target.
Plan this machine →Evidence boundary
A model may technically load and still be too slow, too memory-constrained at long context, or unsuitable for your task. No token-per-second or quality claim appears without a reproducible measurement.
Read the measurement guide →The Local AI index is normalized against the verified model ledger. It adds local planning metadata without creating a second source of truth for model numbers.
A compatibility label is not a promise that a particular app, quantization, prompt, or operating system will behave well.
It can show whether a disclosed or calculated Q4 weight leaves declared memory headroom at 4K, 8K, 32K, or 128K planning contexts. It cannot measure model-specific KV cache, output rate, quality, thermal behavior, or disk use. A tight result is a warning to test deliberately, not an invitation to oversell a machine.
Why no universal “runs on 16GB” badge? A quantization’s metadata, GPU layers, runtime buffers, context length, and offload configuration all move the answer. The planner makes its safety boundary inspectable and sends you to the exact model source before you download anything.
The weights may have no per-token API charge, but local deployment still has hardware, electricity, storage, maintenance, testing, and review costs. License terms also differ by provider and model.
No. Active parameters can help explain compute per token, but the full set of weights still matters for storage and loading memory. The local planner uses total parameters when it must calculate a Q4 weight estimate.
They are artifact and runtime ecosystems, not a permanent model-provider guarantee. This first release links documented runtime routes and keeps provider weights separate from community conversions, which need their own verification.