Local AI reference · updated 15 August 2026

Plan local AI without pretending every machine is the same.

Benchr connects provider facts—parameters, license, context, and published memory figures—to a visible local-deployment method. It is a planning reference, not a benchmark lab or a download directory.

Two concrete local decisions

Use the Explorer when you need to inspect a record. Use the planner when you know the memory available and need an honest shortlist with trade-offs.

Reference database

Explore local model records

Filter source-linked initial records by size, license, modality, native context, and documented runtime route. Every factual row leads back to a provider source.

Open the local model explorer →

Interactive planner

What can I run?

Choose NVIDIA VRAM, Apple unified memory, or a CPU-only memory path, then see what has comfortable headroom, what is tight, and what is not recommended at your context target.

Plan this machine →

Evidence boundary

Weights are not speed

A model may technically load and still be too slow, too memory-constrained at long context, or unsuitable for your task. No token-per-second or quality claim appears without a reproducible measurement.

Read the measurement guide →

What the records carry

The Local AI index is normalized against the verified model ledger. It adds local planning metadata without creating a second source of truth for model numbers.

ParametersTotal and active counts are kept distinct for MoE models.
LicenseProvider terms are shown as published; “open weight” is not silently equated with OSI open source.
ContextNative and documented extended limits are separate, because the latter may need runtime configuration.
MemoryProvider Q4 figures are preserved; otherwise the tool exposes simple Q4 weight arithmetic.

The planning boundary

A compatibility label is not a promise that a particular app, quantization, prompt, or operating system will behave well.

What this planner can say

It can show whether a disclosed or calculated Q4 weight leaves declared memory headroom at 4K, 8K, 32K, or 128K planning contexts. It cannot measure model-specific KV cache, output rate, quality, thermal behavior, or disk use. A tight result is a warning to test deliberately, not an invitation to oversell a machine.

Why no universal “runs on 16GB” badge? A quantization’s metadata, GPU layers, runtime buffers, context length, and offload configuration all move the answer. The planner makes its safety boundary inspectable and sends you to the exact model source before you download anything.

Questions this reference answers

Is a local model free to run?

The weights may have no per-token API charge, but local deployment still has hardware, electricity, storage, maintenance, testing, and review costs. License terms also differ by provider and model.

Does an MoE model need memory for only its active parameters?

No. Active parameters can help explain compute per token, but the full set of weights still matters for storage and loading memory. The local planner uses total parameters when it must calculate a Q4 weight estimate.

Why are GGUF and MLX not shown as a universal availability badge?

They are artifact and runtime ecosystems, not a permanent model-provider guarantee. This first release links documented runtime routes and keeps provider weights separate from community conversions, which need their own verification.