Local AI planner · transparent assumptions
What can I run?
Choose the memory that actually constrains your local setup. The result separates comfortable planning headroom from a tight load and from a configuration we do not recommend. It does not claim performance or model quality.
Plan the starting configuration
The URL updates as you change a control. Share it as a planning conversation, then verify the exact artifact and runtime before treating the outcome as a deployment decision.
Loading the source-linked planner…
Method and limits
This is deliberately a conservative planning tool rather than an imaginary universal benchmark.
Formula shown in every result
For records without an official Q4 memory value, Q4 planning weight = total parameter count × 0.5 GB. It is the theoretical four-bit payload in decimal gigabytes. The tool requires 2 GB of remaining planning headroom at 4K, 3 GB at 8K, 6 GB at 32K, and 12 GB at 128K. Those boundaries are BenchR policy choices exposed in the page, not provider figures.
Why ranges and caveats matter: weights are one component. Runtime buffers, KV cache, GPU layer split, operating-system use, model revision, quantization metadata, and prompt shape can change the real result. “Comfortable” means only that this limited planning check leaves its visible headroom; it does not mean fast, stable, or high quality.
Before you act on a result
Why is my preferred model not listed as comfortable?
It may be larger than the selected memory once the Q4 weights are considered, or the remaining room may not meet the selected context boundary. A tight or not-recommended result does not prove it can never run; it means BenchR will not describe that configuration as a healthy default.
Does Apple unified memory work like GPU VRAM?
No. The tool uses the selected amount as an envelope for a transparent weight-headroom check, not as a claim that Apple unified memory behaves identically to discrete VRAM. Bandwidth, memory pressure, runtime, and other applications matter.
Does context length change the outcome?
Yes. Larger contexts require more cache memory, and some models publish an extended limit only with a special configuration. The planner increases its visible headroom boundary at 32K and 128K and flags a requested context beyond the native limit.