Parameters
Total vs active
For MoE models, active parameters explain a compute path, not necessarily how many weights must be kept available. The table never turns “active” into a memory shortcut.
Local AI reference · source-linked records
Find a starting point by actual deployment constraints—parameter count, terms, context, modality, and a provider-linked runtime route. The memory column is a Q4 planning weight, not a file-size claim or speed result.
Filter state is retained in the URL so a shortlist can be shared without an account. Factual fields are joined from model-figures.json.
Loading source-linked records…
| Model | Parameters | License | Context | Q4 weight plan | Routes and modalities |
|---|---|---|---|---|---|
| The source-linked local model index is loading. | |||||
Parameters
For MoE models, active parameters explain a compute path, not necessarily how many weights must be kept available. The table never turns “active” into a memory shortcut.
Context
A documented extension is displayed after the native limit. It can need YaRN or another runtime setting, can add memory pressure, and may carry a quality trade-off.
Q4 planning
Provider Q4 memory figures remain provider-labelled. Otherwise the result is exactly four-bit weights × total parameters; metadata and runtime overhead stay outside it.
This is a deployment reference. It does not invent a cross-model quality score from parameter size or provider benchmarks. Build a shortlist here, then run the same representative cases through the exact setup you would deploy.
Not as a universal provider-availability claim. The table intentionally distinguishes provider weight routes from ecosystem conversions; later format coverage must be artifact-specific and source-verified.
No. It is a starting documentation route, not a generated configuration. Read the provider or runtime instructions for the exact model revision, platform, and context setting.