The best open-weight model for each GPU, ranked by intelligence, speed and context. Click a GPU to see its full ranking.
A model is listed when one of its builds fits entirely in VRAM: build size + KV cache + 0.6 GB. Builds are the quantized files Hugging Face links to the model (GGUF, NVFP4, AWQ, EXL3, ternary …), at their real sizes. The KV cache comes from each model's config, so hybrid-attention models fit far longer contexts. With context "Any", a model needs to fit at least a 32K context, and the Context column shows the longest that fits in the remaining VRAM. With KV cache "Any", each build is tried with a 16-bit, q8 and q4 cache, and the score decides: a quantized cache costs a little quality (q8 0.2%, q4 1%) and buys a longer context.
The Index maps every benchmark a model has (LMArena, GPQA, MMLU-Pro, HLE, SWE-bench, Terminal-Bench, AIME, the Artificial Analysis Index) onto the LMArena scale, then averages them, weighting each by how well it predicts Arena and by how much its source is trusted: independent measurements (LMArena, Artificial Analysis, EvalEval, official benchmark evaluations) count fully, our own harness runs 85%, and self-reported model-card numbers half. A "?" marks a provisional Index, resting on fewer than two independent measurements. Every score and its source is public in quaedra/local-llm-evals. Quantized builds lose some quality: measured when a publisher reports it (build_evals.json), otherwise estimated from bits per weight.
Decode speed is estimated from memory bandwidth and the bytes read per token (active parameters only, for MoE models), on the reference card named above the list. It is calibrated separately for llama.cpp (GGUF builds) and for vLLM, SGLang and ExLlama (other builds) against runs measured on real cards in the Sybil local-ai-registry, which also appear in each model's details.
The score is intelligence^wi × speed^ws × context^wc. Intelligence and speed are relative to the best in the tier. Context runs on a log scale from 8K (0) to 256K (1). The weights come from the two controls, and intelligence gets the rest (at least 20%). Each model shows the build that scores best, so a smaller build that leaves room for a longer context can win. Context only counts with context set to "Any".
Source: build.py and data on the Hub.