Open Benchmark · 148 Items · 18 Categories · 3 Tiers · PRD tier
What's the best model for private, on-prem agentic inference?
klbench: open benchmarks for self-hosted LLMs — CPU, GPU, and frontier cloud APIs on one exam, measured on real on-prem foundations (Tanzu GenAI tile).
Current capability suite — refreshed Aug 26
GPU
everything above the entry cards — consumer Ampere–Ada through Blackwell / Hopper datacenter silicon, any count; ranked by capability composite
ranked by capability composite · 6 models · 1 foundation
- 1Qwen/Qwen3.8-27B-FP80.96T1
- 2openai/gpt-oss-120b0.92T1
- 3google/gemma-4-31B-it-qat-w4a16-ct0.92T1
ENTRY GPU
Entry GPU (16 GB class)single entry datacenter card — Tesla T4 / L4 class: the 'runs on what you already have' tier; ranked by usability because item budgets bite here
ranked by usability — item budgets bite here · 2 models · 1 foundation
- 1hf.co/unsloth/gpt-oss-20b-GGUF:UD-Q4_K_XL0.76T1
- 2hf.co/google/gemma-4-12B-it-qat-q4_0-gguf:latest0.53T2
CPU
no accelerator — Xeon-class CPU inference; ranked by usability because latency is part of the result
ranked by usability — timeouts count as zero · 2 models · 1 foundation
- 1hf.co/unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF:Q4_K_M0.67T1
- 2hf.co/google/gemma-4-26B-A4B-it-qat-q4_0-gguf:latest0.24T1
Latest Results
- Aug 31hf.co/unsloth/gpt-oss-20b-GGUF:UD-Q4_K_XLcdc/ENTRY GPU · ollama0.76
- Aug 31hf.co/unsloth/Qwen3-Coder-30B-A3B-Instruct-GGUF:Q4_K_Mcdc/CPU · ollama0.67
- Aug 30hf.co/google/gemma-4-26B-A4B-it-qat-q4_0-gguf:latestcdc/CPU · ollama0.24
- Aug 30hf.co/google/gemma-4-12B-it-qat-q4_0-gguf:latestcdc/ENTRY GPU · ollama0.53
- Aug 30Qwen/Qwen3.8-27B-FP8ndc/GPU · vllm0.96