Open Benchmark · 148 Items · 18 Categories · 3 Tiers · PRD tier

What's the best model for private, on-prem agentic inference?

klbench: open benchmarks for self-hosted LLMs — CPU, GPU, and frontier cloud APIs on one exam, measured on real on-prem foundations (Tanzu GenAI tile).

Current capability suite — refreshed Aug 26

Latest Results