From the Newsroom | Updates and Stories from Kamiwaza.ai

Kamiwaza Powers New Signal65 PINNACLE AI Benchmark

Written by Kamiwaza | September 10, 2026

Signal65 PINNACLE — a new independent benchmarking framework co-developed with and powered by Kamiwaza technology — is now generally available for enterprise evaluations. Rather than ranking models or measuring token speeds, PINNACLE measures whether AI systems actually complete enterprise work correctly, how much infrastructure capacity that work requires, and the cost per correct outcome.

Signal65, an independent technology research and testing firm, built PINNACLE to close a gap in how enterprises evaluate AI: traditional benchmarks measure model capability or infrastructure performance separately, leaving buyers without a clear answer on whether their chosen model-and-hardware combination can do the job. The framework evaluates model intelligence, silicon capacity, and solution economics across real enterprise workloads using Kamiwaza's PICARD execution harness and the KAMI and RIKER benchmark work.

According to Luke Norris, CEO and co-founder of Kamiwaza: "These results validate why we've built Kamiwaza around orchestration and choice from the beginning. Enterprises shouldn't have to bet their AI strategy on one model or infrastructure provider."

The initial PINNACLE dataset tested 44 configurations across 30 base models from 12 makers — hosted APIs and open-weight models running on NVIDIA H200, NVIDIA B300, and AMD Instinct MI300X infrastructure. Nine configurations reached at least 95% job completion on clean data; only two did so on messy data. Model rankings shifted by workload, with smaller and open-weight models matching or beating frontier models on certain tasks. And token pricing proved a poor proxy for the cost of successful work, with input tokens accounting for 65–91% of hosted-model task costs.