Completed
- Single-model Gemma 4 results across 2B, 4B, and 12B.
- One shared CB LLM runtime across all three Gemma representations.
- Locked full five-shot MMLU and eight-process serving measurements.
Same model, same GPU, and the same four serving workloads, measured side by side.
Comparing the original BF16 teacher to our CB LLM and a variety of LLM compression methods.
Measured scaling
The CB Converter has one engine that runs Gemma 4 2B, 4B, and 12B. As model size increases, the measured performance advantage of our CB technology grows with it.
CB conversion extends beyond Gemma. See the separately measured Qwen3 4B result below, followed by the growing set of models we have validated through conversion.
Customer-defined proof
Run a scoped comparison on your workload before making an adoption decision.