Most quantisation write-ups have a tell: every number points the same way, and the author’s own tool measured all of them. So we did the opposite. We took the 4-bit build of GLM-5.2, the NVFP4 quant, and put it through the model makers’ own official benchmark harnesses, at their own published settings. No home-grown eval, no scoreboard of our own design, no best result out of five attempts.
The short version: at half the bits, the 4-bit build lands effectively even with the larger FP8 build almost everywhere, competitive with top proprietary agents on public leaderboards, with a single honest weakness we are transparent about. Here is the full picture, including the place we come second.
What we actually tested
The credibility of a benchmark is in the method, so here is ours in plain terms.
- We served the publicly available NVFP4 checkpoint on a standard OpenAI-compatible endpoint, the same kind of endpoint you would call.
- We ran each maker’s official harness at their published temperatures, output limits and native function-calling. If a test ships with a judge, we say which judge.
- We scored once. No best-of-N re-rolling. Tasks that ran and failed are counted as failures.
- Where an external issue blocked a task before it ever reached the model, we disclose it rather than quietly dropping it.
That is the whole point. Anyone can pull the same public harnesses and check our working.
The results: near-parity with FP8
On matched, official harnesses the two quants are effectively tied. The gap between 4-bit and FP8 is a rounding error nearly everywhere.
| Test (official harness) | 4-bit (NVFP4) result | Read |
|---|---|---|
| τ²-Bench telecom (strict official protocol) | 99.1% | Matches or edges the FP8 reference. The cleanest apples-to-apples anchor. |
| Terminal-Bench 2.1 (public leaderboard) | 76.4 | Scores above the field of top proprietary agents on the identical public board, a full generation ahead of the previous model version. |
| SWE-bench Pro (public set) | ~75%, early read | In-progress: a subset has run, tracking well above the published first-party figure. Not a final score. |
| BFCL v4 (Berkeley function-calling) | 71.5 | On par with a comparable open model. |
Two honest caveats travel with those numbers, because they are trust signals, not fine print. SWE-bench Pro is a partial, in-progress read, and we will update it when the full set completes. The Terminal-Bench figure is a conservative floor, measured under shared live-serving conditions against fixed per-task time limits, not the ceiling of the quant. Where we compare against named proprietary models, that comparison is true on the public leaderboard and nowhere else, so that is exactly how we state it.

The one place FP8 still wins
Here is the weakness. On long-context exact recall, pulling one precise fact back from far up a long conversation, an order number, a single clause, FP8 still leads. The 4-bit build is measurably behind on that specific skill.
Two things make it a weakness we can live with, and that you should weigh honestly:
- It is a safe failure. When the 4-bit build cannot retrieve the exact detail, it declines rather than inventing one. For a regulated buyer, a model that abstains beats a model that confidently hallucinates a wrong account number.
- Part of the gap may not be the quant at all. Our FP8 comparison point was a third-party FP8 endpoint, so a slice of the difference could be serving conditions rather than the weights. We are not going to claim precision we do not have.
If your workload leans on exact retrieval across very long inputs, this is the trade to know about. For most coding and agentic work, it does not show up.
Why 4-bit matters to you
Bits are not an academic detail. A 4-bit model has a much smaller memory footprint than an FP8 one, and that footprint is what decides how much frontier-class capability you can serve, how fast, per accelerator.
For TensorX that maths is the whole business. Running frontier models at 4-bit, at near-frontier quality, is how a sovereign, EU-hosted platform serves them quickly and affordably without your data ever leaving the EU. Efficiency is not a cost line here. It is what lets sovereign inference compete on speed and price with providers who keep your data somewhere you cannot see.
We verify what we serve
We do not take a vendor accuracy table on faith. We reproduce it on the live endpoint, with the vendor’s own tooling, and we publish the weak spot alongside the wins. That is the version of a benchmark report a serious buyer can actually use, and it is the same principle that runs through everything we do: independent, verifiable, EU-sovereign, no spin.
You want a provider who shows you the whole picture, including the one place it comes second.
See the models we serve and their independent benchmarks on the models page, or talk to us about EU inference.