z-ai

glm-5-turbo

Faster, lower-cost GLM-5 variant tuned for real-time, high-throughput workloads.

z-ai 198K context

Value rank

#13 of 18

intelligence per dollar in our catalogue

Context rank

#3 tied

198K token window

Benchmarks

Independent scores by Artificial Analysis, compared with the strongest models in our catalogue.

Intelligence index

glm-5.3
44.9
kimi-k3
43.8

Long context

kimi-k3
88.7%
glm-5.3
79.7%
glm-5-turbo
71.7%

Best for

Where this model earns its keep.

Prompt-cached workloads

The numbers

Pricing is live from our platform. Prices per 1M tokens, zero data retention on every request.

Input price$1.20
Cache read price$0.30
Output price$4.00
Context window198K tokens
Intelligence / coding index26.6 / -
Agentic: Terminal-Bench v2.1 / tau2- / 99%
Long-context reasoning72%
GPQA / MMLU-Pro85% / -

Or consider

Close alternatives in the catalogue.

Quick start

OpenAI-compatible. Switch in one line.

# pip install openai
client = OpenAI(base_url="https://api.tensorx.ai/v1", api_key="tsx-...")
r = client.chat.completions.create(
    model="z-ai/glm-5-turbo",
    messages=[{"role": "user", "content": "Hello"}],
)

Benchmark data from the Artificial Analysis Intelligence Index, measured independently. Pricing live from the TensorX platform. All inference on EU-sovereign infrastructure with zero data retention.