GLM 4.7 Fp8
Pricing & Specs
Zhipu AI · $0.45 input / $2 output per 1M tokens · 202.752K context window
No longer offered on OpenMark
GLM 4.7 Fp8 is no longer offered on OpenMark (deprecated or superseded upstream). Its specs are kept here for reference — if you're still running it in production, it's worth benchmarking current alternatives: newer models are often cheaper AND better on the same task.
GLM 4.7 Fp8 API Pricing
| Tokens | Price |
|---|---|
| Input | $0.45 / 1M tokens |
| Output | $2 / 1M tokens |
50% discount available for batch API requests. Batch API pricing is available via Together at 50%; cache pricing is not published.
What that means in practice
GLM 4.7 Fp8 Specs
| Spec | Value |
|---|---|
| Context window | 202.752K tokens |
| Input modalities | text |
| Output modalities | text |
| Latency (measured by OpenMark) | ~2.2s median response |
| Reasoning model | Yes |
| Tool / function calling | Yes |
| JSON mode | Yes |
| Streaming | Yes |
| Prompt caching | No |
| Batch API | Yes |
Notes
Scheduled for deprecation on Together AI: April 02, 2026.
FAQ
How much does GLM 4.7 Fp8 cost?
GLM 4.7 Fp8 costs $0.45 per 1M input tokens and $2 per 1M output tokens.
What is GLM 4.7 Fp8's context window?
GLM 4.7 Fp8 has a 202.752K-token context window.
Can I test GLM 4.7 Fp8 on my own task?
GLM 4.7 Fp8 is no longer offered on OpenMark, but you can benchmark its current alternatives from Zhipu AI and other providers on your own task — free tier available.
Still using GLM 4.7 Fp8?
It's been superseded. Benchmark its current alternatives on your actual
workload — you'll likely find something cheaper AND better. Free tier available.