Search Intelligence Score
55.3
+35.3 lift from search
050100
- Lift from search
- +35.3
- 20.0 without search
- Cost per 1K tasks
- $110
- #5 in Search Efficiency
- Time per task
- 147s
- Search plus inference time to complete one task
Accuracy by benchmark
Accuracy on each eval, with search (solid) and without (light).
- DSQA44.0 vs 7.0
- HLE59.0 vs 42.0
- WISER63.0 vs 11.0
Cost per 1K tasks
How the cost splits between model inference, search calls, and page extraction.
- Inference
- $103
- 94%
- Search
- $3.24
- 3%
- Extract
- $3.79
- 3%
InferenceSearchExtract
Compared with
Other models from the same provider alongside models that score closest.
| Model | Score | Lift | Cost per 1K tasks | Time per task |
|---|---|---|---|---|
| GLM 5.3 | 55.3 | +35.3 | $110 | 147s |
| GLM-5.2GLM-5.2 | 51.7 | +34.7 | $122 | 118s |
| GPT-5.6 LunaGPT-5.6 Luna | 54.7 | +32.3 | $23.0 | 59.4s |
| Kimi K3Kimi K3 | 56.3 | +32.3 | $209 | 198s |
| Gemini 3.7Gemini 3.7 Flash | 54.3 | +25.3 | $61.1 | 79.7s |
Search Intelligence Score averages accuracy on DSQA, HLE, and WISER with equal weight; cost and time are averaged the same way, with cost shown per 1,000 tasks. Latest update September 14, 2026. Full methodology.
# GLM 5.3 · Search Capability Leaderboard
- Search Intelligence Score: 55.3 (#8 in Search Intelligence)
- Without search: 20.0 · Lift from search +35.3
- Cost per 1K tasks: $110 (#5 in Search Efficiency)
- Time per task: 147s
- Lab: Z.ai · Model id: z-ai/glm-5.3-openrouter
## Accuracy by benchmark
| Suite | With search | Without search |
|---|---|---|
| DSQA | 44.0 | 7.0 |
| HLE | 59.0 | 42.0 |
| WISER | 63.0 | 11.0 |
## Compared with
| Model | Score | Lift | Cost per 1K tasks | Time per task |
|---|---|---|---|---|
| GLM 5.3 | 55.3 | +35.3 | $110 | 147s |
| GLM-5.2 | 51.7 | +34.7 | $122 | 118s |
| GPT-5.6 Luna | 54.7 | +32.3 | $23.0 | 59.4s |
| Kimi K3 | 56.3 | +32.3 | $209 | 198s |
| Gemini 3.7 Flash | 54.3 | +25.3 | $61.1 | 79.7s |
Latest update September 14, 2026. Methodology: /leaderboard#methodology