Z.ai: GLM-5.3-Flash
Chatz-ai/glm-5.3-flashGLM-5.3-Flash is the lightweight, high-speed model on Z.ai's international site (api.z.ai), released on 2026-08-26. It takes native multimodal input including images and video and targets efficient coding and long-horizon agent tasks. Thinking is always on and cannot be disabled, but the reasoning effort is selectable across low, high, and max. It supports tool use, prompt caching, and web search. Context window: 1M tokens, output: 131K. Available via OpenAI-compatible and Anthropic protocols; note that the OpenAI-compatible surface exposes chat completions and responses endpoints through Ofox.
Context Window
1M
Max Output Tokens
131K
Released
2026-08-26
Capabilities
VisionFunction CallingReasoningPrompt CachingWeb SearchVideo Input
Available Providers
VolcengineZ.ai
Supported Protocols
openaianthropic
Providers
Volcengine
provider.type: "volcengine"Input Tokens
$0.15/M
Output Tokens
$0.5/M
Cache Read
$0.03/M
Web Search
$0.01/R
Protocols
openai
/v1/chat/completions/v1/responsesanthropic
Z.ai
provider.type: "zai"Input Tokens
$0.15/M
Output Tokens
$0.5/M
Cache Read
$0.03/M
Web Search
$0.01/R
Protocols
openai
/v1/chat/completionsanthropic
Code Examples
from openai import OpenAIclient = OpenAI(base_url="https://api.ofox.run/v1",api_key="YOUR_OFOX_API_KEY",)response = client.chat.completions.create(model="z-ai/glm-5.3-flash",messages=[{"role": "user", "content": "Hello!"}],)print(response.choices[0].message.content)
Uptime & Status
Related Models
Frequently Asked Questions
Z.ai: GLM-5.3-Flash on Ofox.ai costs $0.15/M per million input tokens and $0.5/M per million output tokens. Pay-as-you-go, no monthly fees.
Z.ai: GLM-5.3-Flash supports a context window of 1M tokens with max output of 131K tokens, allowing you to process large documents and maintain long conversations.
Simply set your base URL to https://api.ofox.run/v1 and use your Ofox API key. The API is OpenAI-compatible โ just change the base URL and API key in your existing code.
Z.ai: GLM-5.3-Flash supports the following capabilities: Vision, Function Calling, Reasoning, Prompt Caching, Web Search, Video Input. Access all features through the Ofox.ai unified API.