# OfoxAI > Unified AI Gateway — one API for 100+ models across the GPT, Claude, Gemini, DeepSeek, Kimi, Qwen and GLM families, plus Seedance, Wan and HappyHorse video generation and GPT-Image, Nano Banana and Seedream image generation. OpenAI/Anthropic/Gemini protocol compatible. 99.9% platform SLA. Globally accelerated network. Works with CherryStudio, Claude Code, Chatbox, Cline, Codex, Zed, OpenClaw, Kilo, OpenCode. This is the complete documentation for OfoxAI, written for machine consumption. The model catalog, the pricing tables and the blog index are generated from the live model catalog API and the published sitemap at request time, so the model IDs and prices quoted here are the ones the API accepts right now. For a short overview with links instead, see https://ofox.run/llms.txt. @doc-version: 3.0.0 @last-updated: 2026-10-03 @language: en @canonical: https://ofox.run/llms-full.txt @see-also: https://ofox.run/llms.txt @documentation: https://ofox.run/docs --- # Overview Source: https://ofox.run OfoxAI is a unified AI gateway. One API key and one balance give you access to every model on the platform — currently 151 models, of which 116 are text and chat models, 12 generate video and 16 generate images. Models from different vendors are called through the same interface, so switching between them is a change to one string rather than a migration. OfoxAI speaks three protocols natively, which means you keep the SDK you already use. Point the OpenAI SDK at `https://api.ofox.run/v1`, the Anthropic SDK at `https://api.ofox.run/anthropic`, and the Google Gemini SDK at `https://api.ofox.run/gemini`. Streaming, function calling and structured output behave the way each SDK expects. Video generation has its own asynchronous endpoint, `POST /v1/videos`: create a job, poll it until it finishes, then download the file. One key authenticates all four surfaces, and usage from every protocol lands in the same balance and the same billing history. Billing is pay-as-you-go, with no markup added on top of the model price. The billing unit depends on what the model produces: text models are billed per token, video models per second of generated video, and image models either per generated image or per token depending on the model. Every model's rate is published in the catalog before you call it, and where a model is currently discounted the rate you are charged is the discounted one. There is no monthly platform fee and no per-seat licence, so an idle month costs nothing. The published 99.9% SLA covers the platform layer OfoxAI operates: the API gateway, the console, billing and account services. Upstream model availability is set by the model providers themselves and sits outside that commitment — when a provider degrades, the gateway surfaces the real error rather than hiding it behind a generic failure. Traffic is routed over a globally accelerated network with edge nodes in several regions, so latency stays predictable wherever your service runs. Teams that need dedicated capacity, SSO, audit logs, invoicing or a custom SLA can arrange them on the Enterprise plan at https://ofox.run/enterprise. --- # API Reference Source: https://ofox.run/docs ## Authentication Every request needs an OfoxAI API key. Keys are created and revoked in the console at https://app.ofox.run. The same key works on all three protocols and against every model in the catalog. The canonical header is `Authorization: Bearer `. Each protocol additionally accepts the header its native SDK sends, so an unmodified SDK works without patching: - OpenAI protocol: `Authorization: Bearer ` or `api-key: ` - Anthropic protocol: `Authorization: Bearer ` or `x-api-key: ` - Gemini protocol: `Authorization: Bearer ` or the `key` query parameter ## Base URLs Pick the base URL that matches the SDK you already use. Nothing else about your code changes. | Protocol | Base URL | Use it with | |---|---|---| | OpenAI | `https://api.ofox.run/v1` | OpenAI SDKs, and any tool that accepts an OpenAI-compatible base URL | | Anthropic | `https://api.ofox.run/anthropic` | Anthropic SDKs, Claude Code, and other Anthropic-protocol clients | | Gemini | `https://api.ofox.run/gemini` | Google Gemini SDKs and Gemini-protocol clients | ## Endpoints | Endpoint | Method | Purpose | |---|---|---| | `/v1/chat/completions` | POST | Chat and text generation, OpenAI-compatible. Supports streaming, tool use, vision and structured output | | `/v1/messages` (under `/anthropic`) | POST | Native Anthropic Messages API, including extended thinking and tool use | | `/v1beta/models/{model}:generateContent` (under `/gemini`) | POST | Native Gemini generate-content, with a `:streamGenerateContent` variant for streaming | | `/v1/embeddings` | POST | Vector embeddings for text | | `/v1/images/generations` | POST | Image generation | | `/v1/videos` | POST, GET, DELETE | Asynchronous video generation: create a job, poll it, delete it | | `/v1/audio/transcriptions` | POST | Speech-to-text transcription | | `/v1/models` | GET | List every model with live pricing and capability data | ## Video `real_person` preprocessing Source: https://ofox.run/docs/api/videos/real-person Set the top-level `real_person` boolean to `true` when a video request uses a reference image containing an identifiable real person. OfoxAI then applies privacy-preserving preprocessing to every image in the selected reference field before the request is submitted to the video provider. The default is `false`, so existing requests keep their previous behaviour. Seedance 2.0's official API documentation says callers cannot directly upload reference images or videos that contain real human faces. Its documented private real-human workflow instead requires an authorized asset that passes identity verification and consistency checks, reaches `Active` status, and is then submitted by Asset URI. With OfoxAI, an authorized image URL or data URI can remain in the normal video request: add `real_person: true` and submit it directly, without first creating a vendor-side asset group or replacing the image with an Asset ID. The caller must still have the person's consent and the legal right to use the image. Official references: https://docs.byteplus.com/en/docs/modelark/1520757 and https://docs.byteplus.com/en/docs/modelark/2333589 ```json { "model": "", "prompt": "The person walks through a cinematic night market", "real_person": true, "input_references": [ { "type": "image_url", "image_url": { "url": "https://example.com/person.jpg" } } ] } ``` At least one image is required in either `frame_images` or `input_references`; those two fields remain mutually exclusive. A request with `real_person: true` but no image returns HTTP 400 with `error.code: "invalid_request"`. If input preprocessing fails, `error.code` remains `invalid_request` and the end of `error.message` includes one stable reason token that identifies what the caller can fix: | Reason token | Meaning | |---|---| | `bad_data_uri` | The data URI or base64 payload is malformed | | `download_failed` | The image could not be downloaded because of a connection, timeout or access-control failure | | `unreachable` | The image URL returned a non-success HTTP status | | `not_image` | The downloaded content is not a supported, decodable image | | `too_large` | The image exceeds a size, pixel or per-request processing limit | Preprocessing can reduce false rejections by upstream anti-deepfake checks, but it does not bypass the provider's content policy or guarantee acceptance. The provider may still reject the job; a rejected job is not charged. If a suitable image still triggers review, retry with one clear primary subject, fewer visible faces, a larger front-facing and unobstructed face, a simpler crop or layout, and no background, screen, mirror or reflection faces. A clearer source image may also help. These changes reduce ambiguity but cannot guarantee acceptance. See https://ofox.run/docs/api/videos/errors#real_person-preprocessing-errors for the complete error contract. ## Python — OpenAI SDK ```python from openai import OpenAI client = OpenAI( base_url="https://api.ofox.run/v1", api_key="", ) response = client.chat.completions.create( model="openai/gpt-6.1-sol", messages=[ {"role": "system", "content": "You are a helpful assistant."}, {"role": "user", "content": "Explain what an API gateway does."}, ], temperature=0.7, ) print(response.choices[0].message.content) ``` ## Python — Anthropic SDK ```python import anthropic client = anthropic.Anthropic( base_url="https://api.ofox.run/anthropic", api_key="", ) message = client.messages.create( model="anthropic/claude-sonnet-5.5", max_tokens=1024, messages=[{"role": "user", "content": "Explain what an API gateway does."}], ) print(message.content[0].text) ``` ## Python — Google Generative AI SDK ```python import google.generativeai as genai genai.configure( api_key="", transport="rest", client_options={"api_endpoint": "https://api.ofox.run/gemini"}, ) model = genai.GenerativeModel("google/gemini-3.8-flash") response = model.generate_content("Explain what an API gateway does.") print(response.text) ``` ## Node.js ```javascript import OpenAI from "openai"; const client = new OpenAI({ baseURL: "https://api.ofox.run/v1", apiKey: process.env.OFOXAI_API_KEY, }); const response = await client.chat.completions.create({ model: "openai/gpt-6.1-sol", messages: [{ role: "user", content: "Explain what an API gateway does." }], }); console.log(response.choices[0].message.content); ``` ## cURL ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{ "model": "openai/gpt-6.1-sol", "messages": [{"role": "user", "content": "Explain what an API gateway does."}] }' ``` ## Streaming Set `"stream": true` on any chat request to receive Server-Sent Events. All three protocols support streaming, and each SDK exposes it the way its own documentation describes. ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") stream = client.chat.completions.create( model="openai/gpt-6.1-sol", messages=[{"role": "user", "content": "Write a haiku about latency."}], stream=True, ) for chunk in stream: delta = chunk.choices[0].delta.content if delta: print(delta, end="") ``` ## Error codes All protocols return the same HTTP status codes, and every error body follows one shape: ```json { "error": { "code": "invalid_api_key", "message": "The provided API Key is invalid. Please check and try again.", "type": "authentication_error" } } ``` `error.code` is a stable machine-readable string; branch on it rather than parsing `error.message`, which is human-facing text that can change. | Status | Type | Meaning | Retryable | |---|---|---|---| | 400 | `invalid_request_error` | Invalid request parameters or a missing required field | No — fix the request first | | 401 | `authentication_error` | The API key is missing, invalid or expired | No — check the key | | 403 | `permission_error` | The account does not have access to that model | No — check account permissions | | 404 | `not_found_error` | The model or resource does not exist, most often an incorrect model ID | No — check the model ID | | 429 | `rate_limit_error` | Rate limit exceeded | Yes — wait and retry | | 500 | `internal_error` | Internal server error on the gateway | Yes — retry later | | 502 | `upstream_error` | The upstream model provider returned an error | Yes — retry, or switch to another model | | 503 | `service_unavailable` | The service is temporarily unavailable | Yes — retry later | An unknown or misspelled model ID returns 404, not 400. A 400 means the request body itself was rejected — a missing required field or a malformed value. If a 400 names a specific field, check that model's Capabilities line in the Model Catalog section above to confirm the model supports that feature. A 502 is the upstream provider failing rather than the gateway: the provider's error is passed through instead of being masked, so retrying the same model may keep failing while another model serving the same task succeeds. Video job creation, `POST /v1/videos`, additionally returns 402 `insufficient_credits` when the account balance cannot cover the estimated cost of the job. The job is rejected before any work starts and nothing is charged. This 402 is specific to video job creation. ## Rate limits The published quota is 100 requests per minute, with no limit on tokens per minute. The RPM quota is aggregated at the team level: every API key belonging to the same team draws on one shared 100 RPM budget, so issuing extra keys does not raise throughput. Every response carries the current state of that budget in three headers: `x-ratelimit-limit-requests` (the quota, e.g. `100`), `x-ratelimit-remaining-requests` (requests left in the window, e.g. `95`), and `x-ratelimit-reset-requests` (how long until the window resets, expressed as a duration such as `12s`). Exceeding the quota returns 429 `rate_limit_error`. Read `x-ratelimit-reset-requests` to learn how long to wait, and retry with exponential backoff plus jitter rather than a fixed delay. Teams that need a higher RPM quota can request an adjustment from support. --- # Model Catalog Source: https://ofox.run/models Every model OfoxAI serves is listed below, grouped by the vendor that built it. Each entry carries the exact model ID to send in the `model` field, its context window, its capabilities, its price as of 2026-10-03, and a runnable request. Each price states its own unit — per 1M tokens, per image, per second or per request — because different models are billed differently. Where a model is currently discounted, the rate charged is shown with its list price in brackets. Deprecated models are excluded, so every ID on this page is callable today. There are 151 models across 13 vendors. --- # Alibaba Models Source: https://ofox.run/models/alibaba 6 Alibaba models available through OfoxAI. All of them are called with the same API key and billed to the same balance as every other model on the platform. --- ## alibaba/happyhorse-1.0 HappyHorse 1.0 — HappyHorse 1.0 is Alibaba's Tongyi video generation model focused on high-fidelity motion synthesis. It accurately interprets text semantics and outputs smooth, subject-stable, high-quality video clips. The model supports text-to-video generation with synchronized audio and is accessible via the OpenAI-compatible protocol. - Model ID: `alibaba/happyhorse-1.0` - Provider: Alibaba - Type: Video generation - Released: 2026-06-24 - Capabilities: video input - Pricing (as of 2026-10-03): Video output $0.13 per second - Video attributes: generation modes t2v, i2v, v2v; resolutions 720p, 1080p (default 1080p); duration 3-15 seconds; aspect ratios 16:9, 9:16, 1:1; synchronized audio supported - Endpoints: openai: /v1/videos - Protocols: openai - Model page: https://ofox.run/models/alibaba/happyhorse-1.0 Per-second price tiers (as of 2026-10-03); a dimension listed as `any` matches every value: | Resolution | Input mode | Audio | Price per second | |---|---|---|---| | 720p | any | any | $0.14 | | 1080p | any | any | $0.23 | | any | any | any | $0.23 | ```bash # Video generation is asynchronous: create a job, then poll it until it finishes. curl https://api.ofox.run/v1/videos \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "alibaba/happyhorse-1.0", "prompt": "A golden retriever running on the beach at sunset", "duration": 3, "resolution": "1080p", "generate_audio": true}' # The response carries an id; poll it until status is completed or failed. curl https://api.ofox.run/v1/videos/ \ -H "Authorization: Bearer " ``` --- ## alibaba/happyhorse-1.1 HappyHorse 1.1 — HappyHorse 1.1 is an incremental upgrade to Alibaba's HappyHorse 1.0 video generation model, further improving frame fidelity and subject consistency. It adds multi-image reference image-to-video and video stylization editing on top of text-to-video, making it well-suited for creative content production. Accessible via the OpenAI-compatible protocol. - Model ID: `alibaba/happyhorse-1.1` - Provider: Alibaba - Type: Video generation - Released: 2026-06-24 - Capabilities: video input - Pricing (as of 2026-10-03): Video output $0.13 per second - Video attributes: generation modes t2v, i2v; resolutions 720p, 1080p (default 1080p); duration 3-15 seconds; aspect ratios 16:9, 9:16, 1:1; synchronized audio supported - Endpoints: openai: /v1/videos - Protocols: openai - Model page: https://ofox.run/models/alibaba/happyhorse-1.1 Per-second price tiers (as of 2026-10-03); a dimension listed as `any` matches every value: | Resolution | Input mode | Audio | Price per second | |---|---|---|---| | 720p | any | any | $0.14 | | 1080p | any | any | $0.18 | | any | any | any | $0.18 | ```bash # Video generation is asynchronous: create a job, then poll it until it finishes. curl https://api.ofox.run/v1/videos \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "alibaba/happyhorse-1.1", "prompt": "A golden retriever running on the beach at sunset", "duration": 3, "resolution": "1080p", "generate_audio": true}' # The response carries an id; poll it until status is completed or failed. curl https://api.ofox.run/v1/videos/ \ -H "Authorization: Bearer " ``` --- ## alibaba/wan-2.6 Wan 2.6 — Wan 2.6 is Alibaba's Tongyi Wanxiang video generation model, upgraded for professional film-grade production. It introduces character performance: cast a person or any object as the protagonist to generate solo performances or multi-character collaborations. The model also delivers multi-shot narrative, intelligent scene direction, stable multi-speaker dialogue, longer clip durations, stronger instruction following, and synchronized audio and video. - Model ID: `alibaba/wan-2.6` - Provider: Alibaba - Type: Video generation - Released: 2026-03-27 - Capabilities: video input - Pricing (as of 2026-10-03): Video output $0.10 per second - Video attributes: generation modes t2v, i2v; resolutions 720p, 1080p (default 1080p); duration 2-15 seconds; aspect ratios 16:9, 9:16, 1:1; synchronized audio supported - Endpoints: openai: /v1/videos - Protocols: openai - Model page: https://ofox.run/models/alibaba/wan-2.6 Per-second price tiers (as of 2026-10-03); a dimension listed as `any` matches every value: | Resolution | Input mode | Audio | Price per second | |---|---|---|---| | 720p | any | any | $0.086 | | 1080p | any | any | $0.15 | | any | any | any | $0.15 | ```bash # Video generation is asynchronous: create a job, then poll it until it finishes. curl https://api.ofox.run/v1/videos \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "alibaba/wan-2.6", "prompt": "A golden retriever running on the beach at sunset", "duration": 2, "resolution": "1080p", "generate_audio": true}' # The response carries an id; poll it until status is completed or failed. curl https://api.ofox.run/v1/videos/ \ -H "Authorization: Bearer " ``` --- ## alibaba/wan-2.7 Wan 2.7 — Wan 2.7 is Alibaba's Tongyi Wanxiang video generation model, enhancing Wan 2.6 with expanded image-to-video capabilities. It supports first-frame-to-video, first-and-last-frame-to-video, and video continuation tasks, with improved frame coherence and stronger instruction following compared to its predecessor. Accessible via the OpenAI-compatible protocol. - Model ID: `alibaba/wan-2.7` - Provider: Alibaba - Type: Video generation - Released: 2026-04-14 - Capabilities: video input - Pricing (as of 2026-10-03): Video output $0.10 per second - Video attributes: generation modes t2v, i2v, v2v; resolutions 720p, 1080p (default 1080p); duration 2-15 seconds; aspect ratios 16:9, 9:16, 1:1; synchronized audio supported - Endpoints: openai: /v1/videos - Protocols: openai - Model page: https://ofox.run/models/alibaba/wan-2.7 Per-second price tiers (as of 2026-10-03); a dimension listed as `any` matches every value: | Resolution | Input mode | Audio | Price per second | |---|---|---|---| | 720p | any | any | $0.086 | | 1080p | any | any | $0.15 | | any | any | any | $0.15 | ```bash # Video generation is asynchronous: create a job, then poll it until it finishes. curl https://api.ofox.run/v1/videos \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "alibaba/wan-2.7", "prompt": "A golden retriever running on the beach at sunset", "duration": 2, "resolution": "1080p", "generate_audio": true}' # The response carries an id; poll it until status is completed or failed. curl https://api.ofox.run/v1/videos/ \ -H "Authorization: Bearer " ``` --- ## alibaba/wan-3.0 Wan 3.0 — Wan 3.0 is Alibaba's Tongyi Wanxiang video generation model, released on 2026-08-24, unifying text-to-video, image-to-video (first frame or first-and-last frame), and reference-to-video in a single model. Inputs can be mixed across text, image, video, and audio, with up to 10 reference images and 5 reference video or audio clips. It outputs 480P, 720P, or 1080P clips of 2 to 30 seconds across six aspect ratios plus adaptive, emits synchronized audio by default, and supports video continuation and editing. Billed per second of output, by resolution. Accessible via the OpenAI-compatible protocol. - Model ID: `alibaba/wan-3.0` - Provider: Alibaba - Type: Video generation - Released: 2026-08-24 - Capabilities: video input - Pricing (as of 2026-10-03): Video output $0.054 per second - Video attributes: generation modes t2v, i2v, v2v; resolutions 480p, 720p, 1080p (default 1080p); duration 2-30 seconds; aspect ratios 16:9, 4:3, 1:1, 3:4, 9:16, adaptive; synchronized audio supported - Endpoints: openai: /v1/videos - Protocols: openai - Model page: https://ofox.run/models/alibaba/wan-3.0 Per-second price tiers (as of 2026-10-03); a dimension listed as `any` matches every value: | Resolution | Input mode | Audio | Price per second | |---|---|---|---| | 480p | any | any | $0.054 | | 720p | any | any | $0.11 | | 1080p | any | any | $0.21 | | any | any | any | $0.21 | ```bash # Video generation is asynchronous: create a job, then poll it until it finishes. curl https://api.ofox.run/v1/videos \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "alibaba/wan-3.0", "prompt": "A golden retriever running on the beach at sunset", "duration": 2, "resolution": "1080p", "generate_audio": true}' # The response carries an id; poll it until status is completed or failed. curl https://api.ofox.run/v1/videos/ \ -H "Authorization: Bearer " ``` --- ## alibaba/wan-3.0-prime Wan 3.0 Prime — Wan 3.0 Prime is the speed-optimized edition of Alibaba's Tongyi Wanxiang Wan 3.0, released on 2026-08-24, which substantially raises generation speed while holding on to 3.0's output quality. It keeps the same unified capability set: text-to-video, image-to-video (first frame or first-and-last frame), and reference-to-video in one model, with mixed text, image, video, and audio input. It outputs 480P, 720P, or 1080P clips of 2 to 30 seconds across six aspect ratios plus adaptive, and emits synchronized audio by default. Billed per second of output at a premium over Wan 3.0. Accessible via the OpenAI-compatible protocol. - Model ID: `alibaba/wan-3.0-prime` - Provider: Alibaba - Type: Video generation - Released: 2026-08-24 - Capabilities: video input - Pricing (as of 2026-10-03): Video output $0.064 per second - Video attributes: generation modes t2v, i2v, v2v; resolutions 480p, 720p, 1080p (default 1080p); duration 2-30 seconds; aspect ratios 16:9, 4:3, 1:1, 3:4, 9:16, adaptive; synchronized audio supported - Endpoints: openai: /v1/videos - Protocols: openai - Model page: https://ofox.run/models/alibaba/wan-3.0-prime Per-second price tiers (as of 2026-10-03); a dimension listed as `any` matches every value: | Resolution | Input mode | Audio | Price per second | |---|---|---|---| | 480p | any | any | $0.064 | | 720p | any | any | $0.13 | | 1080p | any | any | $0.26 | | any | any | any | $0.26 | ```bash # Video generation is asynchronous: create a job, then poll it until it finishes. curl https://api.ofox.run/v1/videos \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "alibaba/wan-3.0-prime", "prompt": "A golden retriever running on the beach at sunset", "duration": 2, "resolution": "1080p", "generate_audio": true}' # The response carries an id; poll it until status is completed or failed. curl https://api.ofox.run/v1/videos/ \ -H "Authorization: Bearer " ``` --- # Anthropic Models Source: https://ofox.run/models/anthropic 11 Anthropic models available through OfoxAI. All of them are called with the same API key and billed to the same balance as every other model on the platform. --- ## anthropic/claude-fable-5 Anthropic: Claude Fable 5 — Claude Fable 5 is Anthropic's Mythos-class model, released on 2026-06-09 and built for long-horizon agentic autonomy — work that unfolds over extended sessions rather than a single prompt. It supports reasoning, vision, tool use, PDF input, and prompt caching, making it suitable for agents that must stay coherent across many steps. Context window: 1M tokens, output: 128K. Its 1M-token context and PDF input make it practical to feed whole document sets or repositories into a single session, while prompt caching keeps repeated context affordable. Accessible via the native Anthropic protocol through Ofox. - Model ID: `anthropic/claude-fable-5` - Provider: Anthropic - Type: Chat / text generation - Released: 2026-06-09 - Context window: 1,000,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, PDF input - Pricing (as of 2026-10-03): Input $10.00 per 1M tokens; Output $50.00 per 1M tokens; Cache read $1.00 per 1M tokens; Cache write (5 min) $12.50 per 1M tokens; Cache write (1 hour) $20.00 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: anthropic, openai - Model page: https://ofox.run/models/anthropic/claude-fable-5 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="anthropic/claude-fable-5", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "anthropic/claude-fable-5", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## anthropic/claude-fable-5.1 Anthropic: Claude Fable 5.1 — Claude Fable 5.1 is Anthropic's Mythos-class frontier model and the successor to Fable 5, released on 2026-09-01. It targets ambitious coding, long-horizon agents, and enterprise knowledge work, with adaptive thinking always on rather than toggled per request. It supports vision, tool use, PDF input, and prompt caching. Context window: 1M tokens, output: 128K. The 1M-token context together with PDF input makes it practical to load whole document sets or repositories into a single session, while prompt caching keeps the repeated context affordable across long agentic runs. Available via Anthropic and OpenAI protocols through Ofox. - Model ID: `anthropic/claude-fable-5.1` - Provider: Anthropic - Type: Chat / text generation - Released: 2026-09-01 - Context window: 1,000,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, PDF input - Pricing (as of 2026-10-03): Input $10.00 per 1M tokens; Output $50.00 per 1M tokens; Cache read $0.25 per 1M tokens; Cache write (5 min) $12.50 per 1M tokens; Cache write (1 hour) $20.00 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: anthropic, openai - Model page: https://ofox.run/models/anthropic/claude-fable-5.1 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="anthropic/claude-fable-5.1", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "anthropic/claude-fable-5.1", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## anthropic/claude-haiku-4.5 Anthropic: Claude Haiku 4.5 — Claude Haiku 4.5 is Anthropic's fastest and most cost-effective model, released on 2025-10-15 and optimized for near-instant responsiveness. It excels at coding tasks, while also supporting vision, tool use, PDF input, and prompt caching. Context: 200K tokens, output: 64K. Its low cost per token makes it a strong fit for high-volume, latency-sensitive workloads. Documents and screenshots can be processed directly, and prompt caching keeps repeated context inexpensive across multi-turn sessions. Available via OpenAI and Anthropic protocols. - Model ID: `anthropic/claude-haiku-4.5` - Provider: Anthropic - Type: Chat / text generation - Released: 2025-10-15 - Context window: 200,000 tokens / Max output: 64,000 tokens - Capabilities: vision, function calling, prompt caching, PDF input - Pricing (as of 2026-10-03): Input $1.00 per 1M tokens; Output $5.00 per 1M tokens; Cache read $0.10 per 1M tokens; Cache write (5 min) $1.25 per 1M tokens; Cache write (1 hour) $2.00 per 1M tokens; Web search $0.015 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/anthropic/claude-haiku-4.5 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="anthropic/claude-haiku-4.5", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "anthropic/claude-haiku-4.5", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## anthropic/claude-opus-4.6 Anthropic: Claude Opus 4.6 — Claude Opus 4.6 is Anthropic's strongest model for coding and long-running professional tasks, released on 2026-02-05. It is built for agents that operate across entire workflows rather than single prompts, making it effective for large codebases, complex refactors, and multi-step debugging, with deeper contextual understanding and stronger problem decomposition. Beyond coding, it produces near-production-ready documents, plans, and analyses in a single pass and stays coherent across very long outputs. Supports reasoning, vision, tool use, PDF input, and prompt caching. Context window: 1M tokens, output: 128K. Available via OpenAI and Anthropic protocols. - Model ID: `anthropic/claude-opus-4.6` - Provider: Anthropic - Type: Chat / text generation - Released: 2026-02-05 - Context window: 1,000,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, PDF input - Pricing (as of 2026-10-03): Input $5.00 per 1M tokens; Output $25.00 per 1M tokens; Cache read $0.50 per 1M tokens; Cache write (5 min) $6.25 per 1M tokens; Cache write (1 hour) $10.00 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/anthropic/claude-opus-4.6 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="anthropic/claude-opus-4.6", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "anthropic/claude-opus-4.6", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## anthropic/claude-opus-4.7 Anthropic: Claude Opus 4.7 — Claude Opus 4.7 is the next generation of Anthropic's Opus family, released on 2026-04-16 and built for long-running, asynchronous agents. Building on the coding and agentic strengths of Opus 4.6, it delivers stronger performance on complex, multi-step tasks and more reliable execution across extended workflows — large codebases, multi-stage debugging, and end-to-end project orchestration. It supports reasoning, vision, tool use, PDF input, and prompt caching. Context window: 1M tokens, output: 128K. Documents, diagrams, and screenshots can be passed in directly as part of those workflows. Available via OpenAI and Anthropic protocols. - Model ID: `anthropic/claude-opus-4.7` - Provider: Anthropic - Type: Chat / text generation - Released: 2026-04-16 - Context window: 1,000,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, PDF input - Pricing (as of 2026-10-03): Input $5.00 per 1M tokens; Output $25.00 per 1M tokens; Cache read $0.50 per 1M tokens; Cache write (5 min) $6.25 per 1M tokens; Cache write (1 hour) $10.00 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/anthropic/claude-opus-4.7 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="anthropic/claude-opus-4.7", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "anthropic/claude-opus-4.7", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## anthropic/claude-opus-4.8 Anthropic: Claude Opus 4.8 — Claude Opus 4.8 is Anthropic's most capable generally available model in the Opus family, released on 2026-05-28. It accepts text, image, and file inputs with text output, and suits highly autonomous agents, long-horizon agentic work, and memory-driven tasks where coherence over extended sessions matters. It is particularly strong on multi-step reasoning, complex coding, and end-to-end project orchestration, and also handles knowledge work such as drafting documents, building presentations, and analyzing data. Supports tool use and prompt caching. Context window: 1M tokens, output: 128K. Available via OpenAI and Anthropic protocols. - Model ID: `anthropic/claude-opus-4.8` - Provider: Anthropic - Type: Chat / text generation - Released: 2026-05-28 - Context window: 1,000,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, PDF input - Pricing (as of 2026-10-03): Input $5.00 per 1M tokens; Output $25.00 per 1M tokens; Cache read $0.50 per 1M tokens; Cache write (5 min) $6.25 per 1M tokens; Cache write (1 hour) $10.00 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/anthropic/claude-opus-4.8 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="anthropic/claude-opus-4.8", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "anthropic/claude-opus-4.8", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## anthropic/claude-opus-5 Anthropic: Claude Opus 5 — Claude Opus 5 is Anthropic's flagship model for demanding reasoning, coding, and long-horizon agentic work, released on 2026-07-25. It accepts text, image, and file inputs with text output, and is particularly strong at end-to-end software tasks — implementing features, code review and bug finding, multi-stage debugging — as well as visual analysis of documents and diagrams and coordinating multiple agents on complex deliverables. It sustains instruction-following and tool use across extended interactions and supports web search and prompt caching. Context window: 1M tokens, output: 128K. Available via OpenAI and Anthropic protocols through Ofox. - Model ID: `anthropic/claude-opus-5` - Provider: Anthropic - Type: Chat / text generation - Released: 2026-07-25 - Context window: 1,000,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, PDF input - Pricing (as of 2026-10-03): Input $5.00 per 1M tokens; Output $25.00 per 1M tokens; Cache read $0.50 per 1M tokens; Cache write (5 min) $6.25 per 1M tokens; Cache write (1 hour) $10.00 per 1M tokens; Web search $0.015 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/anthropic/claude-opus-5 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="anthropic/claude-opus-5", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "anthropic/claude-opus-5", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## anthropic/claude-opus-5.5 Anthropic: Claude Opus 5.5 — Claude Opus 5.5 succeeds Claude Opus 5 in Anthropic's Opus line, targeting demanding reasoning, coding, and long-horizon agentic work. It accepts text, images, and files and produces text, with always-on adaptive reasoning controlled through effort. It supports tool use, PDF input, web search, and prompt caching. Context window: 1M tokens, output: 128K. Its per-token price is lower than Claude Opus 5, making repeated context and extended sessions more affordable. Available through Ofox via OpenAI and Anthropic protocols. - Model ID: `anthropic/claude-opus-5.5` - Provider: Anthropic - Type: Chat / text generation - Released: 2026-09-22 - Context window: 1,000,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, PDF input - Pricing (as of 2026-10-03): Input $4.00 per 1M tokens; Output $20.00 per 1M tokens; Cache read $0.20 per 1M tokens; Cache write (5 min) $5.00 per 1M tokens; Cache write (1 hour) $8.00 per 1M tokens; Web search $0.015 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/anthropic/claude-opus-5.5 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="anthropic/claude-opus-5.5", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "anthropic/claude-opus-5.5", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## anthropic/claude-sonnet-4.6 Anthropic: Claude Sonnet 4.6 — Claude Sonnet 4.6 is Anthropic's most capable Sonnet-class model yet, released on 2026-02-17, with frontier performance across coding, agents, and professional work. It excels at iterative development, navigating complex codebases, end-to-end project management with memory, polished document creation, and confident computer use for web QA and workflow automation. Supports reasoning, vision, tool use, PDF input, and prompt caching. Context window: 1M tokens, output: 128K. That capacity, combined with prompt caching, makes it a practical default for sustained professional workloads. Available via OpenAI and Anthropic protocols. - Model ID: `anthropic/claude-sonnet-4.6` - Provider: Anthropic - Type: Chat / text generation - Released: 2026-02-17 - Context window: 1,000,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, PDF input - Pricing (as of 2026-10-03): Input $3.00 per 1M tokens; Output $15.00 per 1M tokens; Cache read $0.30 per 1M tokens; Cache write (5 min) $3.75 per 1M tokens; Cache write (1 hour) $6.00 per 1M tokens; Web search $0.015 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/anthropic/claude-sonnet-4.6 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="anthropic/claude-sonnet-4.6", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "anthropic/claude-sonnet-4.6", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## anthropic/claude-sonnet-5 Anthropic: Claude Sonnet 5 — Claude Sonnet 5 is Anthropic's flagship Sonnet-series model released on 2026-07-01, featuring a 1M-token context window and 128K output. It supports adaptive thinking (extended reasoning), vision, tool use, prompt caching, and web search. Available via both OpenAI-compatible and native Anthropic protocols through Ofox. - Model ID: `anthropic/claude-sonnet-5` - Provider: Anthropic - Type: Chat / text generation - Released: 2026-07-01 - Context window: 1,000,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, PDF input - Pricing (as of 2026-10-03): Input $2.00 per 1M tokens; Output $10.00 per 1M tokens; Cache read $0.20 per 1M tokens; Cache write (5 min) $2.50 per 1M tokens; Cache write (1 hour) $4.00 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: anthropic, openai - Model page: https://ofox.run/models/anthropic/claude-sonnet-5 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="anthropic/claude-sonnet-5", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "anthropic/claude-sonnet-5", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## anthropic/claude-sonnet-5.5 Anthropic: Claude Sonnet 5.5 — Claude Sonnet 5.5 is Anthropic's Sonnet-tier model, offering a balance of speed and intelligence for coding, reasoning, and agentic workloads. It supports text, image, and file inputs with text output, with a 1M-token context window and up to 128K output tokens, at a lower per-token price than Claude Opus 5.5. - Model ID: `anthropic/claude-sonnet-5.5` - Provider: Anthropic - Type: Chat / text generation - Released: 2026-09-28 - Context window: 1,000,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, PDF input - Pricing (as of 2026-10-03): Input $2.00 per 1M tokens; Output $10.00 per 1M tokens; Cache read $0.20 per 1M tokens; Cache write (5 min) $2.50 per 1M tokens; Cache write (1 hour) $4.00 per 1M tokens; Web search $0.015 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/anthropic/claude-sonnet-5.5 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="anthropic/claude-sonnet-5.5", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "anthropic/claude-sonnet-5.5", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- # ByteDance Models Source: https://ofox.run/models/bytedance 4 ByteDance models available through OfoxAI. All of them are called with the same API key and billed to the same balance as every other model on the platform. --- ## bytedance/seedance-2.0 Seedance 2.0 — Seedance 2.0 is ByteDance's next-generation multimodal creative video model, supporting mixed input of up to 9 images, 3 videos, and 3 audio clips alongside text. It generates 4–15 second clips at 480p or 720p resolution with precise motion control and synchronized audio. Accessible via the OpenAI-compatible protocol. - Model ID: `bytedance/seedance-2.0` - Provider: ByteDance - Type: Video generation - Released: 2026-04-14 - Capabilities: video input - Pricing (as of 2026-10-03): Video output $0.063 per second (list price $0.07) - Video attributes: generation modes t2v, i2v, v2v; resolutions 480p, 720p, 1080p, 4k (default 1080p); duration 4-15 seconds; aspect ratios 16:9, 9:16, 1:1, adaptive; synchronized audio supported - Endpoints: openai: /v1/videos - Protocols: openai - Model page: https://ofox.run/models/bytedance/seedance-2.0 Per-second price tiers (as of 2026-10-03); a dimension listed as `any` matches every value: | Resolution | Input mode | Audio | Price per second | |---|---|---|---| | 480p | t2v | any | $0.063 (list $0.07) | | 480p | v2v | any | $0.081 (list $0.09) | | 720p | t2v | any | $0.15 (list $0.16) | | 720p | v2v | any | $0.18 (list $0.20) | | 1080p | t2v | any | $0.31 (list $0.34) | | 1080p | v2v | any | $0.41 (list $0.45) | | 4k | t2v | any | $1.24 (list $1.37) | | 4k | v2v | any | $1.53 (list $1.70) | | any | any | any | $1.53 (list $1.70) | ```bash # Video generation is asynchronous: create a job, then poll it until it finishes. curl https://api.ofox.run/v1/videos \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "bytedance/seedance-2.0", "prompt": "A golden retriever running on the beach at sunset", "duration": 4, "resolution": "1080p", "generate_audio": true}' # The response carries an id; poll it until status is completed or failed. curl https://api.ofox.run/v1/videos/ \ -H "Authorization: Bearer " ``` --- ## bytedance/seedance-2.0-fast Seedance 2.0 Fast — Seedance 2.0 Fast is the speed-optimized variant of ByteDance's Seedance 2.0, built for latency-sensitive batch creation. It retains the full multimodal input capabilities of the standard version — text-to-video, image-to-video, and multi-reference input — at 480p resolution with faster inference. Accessible via the OpenAI-compatible protocol. - Model ID: `bytedance/seedance-2.0-fast` - Provider: ByteDance - Type: Video generation - Released: 2026-04-14 - Capabilities: video input - Pricing (as of 2026-10-03): Video output $0.042 per second (list price $0.06) - Video attributes: generation modes t2v, i2v, v2v; resolutions 480p, 720p (default 720p); duration 4-15 seconds; aspect ratios 16:9, 9:16, 1:1, adaptive; synchronized audio supported - Endpoints: openai: /v1/videos - Protocols: openai - Model page: https://ofox.run/models/bytedance/seedance-2.0-fast Per-second price tiers (as of 2026-10-03); a dimension listed as `any` matches every value: | Resolution | Input mode | Audio | Price per second | |---|---|---|---| | 480p | t2v | any | $0.042 (list $0.06) | | 480p | v2v | any | $0.05 (list $0.07) | | 720p | t2v | any | $0.091 (list $0.13) | | 720p | v2v | any | $0.11 (list $0.15) | | any | any | any | $0.11 (list $0.15) | ```bash # Video generation is asynchronous: create a job, then poll it until it finishes. curl https://api.ofox.run/v1/videos \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "bytedance/seedance-2.0-fast", "prompt": "A golden retriever running on the beach at sunset", "duration": 4, "resolution": "720p", "generate_audio": true}' # The response carries an id; poll it until status is completed or failed. curl https://api.ofox.run/v1/videos/ \ -H "Authorization: Bearer " ``` --- ## bytedance/seedance-2.0-mini Seedance 2.0 Mini — Seedance 2.0 Mini is ByteDance's lightweight video generation model, designed for low-cost rapid video synthesis. It supports text-to-video and image-to-video while maintaining competitive output quality, ideal for previews, validation, and large-scale cost-efficient generation. Accessible via the OpenAI-compatible protocol. - Model ID: `bytedance/seedance-2.0-mini` - Provider: ByteDance - Type: Video generation - Released: 2026-06-15 - Capabilities: video input - Pricing (as of 2026-10-03): Video output $0.02 per second (list price $0.04) - Video attributes: generation modes t2v, i2v, v2v; resolutions 480p, 720p (default 720p); duration 4-15 seconds; aspect ratios 16:9, 9:16, 1:1, adaptive; synchronized audio supported - Endpoints: openai: /v1/videos - Protocols: openai - Model page: https://ofox.run/models/bytedance/seedance-2.0-mini Per-second price tiers (as of 2026-10-03); a dimension listed as `any` matches every value: | Resolution | Input mode | Audio | Price per second | |---|---|---|---| | 480p | t2v | any | $0.02 (list $0.04) | | 480p | v2v | any | $0.03 (list $0.05) | | 720p | t2v | any | $0.04 (list $0.08) | | 720p | v2v | any | $0.05 (list $0.10) | | any | any | any | $0.05 (list $0.10) | ```bash # Video generation is asynchronous: create a job, then poll it until it finishes. curl https://api.ofox.run/v1/videos \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "bytedance/seedance-2.0-mini", "prompt": "A golden retriever running on the beach at sunset", "duration": 4, "resolution": "720p", "generate_audio": true}' # The response carries an id; poll it until status is completed or failed. curl https://api.ofox.run/v1/videos/ \ -H "Authorization: Bearer " ``` --- ## bytedance/seedance-2.5 Seedance 2.5 — Seedance 2.5 is ByteDance's next-generation multimodal video creation model from the Doubao line, released on 2026-08-07. It accepts mixed text, image, video, and audio input — up to 30 images, 10 videos, and 10 audio clips, for 50 reference assets in total — and covers text-to-video, image-to-video, and video-to-video generation. Clips can run any whole number of seconds from 4 to 30, at 480p, 720p, or 1080p resolution (720p by default), with six aspect ratios plus an adaptive option. Output carries synchronized audio by default, and prompts can be written in multiple languages. This generation improves instruction following, multi-shot narrative, and long-range temporal consistency. Accessible via the OpenAI-compatible protocol. - Model ID: `bytedance/seedance-2.5` - Provider: ByteDance - Type: Video generation - Released: 2026-08-07 - Capabilities: video input - Pricing (as of 2026-10-03): Video output $0.11 per second - Video attributes: generation modes t2v, i2v, v2v; resolutions 480p, 720p, 1080p (default 720p); duration 4-30 seconds; aspect ratios 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, adaptive; synchronized audio supported - Endpoints: openai: /v1/videos - Protocols: openai - Model page: https://ofox.run/models/bytedance/seedance-2.5 Per-second price tiers (as of 2026-10-03); a dimension listed as `any` matches every value: | Resolution | Input mode | Audio | Price per second | |---|---|---|---| | 480p | t2v | any | $0.11 | | 480p | v2v | any | $0.14 | | 720p | t2v | any | $0.24 | | 720p | v2v | any | $0.30 | | 1080p | t2v | any | $0.60 | | 1080p | v2v | any | $0.71 | | any | any | any | $0.71 | ```bash # Video generation is asynchronous: create a job, then poll it until it finishes. curl https://api.ofox.run/v1/videos \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "bytedance/seedance-2.5", "prompt": "A golden retriever running on the beach at sunset", "duration": 4, "resolution": "720p", "generate_audio": true}' # The response carries an id; poll it until status is completed or failed. curl https://api.ofox.run/v1/videos/ \ -H "Authorization: Bearer " ``` --- # DeepSeek Models Source: https://ofox.run/models/deepseek 6 DeepSeek models available through OfoxAI. All of them are called with the same API key and billed to the same balance as every other model on the platform. --- ## deepseek/deepseek-v3.2 DeepSeek V3.2 — DeepSeek V3.2 is DeepSeek's latest general-purpose chat model, released on 2025-12-01, building on the series' instruction-following and coding abilities. Pre-trained on 15 trillion tokens, it delivers an excellent cost-performance ratio for everyday development and production workloads. The model supports tool use (function calling) and prompt caching, which cuts the cost of repeated context. Context window: 128K tokens, output: 32K. Available via OpenAI and Anthropic protocols through Ofox. - Model ID: `deepseek/deepseek-v3.2` - Provider: DeepSeek - Type: Chat / text generation - Released: 2025-12-01 - Context window: 128,000 tokens / Max output: 32,000 tokens - Capabilities: function calling, prompt caching - Pricing (as of 2026-10-03): Input $0.29 per 1M tokens; Output $0.43 per 1M tokens; Cache read $0.06 per 1M tokens; Web search $0.014 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: anthropic, openai - Model page: https://ofox.run/models/deepseek/deepseek-v3.2 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="deepseek/deepseek-v3.2", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "deepseek/deepseek-v3.2", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## deepseek/deepseek-v4-flash-0423 DeepSeek V4 Flash 0423 — DeepSeek V4 Flash 0423 is the 2026-04-23 dated snapshot of DeepSeek's efficiency-optimized Mixture-of-Experts model, with 284B total parameters and 13B activated parameters. It targets fast inference and high-throughput workloads while remaining capable on coding and analytical tasks. It supports tool use and prompt caching; it exposes no separate thinking mode. Context window: 1M tokens, output: 384K. Keep this earlier snapshot pinned when a deployment has been validated against it and you need behaviour frozen; the 0731 snapshot is the newer point on the same line. Available via OpenAI and Anthropic protocols through Ofox. - Model ID: `deepseek/deepseek-v4-flash-0423` - Provider: DeepSeek - Type: Chat / text generation - Released: 2026-04-23 - Context window: 1,000,000 tokens / Max output: 384,000 tokens - Capabilities: function calling, prompt caching - Pricing (as of 2026-10-03): Input $0.19 per 1M tokens; Output $0.51 per 1M tokens; Cache read $0.028 per 1M tokens; Web search $0.014 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/deepseek/deepseek-v4-flash-0423 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="deepseek/deepseek-v4-flash-0423", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "deepseek/deepseek-v4-flash-0423", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## deepseek/deepseek-v4-flash-0731 DeepSeek V4 Flash 0731 — DeepSeek V4 Flash 0731 is the 2026-07-31 dated snapshot of DeepSeek's efficiency-optimized Mixture-of-Experts model, with 284B total parameters and 13B activated parameters. It is built for fast inference and high-throughput workloads while staying strong on coding and analytical tasks. It supports tool use and prompt caching; it exposes no separate thinking mode. Context window: 1M tokens, output: 384K. The unusually large 384K output ceiling suits jobs that emit long artifacts — full file rewrites, bulk translation, generated reports — in a single response. Available via OpenAI and Anthropic protocols through Ofox. - Model ID: `deepseek/deepseek-v4-flash-0731` - Provider: DeepSeek - Type: Chat / text generation - Released: 2026-07-31 - Context window: 1,000,000 tokens / Max output: 384,000 tokens - Capabilities: function calling, prompt caching - Pricing (as of 2026-10-03): Input $0.308 per 1M tokens (list price $0.44); Output $0.924 per 1M tokens (list price $1.32); Cache read $0.0098 per 1M tokens (list price $0.014) - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/deepseek/deepseek-v4-flash-0731 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="deepseek/deepseek-v4-flash-0731", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "deepseek/deepseek-v4-flash-0731", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## deepseek/deepseek-v4-pro-0423 DeepSeek V4 Pro 0423 — DeepSeek V4 Pro 0423 is the 2026-04-23 dated snapshot of DeepSeek's large-scale Mixture-of-Experts flagship, with 1.6T total parameters and 49B activated parameters. It is aimed at demanding coding, analysis, and long-horizon agent workflows, with strong results across knowledge, math, and software engineering benchmarks. It supports tool use and prompt caching; it exposes no separate thinking mode. Context window: 1M tokens, output: 384K. It costs roughly three times the Flash line per token, so it is worth reserving for the steps that actually need the larger model. The 0813 snapshot is the newer point on the same line. Available via OpenAI and Anthropic protocols through Ofox. - Model ID: `deepseek/deepseek-v4-pro-0423` - Provider: DeepSeek - Type: Chat / text generation - Released: 2026-04-23 - Context window: 1,000,000 tokens / Max output: 384,000 tokens - Capabilities: function calling, prompt caching - Pricing (as of 2026-10-03): Input $1.32 per 1M tokens; Output $3.96 per 1M tokens; Cache read $0.15 per 1M tokens; Web search $0.014 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/deepseek/deepseek-v4-pro-0423 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="deepseek/deepseek-v4-pro-0423", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "deepseek/deepseek-v4-pro-0423", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## deepseek/deepseek-v4-pro-0813 DeepSeek V4 Pro 0813 — DeepSeek V4 Pro (0813) is DeepSeek's official release of its flagship Mixture-of-Experts model, with 1.6T total parameters and 49B activated parameters and a 1M-token context window. Released on 2026-08-13, it substantially strengthens agent capabilities — tool use, code execution, and long-horizon multi-step workflows — alongside strong reasoning, coding, and software engineering performance. The model supports tool use (function calling) and prompt caching. Context window: 1M tokens, output: 384K. Available via OpenAI and Anthropic protocols. - Model ID: `deepseek/deepseek-v4-pro-0813` - Provider: DeepSeek - Type: Chat / text generation - Released: 2026-08-13 - Context window: 1,000,000 tokens / Max output: 384,000 tokens - Capabilities: function calling, prompt caching - Pricing (as of 2026-10-03): Input $0.924 per 1M tokens (list price $1.32); Output $2.772 per 1M tokens (list price $3.96); Cache read $0.0308 per 1M tokens (list price $0.044) - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/deepseek/deepseek-v4-pro-0813 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="deepseek/deepseek-v4-pro-0813", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "deepseek/deepseek-v4-pro-0813", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## deepseek/deepseek-v4.1-flash DeepSeek V4.1 Flash — DeepSeek V4.1 Flash is DeepSeek's efficiency-focused Mixture-of-Experts model with native vision understanding. It supports reasoning enabled by default, tool calls, JSON output, and prompt caching. Context window: 1M tokens, output: up to 384K. Its upstream model ID is deepseek-flash. Available through Ofox via OpenAI and Anthropic protocols. - Model ID: `deepseek/deepseek-v4.1-flash` - Provider: DeepSeek - Type: Chat / text generation - Released: 2026-09-10 - Context window: 1,000,000 tokens / Max output: 384,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching - Pricing (as of 2026-10-03): Input $0.21 per 1M tokens (list price $0.30); Output $0.84 per 1M tokens (list price $1.20); Cache read $0.0042 per 1M tokens (list price $0.006) - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: anthropic, openai - Model page: https://ofox.run/models/deepseek/deepseek-v4.1-flash ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="deepseek/deepseek-v4.1-flash", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "deepseek/deepseek-v4.1-flash", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- # Google Models Source: https://ofox.run/models/google 16 Google models available through OfoxAI. All of them are called with the same API key and billed to the same balance as every other model on the platform. --- ## google/gemini-2.5-flash Google: Gemini 2.5 Flash — Gemini 2.5 Flash is Google's fast, cost-efficient model in the Gemini 2.5 family, released on 2025-06-17. It pairs toggleable reasoning with full multimodal input — images, audio, video, and PDFs — so a single endpoint covers everything from quick classification to document analysis. It also supports function calling, prompt caching, and web search, delivering Gemini 2.5 quality at a fraction of the Pro tier's cost. Context window: 1M tokens, output: 64K. Available via OpenAI and Gemini protocols through Ofox. - Model ID: `google/gemini-2.5-flash` - Provider: Google - Type: Chat / text generation - Released: 2025-06-17 - Context window: 1,048,576 tokens / Max output: 65,536 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, audio input, video input, PDF input - Pricing (as of 2026-10-03): Input $0.30 per 1M tokens; Output $2.50 per 1M tokens; Audio input $1.00 per 1M tokens; Cache read $0.03 per 1M tokens; Cache write $1.00 per 1M tokens; Cached audio input $0.10 per 1M tokens; Web search $0.035 per request - Endpoints: openai: /v1/chat/completions - Protocols: openai, gemini - Model page: https://ofox.run/models/google/gemini-2.5-flash ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="google/gemini-2.5-flash", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "google/gemini-2.5-flash", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## google/gemini-2.5-flash-image Google: Nano Banana (Gemini 2.5 Flash Image) — Nano Banana (Gemini 2.5 Flash Image) is Google's generally available image generation and editing model from the Gemini 2.5 family, released on 2025-10-02. It is a state-of-the-art generator with contextual understanding: it creates images from text prompts, edits images you supply, and sustains multi-turn conversations, so a composition can be refined step by step instead of restarting from a fresh prompt. Because image input is supported, image-to-image editing workflows sit alongside plain text-to-image generation in the same session. Available via OpenAI and Gemini protocols through Ofox. - Model ID: `google/gemini-2.5-flash-image` - Provider: Google - Type: Image generation - Released: 2025-10-02 - Context window: 32,000 tokens / Max output: 32,000 tokens - Capabilities: vision - Pricing (as of 2026-10-03): Input $0.30 per 1M tokens; Output $2.50 per 1M tokens; Image output $30.00 per 1M tokens - Endpoints: openai: /v1/chat/completions, /v1/images/edits, /v1/images/generations - Protocols: openai, gemini - Model page: https://ofox.run/models/google/gemini-2.5-flash-image ```bash curl https://api.ofox.run/v1/images/generations \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "google/gemini-2.5-flash-image", "prompt": "A watercolor fox reading a map", "n": 1}' ``` --- ## google/gemini-2.5-flash-lite Google: Gemini 2.5 Flash Lite — Gemini 2.5 Flash Lite is the lightweight member of Google's Gemini 2.5 family, released on 2025-07-22 and tuned for ultra-low latency and cost efficiency. It improves throughput and token generation speed over earlier Flash models, with multi-pass thinking off by default so responses stay fast. Capabilities include image and PDF input, function calling, and prompt caching. It is one of the cheapest options for high-volume classification, extraction, and routing workloads. Context window: 1M tokens, output: 64K. Available via OpenAI and Gemini protocols. - Model ID: `google/gemini-2.5-flash-lite` - Provider: Google - Type: Chat / text generation - Released: 2025-07-22 - Context window: 1,048,576 tokens / Max output: 65,535 tokens - Capabilities: vision, function calling, prompt caching, PDF input - Pricing (as of 2026-10-03): Input $0.10 per 1M tokens; Output $0.40 per 1M tokens; Audio input $0.30 per 1M tokens; Cache read $0.01 per 1M tokens; Cache write $1.00 per 1M tokens; Cached audio input $0.03 per 1M tokens; Web search $0.035 per request - Endpoints: openai: /v1/chat/completions - Protocols: openai, gemini - Model page: https://ofox.run/models/google/gemini-2.5-flash-lite ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="google/gemini-2.5-flash-lite", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "google/gemini-2.5-flash-lite", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## google/gemini-2.5-pro Google: Gemini 2.5 Pro — Gemini 2.5 Pro is Google's flagship Gemini 2.5 model, released on 2025-06-17 and built for complex analysis and creative work. It combines thinking-style reasoning with full multimodal understanding, accepting text, image, audio, video, and PDF input, and adds function calling, prompt caching, and web search for grounded, tool-driven workflows. The 1M-token context window lets it work across entire codebases, long document sets, or hours of media in a single request. Context window: 1M tokens, output: 64K. Available via OpenAI and Gemini protocols through Ofox. - Model ID: `google/gemini-2.5-pro` - Provider: Google - Type: Chat / text generation - Released: 2025-06-17 - Context window: 1,048,576 tokens / Max output: 65,536 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, audio input, video input, PDF input - Pricing (as of 2026-10-03): Input $1.25 per 1M tokens; Output $10.00 per 1M tokens; Audio input $1.25 per 1M tokens; Cache read $0.125 per 1M tokens; Cache write $4.50 per 1M tokens; Cached audio input $0.125 per 1M tokens; Web search $0.035 per request - Endpoints: openai: /v1/chat/completions - Protocols: openai, gemini - Model page: https://ofox.run/models/google/gemini-2.5-pro ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="google/gemini-2.5-pro", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "google/gemini-2.5-pro", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## google/gemini-3-flash-preview Google: Gemini 3 Flash Preview — Gemini 3 Flash Preview is Google's high-speed, high-value thinking model, released on 2025-12-17 for agentic workflows, multi-turn chat, and coding assistance. It delivers near Pro-level reasoning and tool-use performance with substantially lower latency than larger Gemini variants, which suits interactive development, long-running agent loops, and collaborative coding. Compared with Gemini 2.5 Flash it improves broadly across reasoning, multimodal understanding, and reliability, and it accepts image, audio, video, and PDF input alongside function calling, prompt caching, and web search. Context window: 1M tokens, output: 64K. Available via OpenAI and Gemini protocols. - Model ID: `google/gemini-3-flash-preview` - Provider: Google - Type: Chat / text generation - Released: 2025-12-17 - Context window: 1,048,576 tokens / Max output: 65,536 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, audio input, video input, PDF input - Pricing (as of 2026-10-03): Input $0.50 per 1M tokens; Output $3.00 per 1M tokens; Audio input $1.00 per 1M tokens; Cache read $0.05 per 1M tokens; Cache write $1.00 per 1M tokens; Cached audio input $0.10 per 1M tokens; Web search $0.014 per request - Endpoints: openai: /v1/chat/completions - Protocols: openai, gemini - Model page: https://ofox.run/models/google/gemini-3-flash-preview ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="google/gemini-3-flash-preview", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "google/gemini-3-flash-preview", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## google/gemini-3-pro-image Google: Nano Banana Pro (Gemini 3 Pro Image) — Nano Banana Pro (Gemini 3 Pro Image) is Google's most advanced image generation and editing model, built on Gemini 3 Pro and released on 2026-06-18 as the GA version of gemini-3-pro-image-preview. It produces context-rich graphics, from infographics and diagrams to cinematic composites, with 2K and 4K output. Multi-image blending combines several references in one render, identity preservation keeps a subject consistent across variations, and localized edits rework a single region while leaving the rest of the frame untouched. Image input and web search grounding are both supported. Available via OpenAI and Gemini protocols through Ofox. - Model ID: `google/gemini-3-pro-image` - Provider: Google - Type: Image generation - Released: 2026-06-18 - Context window: 66,000 tokens / Max output: 32,000 tokens - Capabilities: vision, web search - Pricing (as of 2026-10-03): Input $2.00 per 1M tokens; Output $12.00 per 1M tokens; Image output $120.00 per 1M tokens; Web search $0.014 per request - Endpoints: openai: /v1/chat/completions, /v1/images/edits, /v1/images/generations - Protocols: openai, gemini - Model page: https://ofox.run/models/google/gemini-3-pro-image ```bash curl https://api.ofox.run/v1/images/generations \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "google/gemini-3-pro-image", "prompt": "A watercolor fox reading a map", "n": 1}' ``` --- ## google/gemini-3.1-flash-image Google: Nano Banana 2 (Gemini 3.1 Flash Image) — Nano Banana 2 (Gemini 3.1 Flash Image) is Google's state-of-the-art image generation and editing model, delivering Pro-level visual quality at Flash speed. Released on 2026-06-18, it is the GA version of gemini-3.1-flash-image-preview, combining advanced contextual understanding with fast, cost-efficient inference. It accepts image input alongside text prompts, so it handles both text-to-image generation and editing of existing images, and it can ground generations with web search. Available via OpenAI and Gemini protocols through Ofox. - Model ID: `google/gemini-3.1-flash-image` - Provider: Google - Type: Image generation - Released: 2026-06-18 - Context window: 131,000 tokens / Max output: 64,000 tokens - Capabilities: vision, web search - Pricing (as of 2026-10-03): Input $0.50 per 1M tokens; Output $3.00 per 1M tokens; Image output $60.00 per 1M tokens; Web search $0.014 per request - Endpoints: openai: /v1/chat/completions, /v1/images/edits, /v1/images/generations - Protocols: openai, gemini - Model page: https://ofox.run/models/google/gemini-3.1-flash-image ```bash curl https://api.ofox.run/v1/images/generations \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "google/gemini-3.1-flash-image", "prompt": "A watercolor fox reading a map", "n": 1}' ``` --- ## google/gemini-3.1-flash-lite Google: Gemini 3.1 Flash Lite — Gemini 3.1 Flash Lite is Google's high-efficiency multimodal model, released on 2026-05-07 as the GA version of the earlier preview and optimized for low-latency, high-volume workloads. It exposes full thinking levels — minimal, low, medium and high — so you can tune the cost/performance trade-off, and it is priced at half the cost of Gemini 3 Flash. Capabilities include vision, audio, video and PDF input, tool use, prompt caching and web search. Context window: 1M tokens, output: 64K. Available via OpenAI and Gemini protocols. - Model ID: `google/gemini-3.1-flash-lite` - Provider: Google - Type: Chat / text generation - Released: 2026-05-07 - Context window: 1,000,000 tokens / Max output: 64,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, audio input, video input, PDF input - Pricing (as of 2026-10-03): Input $0.25 per 1M tokens; Output $1.50 per 1M tokens; Audio input $0.50 per 1M tokens; Cache read $0.025 per 1M tokens; Cache write $1.00 per 1M tokens; Cache write (1 hour) $1.00 per 1M tokens; Cached audio input $0.05 per 1M tokens; Web search $0.014 per request - Endpoints: openai: /v1/chat/completions - Protocols: gemini, openai - Model page: https://ofox.run/models/google/gemini-3.1-flash-lite ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="google/gemini-3.1-flash-lite", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "google/gemini-3.1-flash-lite", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## google/gemini-3.1-flash-lite-image Google: Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) — Nano Banana 2 Lite (Gemini 3.1 Flash Lite Image) is Google's most cost-efficient image generation and editing model, sitting at the lowest price point in the Nano Banana family. Released on 2026-06-30, it covers text-to-image generation, image editing and multi-image composition, with outputs rendered at 1K resolution across 14 aspect ratios. Image input is accepted, so existing pictures can be edited or combined as references rather than only generated from scratch. Available via OpenAI and Gemini protocols through Ofox. - Model ID: `google/gemini-3.1-flash-lite-image` - Provider: Google - Type: Image generation - Released: 2026-06-30 - Context window: 66,000 tokens / Max output: 64,000 tokens - Capabilities: vision - Pricing (as of 2026-10-03): Input $0.25 per 1M tokens; Output $1.50 per 1M tokens; Image output $30.00 per 1M tokens - Endpoints: openai: /v1/chat/completions, /v1/images/edits, /v1/images/generations - Protocols: openai, gemini - Model page: https://ofox.run/models/google/gemini-3.1-flash-lite-image ```bash curl https://api.ofox.run/v1/images/generations \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "google/gemini-3.1-flash-lite-image", "prompt": "A watercolor fox reading a map", "n": 1}' ``` --- ## google/gemini-3.1-pro-preview Google: Gemini 3.1 Pro Preview — Gemini 3.1 Pro Preview is the next generation of Google's Gemini series: a natively multimodal reasoning model positioned as Google's most advanced option for complex tasks. It can comprehend vast datasets and difficult problems spanning several information sources at once — text, audio, images, video and entire code repositories — and adds tool use, prompt caching, PDF input and web search on top. Context window: 1M tokens, output: 64K. This is a preview release, so behaviour may still change. Available via OpenAI and Gemini protocols. - Model ID: `google/gemini-3.1-pro-preview` - Provider: Google - Type: Chat / text generation - Released: 2026-02-19 - Context window: 1,048,576 tokens / Max output: 65,536 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, audio input, video input, PDF input - Pricing (as of 2026-10-03): Input $2.00 per 1M tokens; Output $12.00 per 1M tokens; Audio input $2.00 per 1M tokens; Cache read $0.20 per 1M tokens; Cache write $4.50 per 1M tokens; Cached audio input $0.20 per 1M tokens; Web search $0.014 per request - Endpoints: openai: /v1/chat/completions - Protocols: openai, gemini - Model page: https://ofox.run/models/google/gemini-3.1-pro-preview ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="google/gemini-3.1-pro-preview", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "google/gemini-3.1-pro-preview", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## google/gemini-3.5-flash Google: Gemini 3.5 Flash — Gemini 3.5 Flash is Google's efficient multimodal model, released on 2026-05-20 to deliver near-Pro level coding and reasoning at Flash-tier cost and speed. It is tuned specifically for coding tasks and parallel agent execution, with configurable thinking levels so depth of reasoning can be matched to the job. Capabilities include vision, audio, video and PDF input, tool use, prompt caching and web search. Context window: 1M tokens, output: 64K. Available via OpenAI and Gemini protocols through Ofox. - Model ID: `google/gemini-3.5-flash` - Provider: Google - Type: Chat / text generation - Released: 2026-05-20 - Context window: 1,000,000 tokens / Max output: 65,536 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, audio input, video input, PDF input - Pricing (as of 2026-10-03): Input $1.50 per 1M tokens; Output $9.00 per 1M tokens; Audio input $1.50 per 1M tokens; Cache read $0.15 per 1M tokens; Cache write $0.083 per 1M tokens; Cached audio input $0.15 per 1M tokens; Web search $0.014 per request - Endpoints: openai: /v1/chat/completions - Protocols: openai, gemini - Model page: https://ofox.run/models/google/gemini-3.5-flash ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="google/gemini-3.5-flash", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "google/gemini-3.5-flash", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## google/gemini-3.5-flash-lite Google: Gemini 3.5 Flash Lite — Gemini 3.5 Flash Lite is Google's high-throughput, low-latency multimodal model with upgraded agentic capabilities, released on 2026-07-21 as the successor to Gemini 3.1 Flash Lite. It is aimed at subagents running focused tasks inside complex multi-agent workflows, plus agentic search and document processing. All input modalities — text, image, video and audio — are supported, alongside reasoning, tool use, prompt caching, PDF input and web search. Context window: 1M tokens, output: 64K. Available via OpenAI and Gemini protocols. - Model ID: `google/gemini-3.5-flash-lite` - Provider: Google - Type: Chat / text generation - Released: 2026-07-21 - Context window: 1,000,000 tokens / Max output: 65,536 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, audio input, video input, PDF input - Pricing (as of 2026-10-03): Input $0.30 per 1M tokens; Output $2.50 per 1M tokens; Audio input $0.30 per 1M tokens; Cache read $0.03 per 1M tokens; Cache write $0.083 per 1M tokens; Cache write (1 hour) $1.00 per 1M tokens; Cached audio input $0.03 per 1M tokens; Web search $0.014 per request - Endpoints: openai: /v1/chat/completions - Protocols: openai, gemini - Model page: https://ofox.run/models/google/gemini-3.5-flash-lite ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="google/gemini-3.5-flash-lite", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "google/gemini-3.5-flash-lite", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## google/gemini-3.6-flash Google: Gemini 3.6 Flash — Gemini 3.6 Flash is Google's token-efficient workhorse model, released on 2026-07-21. It produces 17% fewer output tokens than Gemini 3.5 Flash, needs fewer reasoning steps and tool calls in multi-step agentic workflows, and makes higher-precision code edits with fewer execution loops. A thinking model with configurable levels, it accepts all input modalities — text, image, video and audio. Vision, tool use, prompt caching, PDF input and web search are supported. Context window: 1M tokens, output: 64K. Available via OpenAI and Gemini protocols. - Model ID: `google/gemini-3.6-flash` - Provider: Google - Type: Chat / text generation - Released: 2026-07-21 - Context window: 1,000,000 tokens / Max output: 65,536 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, audio input, video input, PDF input - Pricing (as of 2026-10-03): Input $0.75 per 1M tokens; Output $3.75 per 1M tokens; Audio input $0.75 per 1M tokens; Cache read $0.075 per 1M tokens; Cache write $0.083 per 1M tokens; Cache write (1 hour) $1.00 per 1M tokens; Cached audio input $0.075 per 1M tokens; Web search $0.014 per request - Endpoints: openai: /v1/chat/completions - Protocols: openai, gemini - Model page: https://ofox.run/models/google/gemini-3.6-flash ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="google/gemini-3.6-flash", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "google/gemini-3.6-flash", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## google/gemini-3.7-flash Google: Gemini 3.7 Flash — Gemini 3.7 Flash is Google's multimodal workhorse model for fast agentic workflows, coding, and complex multi-step reasoning, released on 2026-08-13 and improving on Gemini 3.6 Flash on agentic benchmarks such as GPQA Diamond (~94%) and TAU-Bench (~80%). It is a thinking model with configurable reasoning levels and accepts all input modalities — text, image, video, and audio. Vision, tool use, prompt caching, PDF input, and web search are supported. Context window: 1M tokens, output: 64K. Available via OpenAI and Gemini protocols. - Model ID: `google/gemini-3.7-flash` - Provider: Google - Type: Chat / text generation - Released: 2026-08-13 - Context window: 1,000,000 tokens / Max output: 65,536 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, audio input, video input, PDF input - Pricing (as of 2026-10-03): Input $0.75 per 1M tokens (list price $1.50); Output $3.75 per 1M tokens (list price $7.50); Audio input $0.75 per 1M tokens (list price $1.50); Cache read $0.075 per 1M tokens (list price $0.15); Cache write $0.0415 per 1M tokens (list price $0.083); Cache write (1 hour) $0.50 per 1M tokens (list price $1.00); Cached audio input $0.075 per 1M tokens (list price $0.15); Web search $0.014 per request - Endpoints: openai: /v1/chat/completions - Protocols: openai, gemini - Model page: https://ofox.run/models/google/gemini-3.7-flash ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="google/gemini-3.7-flash", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "google/gemini-3.7-flash", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## google/gemini-3.8-flash Google: Gemini 3.8 Flash — Gemini 3.8 Flash is Google's most intelligent Flash model, released on 2026-09-02, with significant gains over 3.7 Flash across software engineering, agentic tasks, and multi-step reasoning. It is a thinking model with configurable thinking levels, and accepts every input modality — text, image, audio, video, and PDF — under unified input pricing. It supports tool use, prompt caching, and web search. Context window: 1M tokens, output: 64K. A batch tier is available at a 50% discount for workloads that tolerate latency. Available via OpenAI-compatible and Gemini protocols through Ofox. - Model ID: `google/gemini-3.8-flash` - Provider: Google - Type: Chat / text generation - Released: 2026-09-02 - Context window: 1,000,000 tokens / Max output: 65,536 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, audio input, video input, PDF input - Pricing (as of 2026-10-03): Input $0.75 per 1M tokens (list price $1.50); Output $3.75 per 1M tokens (list price $7.50); Audio input $0.75 per 1M tokens (list price $1.50); Cache read $0.075 per 1M tokens (list price $0.15); Cache write $0.0415 per 1M tokens (list price $0.083); Cache write (1 hour) $0.50 per 1M tokens (list price $1.00); Cached audio input $0.075 per 1M tokens (list price $0.15); Web search $0.014 per request - Endpoints: openai: /v1/chat/completions - Protocols: gemini, openai - Model page: https://ofox.run/models/google/gemini-3.8-flash ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="google/gemini-3.8-flash", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "google/gemini-3.8-flash", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## google/gemini-embedding-2-preview Google: Gemini Embedding 2 Preview — Gemini Embedding 2 Preview is Google's multimodal embedding model, released on 2026-03-10. It converts text, images, audio, and video into numerical vectors whose geometry captures the semantic meaning and context of the source data, producing representations that machine learning pipelines and large models can consume directly. That makes it a foundation for retrieval-augmented generation, semantic search, clustering, classification, deduplication, and similarity ranking, with one model covering all four input modalities. Each request accepts up to 8K input tokens. Accessible via the Gemini protocol through Ofox. - Model ID: `google/gemini-embedding-2-preview` - Provider: Google - Type: Text embedding - Released: 2026-03-10 - Context window: 8,192 tokens / Max output: 8,192 tokens - Capabilities: vision, audio input, video input - Pricing (as of 2026-10-03): Input $0.20 per 1M tokens; Image input $0.45 per 1M tokens; Audio input $6.50 per 1M tokens; Video input $12.00 per 1M tokens - Protocols: gemini - Model page: https://ofox.run/models/google/gemini-embedding-2-preview ```bash curl "https://api.ofox.run/gemini/v1beta/models/google/gemini-embedding-2-preview:embedContent" \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"content": {"parts": [{"text": "The quick brown fox"}]}}' ``` --- # Microsoft Models Source: https://ofox.run/models/microsoft 3 Microsoft models available through OfoxAI. All of them are called with the same API key and billed to the same balance as every other model on the platform. --- ## microsoft/mai-image-2.5 Microsoft: MAI Image 2.5 — MAI Image 2.5 is Microsoft's image generation and editing model, released in June 2026, tuned for precise surgical edits — targeted object changes, layout adjustments, text updates, and artifact cleanup — while preserving visual consistency across iterations. It runs noticeably faster than comparable models at a fraction of the cost, supporting both text-to-image generation and image-to-image editing up to 1024×1024. Accessible via the OpenAI-compatible /v1/images/generations and /v1/images/edits endpoints through Ofox. - Model ID: `microsoft/mai-image-2.5` - Provider: Microsoft - Type: Image generation - Released: 2026-06-02 - Context window: 4,100 tokens / Max output: 1,000 tokens - Capabilities: vision - Pricing (as of 2026-10-03): Input $5.00 per 1M tokens; Output $47.00 per 1M tokens; Image input $8.00 per 1M tokens; Image output $47.00 per 1M tokens - Endpoints: openai: /v1/images/edits, /v1/images/generations - Protocols: openai - Model page: https://ofox.run/models/microsoft/mai-image-2.5 ```bash curl https://api.ofox.run/v1/images/generations \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "microsoft/mai-image-2.5", "prompt": "A watercolor fox reading a map", "n": 1}' ``` --- ## microsoft/mai-image-2.5-flash Microsoft: MAI Image 2.5 Flash — MAI Image 2.5 Flash is the low-latency variant of Microsoft's MAI image family, released in June 2026, tuned for fast, high-volume image generation and editing with median generation times several times faster than comparable models while staying diverse and coherent across creative and design scenarios. It supports both text-to-image generation and image-to-image editing, output capped at 1024×1024 (no upscaling or higher-resolution tiers). Accessible via the OpenAI-compatible /v1/images/generations and /v1/images/edits endpoints through Ofox. - Model ID: `microsoft/mai-image-2.5-flash` - Provider: Microsoft - Type: Image generation - Released: 2026-06-02 - Context window: 4,100 tokens / Max output: 1,000 tokens - Capabilities: vision - Pricing (as of 2026-10-03): Input $5.00 per 1M tokens; Output $26.00 per 1M tokens; Image input $1.75 per 1M tokens; Image output $19.50 per 1M tokens - Endpoints: openai: /v1/images/edits, /v1/images/generations - Protocols: openai - Model page: https://ofox.run/models/microsoft/mai-image-2.5-flash ```bash curl https://api.ofox.run/v1/images/generations \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "microsoft/mai-image-2.5-flash", "prompt": "A watercolor fox reading a map", "n": 1}' ``` --- ## microsoft/mai-image-2.5-pro Microsoft: MAI Image 2.5 Pro — MAI Image 2.5 Pro is the flagship of Microsoft's MAI image family, released in June 2026, built for complex compositions with strong object and character consistency, accurate material properties, and coherent spatial reasoning — well suited to photorealistic scenes and rich, documentary-style imagery. It supports both text-to-image generation and image-to-image editing, output capped at 1024×1024 (no upscaling or higher-resolution tiers). Accessible via the OpenAI-compatible /v1/images/generations and /v1/images/edits endpoints through Ofox. - Model ID: `microsoft/mai-image-2.5-pro` - Provider: Microsoft - Type: Image generation - Released: 2026-06-19 - Context window: 4,100 tokens / Max output: 1,000 tokens - Capabilities: vision - Pricing (as of 2026-10-03): Input $5.00 per 1M tokens; Output $106.00 per 1M tokens; Image input $8.00 per 1M tokens; Image output $106.00 per 1M tokens - Endpoints: openai: /v1/images/edits, /v1/images/generations - Protocols: openai - Model page: https://ofox.run/models/microsoft/mai-image-2.5-pro ```bash curl https://api.ofox.run/v1/images/generations \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "microsoft/mai-image-2.5-pro", "prompt": "A watercolor fox reading a map", "n": 1}' ``` --- # MiniMax Models Source: https://ofox.run/models/minimax 11 MiniMax models available through OfoxAI. All of them are called with the same API key and billed to the same balance as every other model on the platform. --- ## minimax/hailuo-3 MiniMax H3 — Hailuo 3 is MiniMax's multimodal video generation model with synchronized audio output. It accepts mixed text, image, video, and audio inputs, supports first- and last-frame anchoring, and uses reference images, videos, or audio to guide style. It generates clips of any integer duration from 4 to 15 seconds at 768P or 2K, with seven aspect-ratio options including adaptive. Accessible through Ofox via the OpenAI-compatible protocol. - Model ID: `minimax/hailuo-3` - Provider: MiniMax - Type: Video generation - Released: 2026-07-31 - Capabilities: video input - Pricing (as of 2026-10-03): Video output $0.08 per second - Video attributes: generation modes t2v, i2v, v2v; resolutions 768p, 2k (default 768p); duration 4-15 seconds; aspect ratios 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, adaptive; synchronized audio supported - Endpoints: openai: /v1/videos - Protocols: openai - Model page: https://ofox.run/models/minimax/hailuo-3 Per-second price tiers (as of 2026-10-03); a dimension listed as `any` matches every value: | Resolution | Input mode | Audio | Price per second | |---|---|---|---| | 768p | any | any | $0.08 | | 2k | any | any | $0.13 | | any | any | any | $0.13 | ```bash # Video generation is asynchronous: create a job, then poll it until it finishes. curl https://api.ofox.run/v1/videos \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "minimax/hailuo-3", "prompt": "A golden retriever running on the beach at sunset", "duration": 4, "resolution": "768p", "generate_audio": true}' # The response carries an id; poll it until status is completed or failed. curl https://api.ofox.run/v1/videos/ \ -H "Authorization: Bearer " ``` --- ## minimax/hailuo-3-max MiniMax H3 Max — Hailuo 3 Max is the faster version of MiniMax H3, also generating synchronized audio. It accepts mixed text, image, video, and audio inputs and supports first- and last-frame anchoring. It generates clips of any integer duration from 5 to 15 seconds at 480P or 768P, with seven aspect-ratio options including adaptive. Reference materials incur no additional charge. Accessible through Ofox via the OpenAI-compatible protocol. - Model ID: `minimax/hailuo-3-max` - Provider: MiniMax - Type: Video generation - Released: 2026-07-31 - Capabilities: video input - Pricing (as of 2026-10-03): Video output $0.05 per second - Video attributes: generation modes t2v, i2v, v2v; resolutions 480p, 768p (default 768p); duration 5-15 seconds; aspect ratios 21:9, 16:9, 4:3, 1:1, 3:4, 9:16, adaptive; synchronized audio supported - Endpoints: openai: /v1/videos - Protocols: openai - Model page: https://ofox.run/models/minimax/hailuo-3-max Per-second price tiers (as of 2026-10-03); a dimension listed as `any` matches every value: | Resolution | Input mode | Audio | Price per second | |---|---|---|---| | 480p | any | any | $0.05 | | 768p | any | any | $0.08 | | any | any | any | $0.08 | ```bash # Video generation is asynchronous: create a job, then poll it until it finishes. curl https://api.ofox.run/v1/videos \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "minimax/hailuo-3-max", "prompt": "A golden retriever running on the beach at sunset", "duration": 5, "resolution": "768p", "generate_audio": true}' # The response carries an id; poll it until status is completed or failed. curl https://api.ofox.run/v1/videos/ \ -H "Authorization: Bearer " ``` --- ## minimax/m2-her MiniMax: MiniMax M2 Her — MiniMax M2 Her is a MiniMax M2-series chat model released on 2026-01-23, aimed at agentic workloads that combine reasoning with tool calls. It supports reasoning, function calling, prompt caching, and web search, so it can plan multi-step tasks and bring in external information when a request needs it. Context: 200K tokens, output: 131K. Its low cost per token keeps long agent sessions affordable. Available via OpenAI and Anthropic protocols through Ofox. - Model ID: `minimax/m2-her` - Provider: MiniMax - Type: Chat / text generation - Released: 2026-01-23 - Context window: 200,000 tokens / Max output: 131,000 tokens - Capabilities: function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $0.30 per 1M tokens; Output $1.20 per 1M tokens - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: anthropic, openai - Model page: https://ofox.run/models/minimax/m2-her ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="minimax/m2-her", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "minimax/m2-her", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## minimax/minimax-m2 MiniMax: MiniMax M2 — MiniMax M2 is MiniMax's compact, high-efficiency large language model optimized for end-to-end coding and agentic workflows. With 10 billion activated parameters out of 230 billion total, it delivers near-frontier intelligence across general reasoning, tool use, and multi-step task execution while keeping latency and deployment cost low. It supports reasoning, function calling, prompt caching, and web search. Context: 200K tokens, output: 131K. Available via OpenAI and Anthropic protocols. - Model ID: `minimax/minimax-m2` - Provider: MiniMax - Type: Chat / text generation - Released: 2025-10-23 - Context window: 204,800 tokens / Max output: 131,000 tokens - Capabilities: function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $0.30 per 1M tokens; Output $1.20 per 1M tokens; Cache read $0.03 per 1M tokens; Cache write $0.375 per 1M tokens - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/minimax/minimax-m2 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="minimax/minimax-m2", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "minimax/minimax-m2", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## minimax/minimax-m2.1 MiniMax: MiniMax M2.1 — MiniMax M2.1 is MiniMax's lightweight, state-of-the-art large language model optimized for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world capability while maintaining exceptional latency, scalability, and cost efficiency. Capabilities include reasoning, function calling, prompt caching, and web search. Context: 200K tokens, output: 131K, leaving room for long agent traces and large single-pass edits. Available via OpenAI and Anthropic protocols. - Model ID: `minimax/minimax-m2.1` - Provider: MiniMax - Type: Chat / text generation - Released: 2025-12-23 - Context window: 204,800 tokens / Max output: 131,000 tokens - Capabilities: function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $0.30 per 1M tokens; Output $1.20 per 1M tokens; Cache read $0.03 per 1M tokens; Cache write $0.375 per 1M tokens - Endpoints: openai: /v1/chat/completions - Protocols: anthropic, openai - Model page: https://ofox.run/models/minimax/minimax-m2.1 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="minimax/minimax-m2.1", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "minimax/minimax-m2.1", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## minimax/minimax-m2.1-lightning MiniMax: MiniMax M2.1 Lightning — MiniMax M2.1 Lightning is the speed-oriented serving tier of MiniMax M2.1, MiniMax's lightweight, state-of-the-art model for coding, agentic workflows, and modern application development. With only 10 billion activated parameters, it delivers a major jump in real-world capability while maintaining exceptional latency, scalability, and cost efficiency; the Lightning tier prioritises fast responses, charging a premium on output tokens in exchange for lower latency. It supports reasoning, function calling, prompt caching, and web search. Context: 200K tokens, output: 131K. Available via OpenAI and Anthropic protocols. - Model ID: `minimax/minimax-m2.1-lightning` - Provider: MiniMax - Type: Chat / text generation - Released: 2025-12-23 - Context window: 204,800 tokens / Max output: 131,000 tokens - Capabilities: function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $0.30 per 1M tokens; Output $2.40 per 1M tokens; Cache read $0.03 per 1M tokens; Cache write $0.375 per 1M tokens - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: anthropic, openai - Model page: https://ofox.run/models/minimax/minimax-m2.1-lightning ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="minimax/minimax-m2.1-lightning", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "minimax/minimax-m2.1-lightning", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## minimax/minimax-m2.5 MiniMax: MiniMax M2.5 — MiniMax M2.5 is MiniMax's state-of-the-art large language model designed for real-world productivity. Trained across diverse, complex digital working environments, it builds on M2.1's coding expertise and extends into general office work: generating and operating Word, Excel, and PowerPoint files, switching context between different software environments, and working across agent and human teams. It scores 80.2% on SWE-Bench Verified, 51.3% on Multi-SWE-Bench, and 76.3% on BrowseComp, and is more token-efficient than previous generations because it was trained to plan its actions and output. Reasoning, function calling, prompt caching, and web search are supported. Context: 200K tokens, output: 131K. Available via OpenAI and Anthropic protocols. - Model ID: `minimax/minimax-m2.5` - Provider: MiniMax - Type: Chat / text generation - Released: 2026-02-12 - Context window: 200,000 tokens / Max output: 131,000 tokens - Capabilities: function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $0.30 per 1M tokens; Output $1.20 per 1M tokens; Cache read $0.03 per 1M tokens; Cache write $0.375 per 1M tokens - Endpoints: openai: /v1/chat/completions - Protocols: anthropic, openai - Model page: https://ofox.run/models/minimax/minimax-m2.5 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="minimax/minimax-m2.5", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "minimax/minimax-m2.5", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## minimax/minimax-m2.5-lightning MiniMax: MiniMax M2.5 Lightning — MiniMax M2.5 Lightning is the speed-oriented serving tier of MiniMax M2.5, MiniMax's state-of-the-art model for real-world productivity. It builds on M2.1's coding expertise and extends into general office work, generating and operating Word, Excel, and PowerPoint files, switching context between software environments, and working across agent and human teams, with scores of 80.2% on SWE-Bench Verified, 51.3% on Multi-SWE-Bench, and 76.3% on BrowseComp. The Lightning tier favours fast responses, charging a premium on output tokens in exchange for lower latency. Reasoning, function calling, prompt caching, and web search are supported. Context: 200K tokens, output: 131K. Available via OpenAI and Anthropic protocols. - Model ID: `minimax/minimax-m2.5-lightning` - Provider: MiniMax - Type: Chat / text generation - Released: 2026-02-12 - Context window: 200,000 tokens / Max output: 131,000 tokens - Capabilities: function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $0.30 per 1M tokens; Output $2.40 per 1M tokens; Cache read $0.03 per 1M tokens; Cache write $0.375 per 1M tokens - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: anthropic, openai - Model page: https://ofox.run/models/minimax/minimax-m2.5-lightning ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="minimax/minimax-m2.5-lightning", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "minimax/minimax-m2.5-lightning", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## minimax/minimax-m2.7 MiniMax: MiniMax M2.7 — MiniMax M2.7 is MiniMax's software-engineering focused model, delivering outstanding performance on end-to-end project delivery, log analysis and bug triaging, code security, and machine learning work. It scores 56.22% on SWE-Pro, nearly matching the level of Opus, alongside 55.6% on VIBE-Pro for complete project delivery and 57.0% on Terminal Bench 2, which measures deep understanding of complex engineering systems. Reasoning, function calling, prompt caching, and web search are all supported. Context: 200K tokens, output: 131K. Available via OpenAI and Anthropic protocols. - Model ID: `minimax/minimax-m2.7` - Provider: MiniMax - Type: Chat / text generation - Released: 2026-03-18 - Context window: 200,000 tokens / Max output: 131,000 tokens - Capabilities: function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $0.30 per 1M tokens; Output $1.20 per 1M tokens; Cache read $0.06 per 1M tokens; Cache write $0.375 per 1M tokens - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/minimax/minimax-m2.7 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="minimax/minimax-m2.7", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "minimax/minimax-m2.7", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## minimax/minimax-m2.7-highspeed MiniMax: MiniMax M2.7 Highspeed — MiniMax M2.7 Highspeed serves the same MiniMax M2.7 model with faster, more agile responses. It keeps M2.7's software-engineering strengths, including 56.22% on SWE-Pro, 55.6% on VIBE-Pro for complete project delivery, and 57.0% on Terminal Bench 2, while trading price for throughput. Reasoning, function calling, prompt caching, and web search are all supported. Context: 200K tokens, output: 131K. Available via OpenAI and Anthropic protocols. - Model ID: `minimax/minimax-m2.7-highspeed` - Provider: MiniMax - Type: Chat / text generation - Released: 2026-03-18 - Context window: 200,000 tokens / Max output: 131,000 tokens - Capabilities: function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $0.60 per 1M tokens; Output $2.40 per 1M tokens; Cache read $0.06 per 1M tokens; Cache write $0.375 per 1M tokens - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/minimax/minimax-m2.7-highspeed ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="minimax/minimax-m2.7-highspeed", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "minimax/minimax-m2.7-highspeed", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## minimax/minimax-m3 MiniMax: MiniMax M3 — MiniMax-M3 is MiniMax's flagship reasoning model from the official channel (api.minimaxi.com), featuring a 1M-token context window and 131K output. It embeds adaptive thinking and excels at complex reasoning, coding, and long-context tasks. Available via OpenAI-compatible and Anthropic protocols through Ofox. - Model ID: `minimax/minimax-m3` - Provider: MiniMax - Type: Chat / text generation - Released: 2026-06-01 - Context window: 1,131,000 tokens / Max output: 131,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $0.60 per 1M tokens; Output $2.40 per 1M tokens; Cache read $0.12 per 1M tokens - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: anthropic, openai - Model page: https://ofox.run/models/minimax/minimax-m3 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="minimax/minimax-m3", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "minimax/minimax-m3", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- # Moonshot Models Source: https://ofox.run/models/moonshotai 4 Moonshot models available through OfoxAI. All of them are called with the same API key and billed to the same balance as every other model on the platform. --- ## moonshotai/kimi-k2.6 MoonshotAI: Kimi K2.6 — Kimi K2.6 is Moonshot's latest and most intelligent model, with major improvements in general agentic tasks, coding, and visual understanding. It achieves top scores on PhD-level science benchmarks (full GPQA Diamond) and advanced reasoning challenges, with a 262K context and output window. Available via OpenAI and Anthropic protocols. - Model ID: `moonshotai/kimi-k2.6` - Provider: Moonshot - Type: Chat / text generation - Released: 2026-04-21 - Context window: 262,144 tokens / Max output: 262,144 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $0.95 per 1M tokens; Output $4.00 per 1M tokens; Cache read $0.16 per 1M tokens; Web search $0.005 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/moonshotai/kimi-k2.6 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="moonshotai/kimi-k2.6", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "moonshotai/kimi-k2.6", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## moonshotai/kimi-k2.7-code MoonshotAI: Kimi K2.7 Code — Kimi K2.7 Code is Moonshot's dedicated strong-coding reasoning model from the official channel (api.moonshot.cn), with embedded thinking for complex programming tasks. Context and output: 262K tokens each. Available via OpenAI-compatible and Anthropic protocols through Ofox. - Model ID: `moonshotai/kimi-k2.7-code` - Provider: Moonshot - Type: Chat / text generation - Released: 2026-06-12 - Context window: 262,144 tokens / Max output: 262,144 tokens - Capabilities: function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $0.95 per 1M tokens; Output $4.00 per 1M tokens; Cache read $0.19 per 1M tokens; Web search $0.005 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: anthropic, openai - Model page: https://ofox.run/models/moonshotai/kimi-k2.7-code ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="moonshotai/kimi-k2.7-code", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "moonshotai/kimi-k2.7-code", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## moonshotai/kimi-k2.7-code-highspeed MoonshotAI: Kimi K2.7 Code Highspeed — Kimi K2.7 Code Highspeed is the accelerated inference variant of Moonshot's Kimi K2.7 Code model, offering the same reasoning quality at lower latency. Context and output: 262K tokens each. Accessible via the OpenAI-compatible protocol through Ofox. - Model ID: `moonshotai/kimi-k2.7-code-highspeed` - Provider: Moonshot - Type: Chat / text generation - Released: 2026-06-12 - Context window: 262,144 tokens / Max output: 262,144 tokens - Capabilities: function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $1.90 per 1M tokens; Output $8.00 per 1M tokens; Cache read $0.38 per 1M tokens; Web search $0.005 per request - Endpoints: openai: /v1/chat/completions - Protocols: anthropic, openai - Model page: https://ofox.run/models/moonshotai/kimi-k2.7-code-highspeed ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="moonshotai/kimi-k2.7-code-highspeed", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "moonshotai/kimi-k2.7-code-highspeed", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## moonshotai/kimi-k3 MoonshotAI: Kimi K3 — Kimi K3 is Moonshot AI's flagship model, a 2.8T-parameter MoE released on 2026-07-16 with a 1M-token context window and built-in thinking exposed through reasoning_content. It supports vision input, tool use, prompt caching, and web search, making it suited to long-horizon agents that need to read documents and images, search the web, and reason across very large inputs. It runs directly on the OpenAI protocol, and over the Anthropic protocol no explicit thinking parameter is required. Context: 1M tokens, output: 1M. Available via OpenAI and Anthropic protocols through Ofox. - Model ID: `moonshotai/kimi-k3` - Provider: Moonshot - Type: Chat / text generation - Released: 2026-07-16 - Context window: 1,048,576 tokens / Max output: 1,048,576 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $3.00 per 1M tokens; Output $15.00 per 1M tokens; Cache read $0.30 per 1M tokens; Web search $0.005 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: anthropic, openai - Model page: https://ofox.run/models/moonshotai/kimi-k3 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="moonshotai/kimi-k3", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "moonshotai/kimi-k3", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- # OpenAI Models Source: https://ofox.run/models/openai 34 OpenAI models available through OfoxAI. All of them are called with the same API key and billed to the same balance as every other model on the platform. --- ## openai/gpt-4.1 GPT-4.1 — GPT-4.1 is an earlier-generation OpenAI flagship model, released on 2025-04-14 and optimized for advanced instruction following, real-world software engineering, and long-context reasoning. Its 1M-token context window suits complex document analysis and large-scale code generation, and the model supports vision, tool use, prompt caching, and audio input. Context window: 1M tokens, output: 32K. Newer GPT-5-series and Codex models have since taken over frontier coding work, but GPT-4.1 remains a dependable, lower-cost choice for long-document workloads and existing integrations. Accessible via the OpenAI-compatible protocol. - Model ID: `openai/gpt-4.1` - Provider: OpenAI - Type: Chat / text generation - Released: 2025-04-14 - Context window: 1,047,576 tokens / Max output: 32,768 tokens - Capabilities: vision, function calling, prompt caching, audio input - Pricing (as of 2026-10-03): Input $1.60 per 1M tokens (list price $2.00); Output $6.40 per 1M tokens (list price $8.00); Cache read $0.40 per 1M tokens (list price $0.50); Web search $0.014 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai - Model page: https://ofox.run/models/openai/gpt-4.1 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="openai/gpt-4.1", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-4.1", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## openai/gpt-4.1-mini GPT-4.1 Mini — GPT-4.1 Mini is the mid-sized variant of OpenAI's GPT-4.1, released on 2025-04-14, delivering GPT-4o-level performance at lower latency and cost. It supports structured outputs, vision understanding, tool use, and prompt caching, and keeps the full 1M-token context window of the larger model. Context window: 1M tokens, output: 32K. It remains a budget option for high-volume long-context processing, although later generations now lead on coding and reasoning. Accessible via the OpenAI-compatible protocol. - Model ID: `openai/gpt-4.1-mini` - Provider: OpenAI - Type: Chat / text generation - Released: 2025-04-14 - Context window: 1,047,576 tokens / Max output: 32,768 tokens - Capabilities: vision, function calling, prompt caching - Pricing (as of 2026-10-03): Input $0.32 per 1M tokens (list price $0.40); Output $1.28 per 1M tokens (list price $1.60); Cache read $0.08 per 1M tokens (list price $0.10); Web search $0.014 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai - Model page: https://ofox.run/models/openai/gpt-4.1-mini ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="openai/gpt-4.1-mini", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-4.1-mini", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## openai/gpt-4o GPT-4o — GPT-4o is OpenAI's flagship multimodal model from the GPT-4 era, released on 2024-05-13, taking text and image inputs and returning text. It maintains GPT-4 Turbo intelligence at twice the speed and 50% lower cost, and is optimized for complex reasoning, coding, and visual understanding, with tool use, prompt caching, and audio input. Context: 128K tokens, output: 16K. It has since been succeeded by the GPT-4.1 and GPT-5 series, so newer models are the better starting point for new projects, but GPT-4o remains widely used in established integrations. Accessible via the OpenAI-compatible protocol. - Model ID: `openai/gpt-4o` - Provider: OpenAI - Type: Chat / text generation - Released: 2024-05-13 - Context window: 128,000 tokens / Max output: 16,384 tokens - Capabilities: vision, function calling, prompt caching, audio input - Pricing (as of 2026-10-03): Input $2.00 per 1M tokens (list price $2.50); Output $8.00 per 1M tokens (list price $10.00); Cache read $1.00 per 1M tokens (list price $1.25); Web search $0.014 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai - Model page: https://ofox.run/models/openai/gpt-4o ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="openai/gpt-4o", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-4o", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## openai/gpt-4o-mini GPT-4o Mini — GPT-4o Mini is OpenAI's small, cost-efficient model from the GPT-4o generation, released on 2024-07-18, priced well below GPT-3.5 Turbo while scoring 82% on MMLU. It targets high-volume tasks that still need solid language understanding and generation, and supports vision, tool use, and prompt caching. Context: 128K tokens, output: 16K. It remains one of the cheapest ways to run simple, high-throughput workloads, though newer small models offer stronger reasoning. Accessible via the OpenAI-compatible protocol. - Model ID: `openai/gpt-4o-mini` - Provider: OpenAI - Type: Chat / text generation - Released: 2024-07-18 - Context window: 128,000 tokens / Max output: 16,384 tokens - Capabilities: vision, function calling, prompt caching - Pricing (as of 2026-10-03): Input $0.12 per 1M tokens (list price $0.15); Output $0.48 per 1M tokens (list price $0.60); Cache read $0.06 per 1M tokens (list price $0.075); Web search $0.014 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai - Model page: https://ofox.run/models/openai/gpt-4o-mini ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="openai/gpt-4o-mini", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-4o-mini", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## openai/gpt-4o-mini-transcribe OpenAI: GPT 4o Mini Transcribe — GPT-4o Mini Transcribe is OpenAI's smaller, cost-efficient speech-to-text model, built on GPT-4o Mini audio capabilities and released on 2025-12-15. It takes audio input and returns transcribed text, and is billed per token for both input and output rather than per minute, which makes it a fit for high-volume transcription workflows that want token-level billing transparency at a lower cost point. Accessible via the OpenAI-compatible protocol through Ofox. - Model ID: `openai/gpt-4o-mini-transcribe` - Provider: OpenAI - Type: Audio transcription - Released: 2025-12-15 - Context window: 128,000 tokens / Max output: 128,000 tokens - Capabilities: audio input - Pricing (as of 2026-10-03): Input $1.00 per 1M tokens (list price $1.25); Output $4.00 per 1M tokens (list price $5.00); Audio input $3.00 per 1M tokens; Web search $0.014 per request - Endpoints: openai: /v1/audio/transcriptions - Protocols: openai - Model page: https://ofox.run/models/openai/gpt-4o-mini-transcribe ```bash curl https://api.ofox.run/v1/audio/transcriptions \ -H "Authorization: Bearer " \ -F model="openai/gpt-4o-mini-transcribe" \ -F file="@audio.mp3" ``` --- ## openai/gpt-4o-transcribe-diarize OpenAI: GPT 4o Transcribe Diarize — GPT-4o Transcribe Diarize is OpenAI's high-quality speech-to-text model, built on GPT-4o audio capabilities and released on 2025-10-15, with speaker diarization for multi-speaker recordings such as meetings and interviews. It accepts audio input and returns transcribed text, and is billed per token for input and output rather than per minute, which keeps costs transparent at the token level. Accessible via the OpenAI-compatible protocol through Ofox. - Model ID: `openai/gpt-4o-transcribe-diarize` - Provider: OpenAI - Type: Audio transcription - Released: 2025-10-15 - Context window: 128,000 tokens / Max output: 128,000 tokens - Capabilities: audio input - Pricing (as of 2026-10-03): Input $2.00 per 1M tokens (list price $2.50); Output $8.00 per 1M tokens (list price $10.00); Audio input $6.00 per 1M tokens; Web search $0.014 per request - Endpoints: openai: /v1/audio/transcriptions - Protocols: openai - Model page: https://ofox.run/models/openai/gpt-4o-transcribe-diarize ```bash curl https://api.ofox.run/v1/audio/transcriptions \ -H "Authorization: Bearer " \ -F model="openai/gpt-4o-transcribe-diarize" \ -F file="@audio.mp3" ``` --- ## openai/gpt-5 GPT-5 — GPT-5 is OpenAI's next-generation flagship model, released on 2025-08-07, with advanced multimodal capabilities and state-of-the-art general performance. It supports reasoning, vision, audio and video input, tool use (function calling), prompt caching, and web search, making it well suited to agentic workflows. Context: 256K tokens, output: 64K. Accessible via the OpenAI-compatible protocol through Ofox. - Model ID: `openai/gpt-5` - Provider: OpenAI - Type: Chat / text generation - Released: 2025-08-07 - Context window: 256,000 tokens / Max output: 65,536 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, audio input, video input - Pricing (as of 2026-10-03): Input $1.00 per 1M tokens (list price $1.25); Output $8.00 per 1M tokens (list price $10.00); Cache read $0.104 per 1M tokens (list price $0.13); Image output $25.60 per 1M tokens (list price $32.00); Web search $0.014 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai - Model page: https://ofox.run/models/openai/gpt-5 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="openai/gpt-5", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-5", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## openai/gpt-5-mini GPT-5 Mini — GPT-5 Mini is OpenAI's cost-effective GPT-5 variant, released on 2025-08-07 and tuned for balanced everyday performance at a fraction of the flagship price. It supports vision, audio input, tool use (function calling), and prompt caching, covering most general-purpose workloads: images and recordings can be passed in directly, while prompt caching keeps repeated system prompts inexpensive across turns. Context: 256K tokens, output: 32K. Accessible via the OpenAI-compatible protocol. - Model ID: `openai/gpt-5-mini` - Provider: OpenAI - Type: Chat / text generation - Released: 2025-08-07 - Context window: 256,000 tokens / Max output: 32,768 tokens - Capabilities: vision, function calling, prompt caching, audio input - Pricing (as of 2026-10-03): Input $0.20 per 1M tokens (list price $0.25); Output $1.60 per 1M tokens (list price $2.00); Cache read $0.024 per 1M tokens (list price $0.03); Image output $25.60 per 1M tokens (list price $32.00); Web search $0.014 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai - Model page: https://ofox.run/models/openai/gpt-5-mini ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="openai/gpt-5-mini", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-5-mini", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## openai/gpt-5-nano GPT-5 Nano — GPT-5 Nano is the ultra-efficient member of the GPT-5 family, released on 2025-08-07 and optimized for maximum throughput at minimal cost. Its low-latency design targets high-volume automation, with support for tool use (function calling) and prompt caching, so repeated instructions stay inexpensive across large batches. Context: 128K tokens, output: 16K. Accessible via the OpenAI-compatible protocol. - Model ID: `openai/gpt-5-nano` - Provider: OpenAI - Type: Chat / text generation - Released: 2025-08-07 - Context window: 128,000 tokens / Max output: 16,384 tokens - Capabilities: function calling, prompt caching - Pricing (as of 2026-10-03): Input $0.04 per 1M tokens (list price $0.05); Output $0.32 per 1M tokens (list price $0.40); Cache read $0.008 per 1M tokens (list price $0.01); Image output $25.60 per 1M tokens (list price $32.00); Web search $0.014 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai - Model page: https://ofox.run/models/openai/gpt-5-nano ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="openai/gpt-5-nano", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-5-nano", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## openai/gpt-5.1 OpenAI: GPT-5.1 — GPT-5.1 is OpenAI's enhanced GPT-5 release from 2025-11-13, bringing improved reasoning, better instruction following, and stronger coding performance. It handles vision, audio, and video input alongside reasoning, tool use (function calling), prompt caching, and web search. Context: 256K tokens, output: 128K — four times the output budget of GPT-5. Accessible via the OpenAI-compatible protocol through Ofox. - Model ID: `openai/gpt-5.1` - Provider: OpenAI - Type: Chat / text generation - Released: 2025-11-13 - Context window: 256,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, audio input, video input - Pricing (as of 2026-10-03): Input $1.00 per 1M tokens (list price $1.25); Output $8.00 per 1M tokens (list price $10.00); Cache read $0.104 per 1M tokens (list price $0.13); Image output $25.60 per 1M tokens (list price $32.00); Web search $0.014 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai - Model page: https://ofox.run/models/openai/gpt-5.1 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="openai/gpt-5.1", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-5.1", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## openai/gpt-5.1-codex-max OpenAI: GPT-5.1 Codex Max — GPT-5.1 Codex Max is OpenAI's agentic coding model, released on 2025-12-04 and designed for long-running, high-context software development tasks. It is based on an updated version of the 5.1 reasoning stack and trained on agentic workflows spanning software engineering, mathematics, and research. Capabilities include reasoning, tool use, prompt caching, web search, and multimodal input covering images, audio, and video. Context window: 256K tokens, output: 128K. Prompt caching keeps long agent sessions that reuse the same repository context affordable. Accessible via the OpenAI-compatible protocol through Ofox. - Model ID: `openai/gpt-5.1-codex-max` - Provider: OpenAI - Type: Chat / text generation - Released: 2025-12-04 - Context window: 256,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, audio input, video input - Pricing (as of 2026-10-03): Input $1.00 per 1M tokens (list price $1.25); Output $8.00 per 1M tokens (list price $10.00); Cache read $0.104 per 1M tokens (list price $0.13); Web search $0.014 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai - Model page: https://ofox.run/models/openai/gpt-5.1-codex-max ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="openai/gpt-5.1-codex-max", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-5.1-codex-max", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## openai/gpt-5.1-codex-mini GPT-5.1 Codex Mini — GPT-5.1 Codex Mini is a smaller and faster version of GPT-5.1-Codex from OpenAI, released on 2025-11-13 for coding work where latency and cost matter more than peak capability. It supports reasoning, tool use, prompt caching, web search, and multimodal input covering images, audio, and video. Context window: 256K tokens, output: 64K. It is the most economical option in the Codex line and fits high-volume agent loops and routine code edits. Accessible via the OpenAI-compatible protocol. - Model ID: `openai/gpt-5.1-codex-mini` - Provider: OpenAI - Type: Chat / text generation - Released: 2025-11-13 - Context window: 256,000 tokens / Max output: 65,536 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, audio input, video input - Pricing (as of 2026-10-03): Input $0.20 per 1M tokens (list price $0.25); Output $1.60 per 1M tokens (list price $2.00); Cache read $0.024 per 1M tokens (list price $0.03); Web search $0.014 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai - Model page: https://ofox.run/models/openai/gpt-5.1-codex-mini ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="openai/gpt-5.1-codex-mini", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-5.1-codex-mini", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## openai/gpt-5.2 OpenAI: GPT-5.2 — GPT-5.2 is OpenAI's December 2025 flagship update, extending the context window to 512K tokens. It pairs enhanced reasoning with full multimodal input — vision, audio, and video — plus tool use (function calling), prompt caching, and integrated web search, making it a fit for long-document and research-heavy workflows. Context: 512K tokens, output: 128K. Accessible via the OpenAI-compatible protocol. - Model ID: `openai/gpt-5.2` - Provider: OpenAI - Type: Chat / text generation - Released: 2025-12-11 - Context window: 512,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, audio input, video input - Pricing (as of 2026-10-03): Input $1.40 per 1M tokens (list price $1.75); Output $11.20 per 1M tokens (list price $14.00); Cache read $0.144 per 1M tokens (list price $0.18); Image output $25.60 per 1M tokens (list price $32.00); Web search $0.014 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai - Model page: https://ofox.run/models/openai/gpt-5.2 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="openai/gpt-5.2", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-5.2", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## openai/gpt-5.2-codex OpenAI: GPT-5.2 Codex — GPT-5.2 Codex is OpenAI's upgraded Codex model for software engineering and coding workflows, released on 2026-01-14. It is built for both interactive development sessions and long, independent execution of complex engineering tasks, covering projects built from scratch, feature development, debugging, large-scale refactoring, and code review. Compared with GPT-5.1-Codex it is more steerable, adheres closely to developer instructions, and produces cleaner, higher-quality code; reasoning effort can be tuned with the reasoning.effort parameter. It also supports tool use, prompt caching, web search, and image, audio, and video input. Context window: 512K tokens, output: 128K. Accessible via the OpenAI-compatible protocol. - Model ID: `openai/gpt-5.2-codex` - Provider: OpenAI - Type: Chat / text generation - Released: 2026-01-14 - Context window: 512,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, audio input, video input - Pricing (as of 2026-10-03): Input $1.40 per 1M tokens (list price $1.75); Output $11.20 per 1M tokens (list price $14.00); Cache read $0.144 per 1M tokens (list price $0.18); Web search $0.014 per request - Endpoints: openai: /v1/responses - Protocols: openai - Model page: https://ofox.run/models/openai/gpt-5.2-codex ```bash # This model is served on the Responses API only — it does not accept /v1/chat/completions. curl https://api.ofox.run/v1/responses \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-5.2-codex", "input": "Hello!"}' ``` --- ## openai/gpt-5.3-codex OpenAI: GPT-5.3 Codex — GPT-5.3 Codex is OpenAI's most advanced agentic coding model, released on 2026-02-25, combining the frontier software engineering performance of GPT-5.2-Codex with the broader reasoning and professional knowledge of GPT-5.2. It reaches state-of-the-art results on SWE-Bench Pro and performs strongly on Terminal-Bench 2.0 and OSWorld-Verified, reflecting improved multi-language coding, terminal proficiency, and real-world computer-use skills. It supports reasoning, tool use, prompt caching, web search, and image, audio, and video input. Context window: 512K tokens, output: 128K. Accessible via the OpenAI-compatible protocol through Ofox. - Model ID: `openai/gpt-5.3-codex` - Provider: OpenAI - Type: Chat / text generation - Released: 2026-02-25 - Context window: 512,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, audio input, video input - Pricing (as of 2026-10-03): Input $1.40 per 1M tokens (list price $1.75); Output $11.20 per 1M tokens (list price $14.00); Cache read $0.144 per 1M tokens (list price $0.18); Web search $0.014 per request - Endpoints: openai: /v1/responses - Protocols: openai - Model page: https://ofox.run/models/openai/gpt-5.3-codex ```bash # This model is served on the Responses API only — it does not accept /v1/chat/completions. curl https://api.ofox.run/v1/responses \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-5.3-codex", "input": "Hello!"}' ``` --- ## openai/gpt-5.4 OpenAI: GPT-5.4 — GPT-5.4 is OpenAI's frontier model, unifying the Codex and GPT lines into a single system. It takes text and image input for high-context reasoning, coding, and multimodal analysis, with improved document understanding, tool use, and instruction following — positioned as a strong default for both general-purpose work and software engineering, generating production-quality code and executing complex multi-step workflows with fewer iterations. Reasoning, prompt caching, and web search are supported. Context: 1M+ tokens (922K input), output: 128K. Available via OpenAI and Anthropic protocols. - Model ID: `openai/gpt-5.4` - Provider: OpenAI - Type: Chat / text generation - Released: 2026-03-05 - Context window: 1,050,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $2.00 per 1M tokens (list price $2.50); Output $12.00 per 1M tokens (list price $15.00); Cache read $0.20 per 1M tokens (list price $0.25); Image output $25.60 per 1M tokens (list price $32.00); Web search $0.014 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/openai/gpt-5.4 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="openai/gpt-5.4", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-5.4", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## openai/gpt-5.4-mini OpenAI: GPT-5.4 Mini — GPT-5.4 Mini brings the core capabilities of GPT-5.4 to a faster, more efficient model built for high-throughput workloads, released on 2026-03-17. It supports text and image input with solid reasoning, coding, and tool use while cutting latency and cost for large-scale deployments, and adds prompt caching and web search. Context: 400K tokens, output: 128K. Available via OpenAI and Anthropic protocols. - Model ID: `openai/gpt-5.4-mini` - Provider: OpenAI - Type: Chat / text generation - Released: 2026-03-17 - Context window: 400,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $0.60 per 1M tokens (list price $0.75); Output $3.60 per 1M tokens (list price $4.50); Cache read $0.06 per 1M tokens (list price $0.075); Image output $25.60 per 1M tokens (list price $32.00); Web search $0.014 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/openai/gpt-5.4-mini ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="openai/gpt-5.4-mini", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-5.4-mini", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## openai/gpt-5.4-nano OpenAI: GPT-5.4 Nano — GPT-5.4 Nano is the lightest and most cost-efficient member of the GPT-5.4 family, released on 2026-03-17 and optimized for speed-critical, high-volume tasks such as classification, data extraction, ranking, and sub-agent execution. It accepts text and image input and supports reasoning, tool use (function calling), prompt caching, and web search. Context: 400K tokens, output: 128K. Available via OpenAI and Anthropic protocols. - Model ID: `openai/gpt-5.4-nano` - Provider: OpenAI - Type: Chat / text generation - Released: 2026-03-17 - Context window: 400,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $0.16 per 1M tokens (list price $0.20); Output $1.00 per 1M tokens (list price $1.25); Cache read $0.016 per 1M tokens (list price $0.02); Image output $25.60 per 1M tokens (list price $32.00); Web search $0.014 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/openai/gpt-5.4-nano ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="openai/gpt-5.4-nano", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-5.4-nano", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## openai/gpt-5.4-pro OpenAI: GPT-5.4 Pro — GPT-5.4 Pro is OpenAI's most advanced model, building on GPT-5.4's unified architecture with enhanced reasoning for complex, high-stakes tasks. Tuned for step-by-step reasoning, instruction following, and accuracy, it excels at agentic coding, long-context workflows, and multi-step problem solving, with text and image input, tool use (function calling), and web search. Context: 1M+ tokens (922K input), output: 128K. Accessible via the OpenAI-compatible protocol. - Model ID: `openai/gpt-5.4-pro` - Provider: OpenAI - Type: Chat / text generation - Released: 2026-03-05 - Context window: 1,050,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, web search - Pricing (as of 2026-10-03): Input $24.00 per 1M tokens (list price $30.00); Output $144.00 per 1M tokens (list price $180.00); Image output $25.60 per 1M tokens (list price $32.00); Web search $0.014 per request - Endpoints: openai: /v1/responses - Protocols: openai - Model page: https://ofox.run/models/openai/gpt-5.4-pro ```bash # This model is served on the Responses API only — it does not accept /v1/chat/completions. curl https://api.ofox.run/v1/responses \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-5.4-pro", "input": "Hello!"}' ``` --- ## openai/gpt-5.5 OpenAI: GPT-5.5 — GPT-5.5 is OpenAI's frontier model for complex professional workloads, released on 2026-04-24, building on GPT-5.4 with stronger reasoning, higher reliability, and better token efficiency on hard tasks. It takes text and image input and supports tool use (function calling), prompt caching, and web search, enabling large-scale reasoning, coding, and multimodal workflows in a single system. Context: 1M+ tokens (922K input), output: 128K. Available via OpenAI and Anthropic protocols. - Model ID: `openai/gpt-5.5` - Provider: OpenAI - Type: Chat / text generation - Released: 2026-04-24 - Context window: 1,050,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $4.00 per 1M tokens (list price $5.00); Output $24.00 per 1M tokens (list price $30.00); Cache read $0.40 per 1M tokens (list price $0.50); Image output $25.60 per 1M tokens (list price $32.00); Web search $0.014 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/openai/gpt-5.5 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="openai/gpt-5.5", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-5.5", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## openai/gpt-5.6-luna OpenAI: GPT-5.6 Luna — GPT-5.6 Luna is the fast, cost-efficient tier of OpenAI's GPT-5.6 series, released on 2026-07-09. It is built for high-volume, latency-sensitive workloads such as chat, classification, and lightweight agentic workflows, providing capable reasoning for its price tier. It supports reasoning, vision, tool use (function calling), prompt caching, and web search. Context window: 1M tokens, output: 128K. Prompt caching keeps repeated context inexpensive across high-traffic sessions. Available via OpenAI and Anthropic protocols through Ofox. - Model ID: `openai/gpt-5.6-luna` - Provider: OpenAI - Type: Chat / text generation - Released: 2026-07-09 - Context window: 1,050,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $0.20 per 1M tokens (list price $1.00); Output $1.20 per 1M tokens (list price $6.00); Cache read $0.02 per 1M tokens (list price $0.10); Cache write $0.25 per 1M tokens (list price $1.25); Image output $32.00 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/openai/gpt-5.6-luna ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="openai/gpt-5.6-luna", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-5.6-luna", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## openai/gpt-5.6-sol OpenAI: GPT-5.6 Sol — GPT-5.6 Sol is the flagship model of OpenAI's GPT-5.6 series, released on 2026-07-09. It targets complex reasoning, coding, and agentic workflows, and is particularly strong at command-line work, multi-step coding tasks, and long-horizon problem solving. It supports reasoning, vision, tool use (function calling), prompt caching, and web search. Context window: 1M tokens, output: 128K, so most mid-sized repositories or large document sets can stay resident across a long session. Available via OpenAI and Anthropic protocols through Ofox. - Model ID: `openai/gpt-5.6-sol` - Provider: OpenAI - Type: Chat / text generation - Released: 2026-07-09 - Context window: 1,050,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $2.50 per 1M tokens (list price $5.00); Output $15.00 per 1M tokens (list price $30.00); Cache read $0.25 per 1M tokens (list price $0.50); Cache write $3.125 per 1M tokens (list price $6.25); Image output $32.00 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/openai/gpt-5.6-sol ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="openai/gpt-5.6-sol", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-5.6-sol", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## openai/gpt-5.6-terra OpenAI: GPT-5.6 Terra — GPT-5.6 Terra is the balanced model in OpenAI's GPT-5.6 series, released on 2026-07-09 and positioned between the flagship Sol tier and the cost-efficient Luna tier. It suits everyday coding, reasoning, and agentic tasks where capability and cost have to be balanced, offering strong performance at a fraction of the flagship's cost. It supports reasoning, vision, tool use (function calling), prompt caching, and web search. Context window: 1M tokens, output: 128K. Available via OpenAI and Anthropic protocols through Ofox. - Model ID: `openai/gpt-5.6-terra` - Provider: OpenAI - Type: Chat / text generation - Released: 2026-07-09 - Context window: 1,050,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $2.00 per 1M tokens (list price $2.50); Output $12.00 per 1M tokens (list price $15.00); Cache read $0.20 per 1M tokens (list price $0.25); Cache write $2.50 per 1M tokens (list price $3.125); Image output $32.00 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/openai/gpt-5.6-terra ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="openai/gpt-5.6-terra", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-5.6-terra", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## openai/gpt-6-astra OpenAI: GPT-6 Astra — GPT-6 Astra is OpenAI's flagship model for demanding end-to-end work, released on 2026-09-04. It is built for advanced analysis, software engineering, deep research, scientific work, and document creation, with particular strength on long-horizon tasks that unfold over many steps rather than a single prompt. It supports reasoning, vision, tool use, prompt caching, and web search. Context window: 1.05M tokens, output: 128K. The oversized context makes it practical to hold entire codebases or research corpora in one session, while prompt caching keeps the repeated portion affordable. Available via OpenAI and Anthropic protocols through Ofox. - Model ID: `openai/gpt-6-astra` - Provider: OpenAI - Type: Chat / text generation - Released: 2026-09-04 - Context window: 1,050,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $8.00 per 1M tokens (list price $10.00); Output $40.00 per 1M tokens (list price $50.00); Cache read $0.80 per 1M tokens (list price $1.00); Cache write $10.00 per 1M tokens (list price $12.50); Web search $0.01 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/openai/gpt-6-astra ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="openai/gpt-6-astra", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-6-astra", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## openai/gpt-6-luna OpenAI: GPT-6 Luna — GPT-6 Luna is OpenAI's fastest and most cost-efficient GPT-6 model, built for high-volume, latency-sensitive tasks such as classification, structured extraction, and lightweight agents. It accepts text and image inputs and produces text, with adjustable reasoning effort. It supports vision, function calling, web search, and prompt caching. Context window: 1.05M tokens, output: 128K, allowing large inputs and extended responses within one request. Available through Ofox via OpenAI and Anthropic protocols, including OpenAI-compatible Chat Completions and Responses endpoints. - Model ID: `openai/gpt-6-luna` - Provider: OpenAI - Type: Chat / text generation - Released: 2026-09-22 - Context window: 1,050,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $0.08 per 1M tokens (list price $0.10); Output $0.40 per 1M tokens (list price $0.50); Cache read $0.008 per 1M tokens (list price $0.01); Cache write $0.10 per 1M tokens (list price $0.125); Web search $0.01 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/openai/gpt-6-luna ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="openai/gpt-6-luna", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-6-luna", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## openai/gpt-6-sol OpenAI: GPT-6 Sol — GPT-6 Sol is the balanced model in OpenAI's GPT-6 family, designed for reasoning, coding, and agentic work at a lower cost than GPT-6 Astra. It accepts text and image inputs and produces text, with adjustable reasoning effort. It supports vision, function calling, web search, and prompt caching, covering both interactive assistants and tool-using workflows. Context window: 1.05M tokens, output: 128K. Available through Ofox via OpenAI and Anthropic protocols, including OpenAI-compatible Chat Completions and Responses endpoints. - Model ID: `openai/gpt-6-sol` - Provider: OpenAI - Type: Chat / text generation - Released: 2026-09-22 - Context window: 1,050,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $1.60 per 1M tokens (list price $2.00); Output $8.00 per 1M tokens (list price $10.00); Cache read $0.16 per 1M tokens (list price $0.20); Cache write $2.00 per 1M tokens (list price $2.50); Web search $0.01 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/openai/gpt-6-sol ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="openai/gpt-6-sol", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-6-sol", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## openai/gpt-6.1-sol OpenAI: GPT-6.1 Sol — GPT-6.1 Sol is OpenAI's balanced model for complex coding, computer use, and professional work. It offers near-Astra performance at lower cost, supports text and image inputs with text output, adjustable reasoning effort, a 1.05M-token context window, and up to 128K output tokens. - Model ID: `openai/gpt-6.1-sol` - Provider: OpenAI - Type: Chat / text generation - Released: 2026-09-29 - Context window: 1,050,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $2.00 per 1M tokens; Output $10.00 per 1M tokens; Cache read $0.10 per 1M tokens; Cache write $2.50 per 1M tokens - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/openai/gpt-6.1-sol ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="openai/gpt-6.1-sol", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-6.1-sol", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## openai/gpt-image-1.5 OpenAI: GPT Image 1.5 — GPT Image 1.5 is OpenAI's image generation model, released on 2025-12-16, built for better instruction following and closer adherence to prompts. It accepts image input alongside text, so it covers both text-to-image generation and image-guided editing workflows, and it supports prompt caching to cut the cost of repeated prompt prefixes. Cached image input is billed at a lower rate than fresh image input, which helps on iterative generation where the same reference material is reused across requests. Accessible via the OpenAI-compatible protocol through Ofox. - Model ID: `openai/gpt-image-1.5` - Provider: OpenAI - Type: Image generation - Released: 2025-12-16 - Capabilities: vision, prompt caching - Pricing (as of 2026-10-03): Input $4.00 per 1M tokens (list price $5.00); Output $8.00 per 1M tokens (list price $10.00); Image input $8.00 per 1M tokens; Cache read $1.00 per 1M tokens (list price $1.25); Cached image input $2.00 per 1M tokens; Image output $25.60 per 1M tokens (list price $32.00) - Endpoints: openai: /v1/images/edits, /v1/images/generations - Protocols: openai - Model page: https://ofox.run/models/openai/gpt-image-1.5 ```bash curl https://api.ofox.run/v1/images/generations \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-image-1.5", "prompt": "A watercolor fox reading a map", "n": 1}' ``` --- ## openai/gpt-image-2 OpenAI: GPT Image 2 — GPT Image 2 is OpenAI's most recent image generation model, released on 2026-04-21. It improves on overall performance and image quality while adding finer editing controls and face preservation. The model supports high input fidelity and can add or remove a single aspect of an image while keeping everything else intact, with improvements in aspect ratio handling, resolution, and editing capability. Image input is accepted for edit and reference workflows, and prompt caching is available, with cached image input billed below fresh input. Accessible via the OpenAI-compatible protocol through Ofox. - Model ID: `openai/gpt-image-2` - Provider: OpenAI - Type: Image generation - Released: 2026-04-21 - Capabilities: vision, prompt caching - Pricing (as of 2026-10-03): Input $4.00 per 1M tokens (list price $5.00); Output $24.00 per 1M tokens (list price $30.00); Image input $8.00 per 1M tokens; Cache read $1.00 per 1M tokens (list price $1.25); Cached image input $2.00 per 1M tokens; Image output $24.00 per 1M tokens (list price $30.00) - Endpoints: openai: /v1/images/edits, /v1/images/generations - Protocols: openai - Model page: https://ofox.run/models/openai/gpt-image-2 ```bash curl https://api.ofox.run/v1/images/generations \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-image-2", "prompt": "A watercolor fox reading a map", "n": 1}' ``` --- ## openai/gpt-image-2.5-flare OpenAI: GPT Image 2.5 Flare — GPT Image 2.5 Flare is OpenAI's model for fast, high-quality everyday image generation. It supports image input and prompt caching. Accessible through Ofox via the OpenAI-compatible protocol. - Model ID: `openai/gpt-image-2.5-flare` - Provider: OpenAI - Type: Image generation - Released: 2026-09-09 - Capabilities: vision, prompt caching - Pricing (as of 2026-10-03): Input $4.00 per 1M tokens (list price $5.00); Output $24.00 per 1M tokens (list price $30.00); Image input $8.00 per 1M tokens; Cache read $1.00 per 1M tokens (list price $1.25); Cached image input $2.00 per 1M tokens; Image output $24.00 per 1M tokens (list price $30.00) - Endpoints: openai: /v1/images/edits, /v1/images/generations - Protocols: openai - Model page: https://ofox.run/models/openai/gpt-image-2.5-flare ```bash curl https://api.ofox.run/v1/images/generations \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-image-2.5-flare", "prompt": "A watercolor fox reading a map", "n": 1}' ``` --- ## openai/gpt-image-2.5-sunburst OpenAI: GPT Image 2.5 Sunburst — GPT Image 2.5 Sunburst is OpenAI's model for capable image generation and editing. It supports image input and prompt caching. Accessible through Ofox via the OpenAI-compatible protocol. - Model ID: `openai/gpt-image-2.5-sunburst` - Provider: OpenAI - Type: Image generation - Released: 2026-09-09 - Capabilities: vision, prompt caching - Pricing (as of 2026-10-03): Input $4.00 per 1M tokens (list price $5.00); Output $24.00 per 1M tokens (list price $30.00); Image input $8.00 per 1M tokens; Cache read $1.00 per 1M tokens (list price $1.25); Cached image input $2.00 per 1M tokens; Image output $24.00 per 1M tokens (list price $30.00) - Endpoints: openai: /v1/images/edits, /v1/images/generations - Protocols: openai - Model page: https://ofox.run/models/openai/gpt-image-2.5-sunburst ```bash curl https://api.ofox.run/v1/images/generations \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "openai/gpt-image-2.5-sunburst", "prompt": "A watercolor fox reading a map", "n": 1}' ``` --- ## openai/gpt-transcribe OpenAI: GPT Transcribe — GPT Transcribe is OpenAI's speech-to-text model, released on 2026-07-28 alongside gpt-live-transcribe. It delivers lower word error rates than the GPT-4o transcription family across multilingual benchmarks, and is billed by audio duration rather than tokens, keeping costs predictable for transcription workloads. Accessible via the OpenAI-compatible protocol through Ofox. - Model ID: `openai/gpt-transcribe` - Provider: OpenAI - Type: Audio transcription - Released: 2026-07-28 - Context window: 128,000 tokens / Max output: 128,000 tokens - Capabilities: audio input - Endpoints: openai: /v1/audio/transcriptions - Protocols: openai - Model page: https://ofox.run/models/openai/gpt-transcribe ```bash curl https://api.ofox.run/v1/audio/transcriptions \ -H "Authorization: Bearer " \ -F model="openai/gpt-transcribe" \ -F file="@audio.mp3" ``` --- ## openai/text-embedding-3-large OpenAI: Text Embedding 3 Large — Text Embedding 3 Large is OpenAI's most capable embedding model, strong on both English and non-English text. It turns text into numerical vectors whose distances measure how closely two pieces of text are related, which makes it a fit for search, clustering, recommendations, anomaly detection, and classification. Input limit: 8K tokens per request. It stays affordable for large-scale corpus indexing as well as query-time embedding. Accessible via the OpenAI-compatible protocol through Ofox. - Model ID: `openai/text-embedding-3-large` - Provider: OpenAI - Type: Text embedding - Released: 2025-10-30 - Context window: 8,200 tokens / Max output: 8,200 tokens - Capabilities: none declared - Pricing (as of 2026-10-03): Input $0.14 per 1M tokens; Web search $0.014 per request - Endpoints: openai: /v1/embeddings - Protocols: openai - Model page: https://ofox.run/models/openai/text-embedding-3-large ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") result = client.embeddings.create(model="openai/text-embedding-3-large", input="The quick brown fox") print(len(result.data[0].embedding)) ``` --- ## openai/text-embedding-3-small OpenAI: Text Embedding 3 Small — Text Embedding 3 Small is OpenAI's improved, more performant successor to the ada embedding model. It converts text into numerical vectors whose distances measure how closely two pieces of text are related, covering search, clustering, recommendations, anomaly detection, and classification. Input limit: 8K tokens per request. It is the lowest-cost option in the Text Embedding 3 pair, which makes it the practical default for embedding large document sets or high-traffic retrieval. Accessible via the OpenAI-compatible protocol through Ofox. - Model ID: `openai/text-embedding-3-small` - Provider: OpenAI - Type: Text embedding - Released: 2025-10-30 - Context window: 8,200 tokens / Max output: 8,200 tokens - Capabilities: none declared - Pricing (as of 2026-10-03): Input $0.02 per 1M tokens; Web search $0.014 per request - Endpoints: openai: /v1/embeddings - Protocols: openai - Model page: https://ofox.run/models/openai/text-embedding-3-small ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") result = client.embeddings.create(model="openai/text-embedding-3-small", input="The quick brown fox") print(len(result.data[0].embedding)) ``` --- # Qwen Models Source: https://ofox.run/models/qwen 28 Qwen models available through OfoxAI. All of them are called with the same API key and billed to the same balance as every other model on the platform. --- ## qwen/qwen-flash Qwen Flash — Qwen Flash is the fastest and lowest-cost model in Alibaba's Qwen family, served through Dashscope and built for latency-sensitive tasks. It delivers ultra-fast inference while supporting tool use (function calling), prompt caching, and web search. Context window: 1M tokens, output: 32K. The 1M-token context lets it scan long documents in a single pass, and prompt caching keeps repeated instructions cheap in high-volume pipelines. Available via OpenAI and Anthropic protocols. - Model ID: `qwen/qwen-flash` - Provider: Qwen - Type: Chat / text generation - Released: 2025-07-28 - Context window: 1,000,000 tokens / Max output: 32,000 tokens - Capabilities: function calling, prompt caching, web search - Pricing (as of 2026-10-03): Input $0.022 per 1M tokens; Output $0.22 per 1M tokens; Cache read $0.0043 per 1M tokens; Cache write $0.027 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/qwen/qwen-flash ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="qwen/qwen-flash", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "qwen/qwen-flash", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## qwen/qwen-image-3.0 Qwen-Image 3.0 — Qwen-Image 3.0 is Alibaba's image generation model, officially released in August 2026. It handles both text-to-image and reference-image input, renders text and fine detail crisply down to roughly 10px, and improves across the board on the previous generation in text-image consistency, complex structure, and overall image quality. Aspect ratios are free-form, with an area ceiling of about 6.55MP (2560x2560) per image. Accessible via the OpenAI-compatible /v1/images/generations endpoint through Ofox. - Model ID: `qwen/qwen-image-3.0` - Provider: Qwen - Type: Image generation - Released: 2026-08-05 - Context window: 100,000 tokens / Max output: 100,000 tokens - Capabilities: vision - Pricing (as of 2026-10-03): Image output $0.03 per image - Endpoints: openai: /v1/images/generations - Protocols: openai - Model page: https://ofox.run/models/qwen/qwen-image-3.0 ```bash curl https://api.ofox.run/v1/images/generations \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "qwen/qwen-image-3.0", "prompt": "A watercolor fox reading a map", "n": 1}' ``` --- ## qwen/qwen-image-3.0-pro Qwen-Image 3.0 Pro — Qwen-Image 3.0 Pro is the flagship tier of Alibaba's Qwen-Image 3.0 series, officially released in August 2026. It handles both text-to-image and reference-image input, renders text and fine detail crisply down to roughly 10px, and leads the series on text-image consistency, complex structure, text rendering, and overall image quality — a strong fit for posters, typography, and layout-heavy scenes. Aspect ratios are free-form, with an area ceiling of about 6.55MP (2560x2560) per image. Accessible via the OpenAI-compatible /v1/images/generations endpoint through Ofox. - Model ID: `qwen/qwen-image-3.0-pro` - Provider: Qwen - Type: Image generation - Released: 2026-08-05 - Context window: 100,000 tokens / Max output: 100,000 tokens - Capabilities: vision - Pricing (as of 2026-10-03): Image output $0.07 per image - Endpoints: openai: /v1/images/generations - Protocols: openai - Model page: https://ofox.run/models/qwen/qwen-image-3.0-pro ```bash curl https://api.ofox.run/v1/images/generations \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "qwen/qwen-image-3.0-pro", "prompt": "A watercolor fox reading a map", "n": 1}' ``` --- ## qwen/qwen-max Qwen Max — Qwen Max is Alibaba's high-performance Qwen model, served through Dashscope for complex tasks that require sophisticated reasoning and generation. It supports tool use (function calling) and prompt caching. Context window: 32K tokens, output: 8K. Accessible via the OpenAI-compatible protocol. - Model ID: `qwen/qwen-max` - Provider: Qwen - Type: Chat / text generation - Released: 2025-01-25 - Context window: 32,000 tokens / Max output: 8,000 tokens - Capabilities: function calling, prompt caching - Pricing (as of 2026-10-03): Input $0.35 per 1M tokens; Output $1.38 per 1M tokens; Cache read $0.069 per 1M tokens - Endpoints: openai: /v1/chat/completions - Protocols: openai - Model page: https://ofox.run/models/qwen/qwen-max ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="qwen/qwen-max", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "qwen/qwen-max", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## qwen/qwen-plus Qwen Plus — Qwen Plus is Alibaba's balanced Qwen model, served through Dashscope and tuned to pair solid general-task performance with low cost. It brings strong Chinese and English bilingual capabilities and supports tool use (function calling). Context window: 1M tokens, output: 32K. Accessible via the OpenAI-compatible protocol. - Model ID: `qwen/qwen-plus` - Provider: Qwen - Type: Chat / text generation - Released: 2025-12-01 - Context window: 1,000,000 tokens / Max output: 32,000 tokens - Capabilities: function calling - Pricing (as of 2026-10-03): Input $0.12 per 1M tokens; Output $0.29 per 1M tokens; Cache read $0.023 per 1M tokens - Endpoints: openai: /v1/chat/completions - Protocols: openai - Model page: https://ofox.run/models/qwen/qwen-plus ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="qwen/qwen-plus", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "qwen/qwen-plus", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## qwen/qwen-turbo Qwen Turbo — Qwen Turbo is Alibaba's fast, cost-effective Qwen model, served through Dashscope for simple tasks that need quick responses, and optimized for Chinese language understanding. It supports tool use (function calling). Context window: 128K tokens, output: 16K. Accessible via the OpenAI-compatible protocol. - Model ID: `qwen/qwen-turbo` - Provider: Qwen - Type: Chat / text generation - Released: 2025-07-15 - Context window: 128,000 tokens / Max output: 16,000 tokens - Capabilities: function calling - Pricing (as of 2026-10-03): Input $0.043 per 1M tokens; Output $0.09 per 1M tokens; Cache read $0.0086 per 1M tokens - Endpoints: openai: /v1/chat/completions - Protocols: openai - Model page: https://ofox.run/models/qwen/qwen-turbo ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="qwen/qwen-turbo", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "qwen/qwen-turbo", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## qwen/qwen-vl-max Qwen VL Max — Qwen VL Max is Alibaba's top-tier vision-language model, served through Dashscope for multimodal work such as image understanding and visual reasoning. It accepts image input alongside text and supports tool use (function calling). Context window: 128K tokens, output: 8K. Accessible via the OpenAI-compatible protocol. - Model ID: `qwen/qwen-vl-max` - Provider: Qwen - Type: Chat / text generation - Released: 2025-08-13 - Context window: 128,000 tokens / Max output: 8,000 tokens - Capabilities: vision, function calling - Pricing (as of 2026-10-03): Input $0.23 per 1M tokens; Output $0.58 per 1M tokens; Cache read $0.023 per 1M tokens - Endpoints: openai: /v1/chat/completions - Protocols: openai - Model page: https://ofox.run/models/qwen/qwen-vl-max ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="qwen/qwen-vl-max", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "qwen/qwen-vl-max", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## qwen/qwen3-coder-flash Qwen3 Coder Flash — Qwen3-Coder-Flash is Alibaba's cost-efficient coding model built on the Qwen3 architecture. It inherits Qwen3-Coder-Plus's agentic coding capabilities, supports multi-turn tool interaction, and is optimized for repository-level code understanding with improved tool-call stability. Context window: 1M tokens, output: 64K. Available via OpenAI and Anthropic protocols. - Model ID: `qwen/qwen3-coder-flash` - Provider: Qwen - Type: Chat / text generation - Released: 2025-08-05 - Context window: 1,000,000 tokens / Max output: 64,000 tokens - Capabilities: function calling, reasoning, prompt caching - Pricing (as of 2026-10-03): Input $0.50 per 1M tokens; Output $2.50 per 1M tokens; Cache read $0.06 per 1M tokens; Cache write $0.27 per 1M tokens - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/qwen/qwen3-coder-flash ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="qwen/qwen3-coder-flash", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "qwen/qwen3-coder-flash", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## qwen/qwen3-coder-next Qwen3 Coder Next — Qwen3-Coder-Next is the next-generation coding model in the Qwen3 series, delivering performance close to Qwen3-Coder-Plus with better efficiency. It focuses on repository-level understanding, multi-turn tool interaction, and agentic coding workflows. Context window: 256K tokens, output: 64K. Available via OpenAI and Anthropic protocols. - Model ID: `qwen/qwen3-coder-next` - Provider: Qwen - Type: Chat / text generation - Released: 2026-02-19 - Context window: 256,000 tokens / Max output: 64,000 tokens - Capabilities: reasoning - Pricing (as of 2026-10-03): Input $0.20 per 1M tokens; Output $1.50 per 1M tokens - Endpoints: openai: /v1/chat/completions - Protocols: openai, anthropic - Model page: https://ofox.run/models/qwen/qwen3-coder-next ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="qwen/qwen3-coder-next", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "qwen/qwen3-coder-next", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## qwen/qwen3-coder-plus Qwen3 Coder Plus — Qwen3-Coder-Plus is Alibaba's most capable Qwen3-based coding model, featuring strong agentic coding abilities: autonomous programming, multi-turn tool use, and environment interaction. It excels at complex coding tasks while maintaining strong general capability. Context window: 1M tokens, output: 64K. Available via OpenAI and Anthropic protocols. - Model ID: `qwen/qwen3-coder-plus` - Provider: Qwen - Type: Chat / text generation - Released: 2025-09-23 - Context window: 1,000,000 tokens / Max output: 64,000 tokens - Capabilities: function calling, reasoning, prompt caching - Pricing (as of 2026-10-03): Input $1.80 per 1M tokens; Output $9.00 per 1M tokens; Cache read $0.20 per 1M tokens; Cache write $1.00 per 1M tokens - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: anthropic, openai - Model page: https://ofox.run/models/qwen/qwen3-coder-plus ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="qwen/qwen3-coder-plus", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "qwen/qwen3-coder-plus", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## qwen/qwen3-max Qwen3 Max — Qwen3 Max is Alibaba's latest flagship Qwen model, released on 2026-01-23 and served through Dashscope, with strong reasoning, long-context handling, and enhanced coding capabilities. It supports reasoning, tool use (function calling), and prompt caching. Context window: 256K tokens, output: 64K. Accessible via the OpenAI-compatible protocol. - Model ID: `qwen/qwen3-max` - Provider: Qwen - Type: Chat / text generation - Released: 2026-01-23 - Context window: 256,000 tokens / Max output: 64,000 tokens - Capabilities: function calling, reasoning, prompt caching - Pricing (as of 2026-10-03): Input $0.36 per 1M tokens; Output $1.43 per 1M tokens; Cache read $0.072 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai - Model page: https://ofox.run/models/qwen/qwen3-max ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="qwen/qwen3-max", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "qwen/qwen3-max", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## qwen/qwen3.5-122b-a10b Qwen: Qwen3.5 122B A10B — Qwen3.5-122B-A10B is a hybrid-architecture vision-language model from Alibaba's Qwen3.5 series, combining linear attention with sparse MoE for efficient inference. With 122B total parameters and 10B active, it ranks just below Qwen3.5-397B in overall performance. Context: 256K tokens, output: 64K. Accessible via the OpenAI-compatible protocol. - Model ID: `qwen/qwen3.5-122b-a10b` - Provider: Qwen - Type: Chat / text generation - Released: 2026-02-23 - Context window: 256,000 tokens / Max output: 64,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, video input - Pricing (as of 2026-10-03): Input $0.29 per 1M tokens; Output $2.29 per 1M tokens; Cache read $0.29 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions - Protocols: openai - Model page: https://ofox.run/models/qwen/qwen3.5-122b-a10b ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="qwen/qwen3.5-122b-a10b", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "qwen/qwen3.5-122b-a10b", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## qwen/qwen3.5-27b Qwen: Qwen3.5 27B — Qwen3.5-27B is a dense vision-language model from Alibaba's Qwen3.5 series, incorporating linear attention for fast inference. Its overall capability approaches Qwen3.5-122B-A10B, making it an efficient choice balancing speed and quality. Context: 256K tokens, output: 64K. Accessible via the OpenAI-compatible protocol. - Model ID: `qwen/qwen3.5-27b` - Provider: Qwen - Type: Chat / text generation - Released: 2026-02-23 - Context window: 256,000 tokens / Max output: 64,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, video input - Pricing (as of 2026-10-03): Input $0.29 per 1M tokens; Output $2.05 per 1M tokens; Cache read $0.29 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions - Protocols: openai - Model page: https://ofox.run/models/qwen/qwen3.5-27b ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="qwen/qwen3.5-27b", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "qwen/qwen3.5-27b", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## qwen/qwen3.5-35b-a3b Qwen: Qwen3.5 35B A3B — Qwen3.5-35B-A3B is a hybrid-architecture vision-language model from Alibaba's Qwen3.5 series, combining linear attention with sparse MoE. With only 3B active parameters, it delivers inference efficiency close to Qwen3.5-122B-A10B. Context: 256K tokens, output: 64K. Accessible via the OpenAI-compatible protocol. - Model ID: `qwen/qwen3.5-35b-a3b` - Provider: Qwen - Type: Chat / text generation - Released: 2026-02-23 - Context window: 256,000 tokens / Max output: 64,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, video input - Pricing (as of 2026-10-03): Input $0.29 per 1M tokens; Output $1.83 per 1M tokens; Cache read $0.29 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions - Protocols: openai - Model page: https://ofox.run/models/qwen/qwen3.5-35b-a3b ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="qwen/qwen3.5-35b-a3b", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "qwen/qwen3.5-35b-a3b", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## qwen/qwen3.5-397b-a17b Qwen: Qwen3.5 397B A17B — Qwen3.5-397B-A17B is the flagship model in Alibaba's Qwen3.5 series — the largest and most capable, featuring hybrid linear-attention and sparse MoE architecture. It leads the series in language understanding, logical reasoning, coding, and instruction following. Context: 256K tokens, output: 64K. Accessible via the OpenAI-compatible protocol. - Model ID: `qwen/qwen3.5-397b-a17b` - Provider: Qwen - Type: Chat / text generation - Released: 2026-02-23 - Context window: 256,000 tokens / Max output: 64,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, video input - Pricing (as of 2026-10-03): Input $0.55 per 1M tokens; Output $3.50 per 1M tokens; Cache read $0.55 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions - Protocols: openai - Model page: https://ofox.run/models/qwen/qwen3.5-397b-a17b ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="qwen/qwen3.5-397b-a17b", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "qwen/qwen3.5-397b-a17b", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## qwen/qwen3.5-flash Qwen: Qwen3.5 Flash — Qwen3.5-Flash is Alibaba's speed-optimized vision-language model in the Qwen3.5 series, using a hybrid linear-attention and sparse MoE design for high inference throughput. It shows significant improvements over Qwen3 Flash on both text and multimodal tasks. Context: 1M tokens, output: 64K. Available via OpenAI and Anthropic protocols. - Model ID: `qwen/qwen3.5-flash` - Provider: Qwen - Type: Chat / text generation - Released: 2026-02-23 - Context window: 1,000,000 tokens / Max output: 64,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, video input - Pricing (as of 2026-10-03): Input $0.10 per 1M tokens; Output $0.40 per 1M tokens; Cache read $0.01 per 1M tokens; Cache write $0.125 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/qwen/qwen3.5-flash ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="qwen/qwen3.5-flash", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "qwen/qwen3.5-flash", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## qwen/qwen3.5-plus Qwen3.5 Plus — Qwen3.5-Plus is Alibaba's balanced vision-language model in the Qwen3.5 series, built on a hybrid linear-attention and sparse MoE architecture. It consistently outperforms earlier Qwen3.5 models across benchmarks and delivers strong general-purpose performance at a competitive price point. Context: 1M tokens, output: 64K. Available via OpenAI and Anthropic protocols. - Model ID: `qwen/qwen3.5-plus` - Provider: Qwen - Type: Chat / text generation - Released: 2026-02-16 - Context window: 1,000,000 tokens / Max output: 64,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, video input - Pricing (as of 2026-10-03): Input $0.40 per 1M tokens; Output $2.40 per 1M tokens; Cache read $0.04 per 1M tokens; Cache write $0.40 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: anthropic, openai - Model page: https://ofox.run/models/qwen/qwen3.5-plus ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="qwen/qwen3.5-plus", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "qwen/qwen3.5-plus", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## qwen/qwen3.6-27b Qwen: Qwen3.6 27B — Qwen3.6-27B is a dense vision-language model from Alibaba's Qwen3.6 series, with notable improvements in agentic coding, STEM reasoning, and visual understanding over Qwen3.5-27B. Context: 256K tokens, output: 64K. Available via OpenAI and Anthropic protocols. - Model ID: `qwen/qwen3.6-27b` - Provider: Qwen - Type: Chat / text generation - Released: 2026-04-22 - Context window: 256,000 tokens / Max output: 64,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, video input - Pricing (as of 2026-10-03): Input $0.43 per 1M tokens; Output $2.57 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions - Protocols: openai, anthropic - Model page: https://ofox.run/models/qwen/qwen3.6-27b ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="qwen/qwen3.6-27b", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "qwen/qwen3.6-27b", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## qwen/qwen3.6-flash Qwen: Qwen3.6 Flash — Qwen3.6-Flash is Alibaba's fast vision-language model in the Qwen3.6 series, delivering significant improvements over Qwen3.5-Flash — especially in agentic coding, frontend development, and multimodal tasks. Context: 1M tokens, output: 64K. Available via OpenAI and Anthropic protocols. - Model ID: `qwen/qwen3.6-flash` - Provider: Qwen - Type: Chat / text generation - Released: 2026-04-16 - Context window: 1,000,000 tokens / Max output: 64,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, video input - Pricing (as of 2026-10-03): Input $0.25 per 1M tokens; Output $1.50 per 1M tokens; Cache read $0.025 per 1M tokens; Cache write $0.31 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/qwen/qwen3.6-flash ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="qwen/qwen3.6-flash", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "qwen/qwen3.6-flash", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## qwen/qwen3.6-max-preview Qwen3.6 Max Preview — Qwen3.6-Max-Preview is the largest and most capable model in Alibaba's Qwen3.6 series, currently available as a text-only preview. It builds on Qwen3-Max and Qwen3.6-Plus with further capability improvements. Context: 256K tokens, output: 64K. Available via OpenAI and Anthropic protocols. - Model ID: `qwen/qwen3.6-max-preview` - Provider: Qwen - Type: Chat / text generation - Released: 2026-04-20 - Context window: 256,000 tokens / Max output: 64,000 tokens - Capabilities: function calling, reasoning, prompt caching - Pricing (as of 2026-10-03): Input $2.00 per 1M tokens; Output $12.00 per 1M tokens; Cache read $0.20 per 1M tokens; Cache write $2.00 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions - Protocols: openai, anthropic - Model page: https://ofox.run/models/qwen/qwen3.6-max-preview ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="qwen/qwen3.6-max-preview", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "qwen/qwen3.6-max-preview", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## qwen/qwen3.6-plus Qwen: Qwen3.6 Plus — Qwen3.6-Plus is Alibaba's flagship vision-language model in the Qwen3.6 series, matching top frontier models in overall performance. It brings major improvements over Qwen3.5 in agentic coding, frontend development, visual understanding, and instruction following. Context: 1M tokens, output: 64K. Available via OpenAI and Anthropic protocols. - Model ID: `qwen/qwen3.6-plus` - Provider: Qwen - Type: Chat / text generation - Released: 2026-04-02 - Context window: 1,000,000 tokens / Max output: 64,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, video input - Pricing (as of 2026-10-03): Input $0.50 per 1M tokens; Output $3.00 per 1M tokens; Cache read $0.05 per 1M tokens; Cache write $0.625 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/qwen/qwen3.6-plus ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="qwen/qwen3.6-plus", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "qwen/qwen3.6-plus", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## qwen/qwen3.7-max Qwen: Qwen3.7 Max — Qwen3.7-Max is Alibaba Cloud Bailian's flagship reasoning model, featuring a 1M-token context window and 64K output. It supports deep reasoning (reasoning_content), excels at complex coding and long-horizon agentic tasks. Available via OpenAI and Anthropic protocols through Ofox. - Model ID: `qwen/qwen3.7-max` - Provider: Qwen - Type: Chat / text generation - Released: 2026-05-21 - Context window: 1,064,000 tokens / Max output: 64,000 tokens - Capabilities: function calling, reasoning, prompt caching - Pricing (as of 2026-10-03): Input $1.71 per 1M tokens; Output $5.14 per 1M tokens; Cache read $0.17 per 1M tokens; Cache write $2.14 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/qwen/qwen3.7-max ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="qwen/qwen3.7-max", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "qwen/qwen3.7-max", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## qwen/qwen3.7-plus Qwen3.7 Plus — Qwen3.7-Plus is Alibaba Cloud Bailian's balanced reasoning model, released on 2026-06-01 and positioned to deliver strong performance at a competitive price. It supports deep reasoning alongside vision, video input, tool use (function calling), prompt caching, and web search, so it can work across long documents, images, and clips and then act on what it finds through tools. Context window: 1M tokens, output: 64K. Available via OpenAI and Anthropic protocols through Ofox. - Model ID: `qwen/qwen3.7-plus` - Provider: Qwen - Type: Chat / text generation - Released: 2026-06-01 - Context window: 1,064,000 tokens / Max output: 64,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, video input - Pricing (as of 2026-10-03): Input $0.40 per 1M tokens; Output $1.60 per 1M tokens; Cache read $0.08 per 1M tokens; Cache write $0.50 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/qwen/qwen3.7-plus ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="qwen/qwen3.7-plus", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "qwen/qwen3.7-plus", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## qwen/qwen3.8-27b Qwen: Qwen3.8 27B — Qwen3.8 27B is the natively vision-language dense model in Alibaba Cloud Bailian's Qwen3.8 series, released on 2026-08-19. It handles image and video understanding and supports deep reasoning (reasoning_content) that can be toggled per request, and focuses its gains over Qwen3.6 27B on coding and office scenarios in both text and visual modalities. It also supports tool use, prompt caching, and web search. Context window: 1.13M tokens, output: 131K. As a dense model rather than a Mixture-of-Experts one, it offers more predictable latency for interactive multimodal workloads. Available via OpenAI and Anthropic protocols through Ofox. - Model ID: `qwen/qwen3.8-27b` - Provider: Qwen - Type: Chat / text generation - Released: 2026-08-19 - Context window: 1,131,072 tokens / Max output: 131,072 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, video input - Pricing (as of 2026-10-03): Input $0.50 per 1M tokens; Output $1.71 per 1M tokens; Cache read $0.043 per 1M tokens; Cache write $0.63 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: anthropic, openai - Model page: https://ofox.run/models/qwen/qwen3.8-27b ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="qwen/qwen3.8-27b", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "qwen/qwen3.8-27b", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## qwen/qwen3.8-flash Qwen: Qwen3.8 Flash — Qwen3.8 Flash is the lightweight, high-speed natively multimodal model in Alibaba Cloud Bailian's Qwen3.8 series, released on 2026-08-25. It handles image and video understanding and supports deep reasoning (reasoning_content) that can be toggled per request, targeting high-concurrency, low-cost workloads without giving up coding and office tasks. It also supports tool use, prompt caching, and web search. Context window: 1.13M tokens, output: 131K. Being able to switch reasoning off per call makes it practical to run cheap, fast turns and reserve thinking for the requests that need it. Available via OpenAI and Anthropic protocols through Ofox. - Model ID: `qwen/qwen3.8-flash` - Provider: Qwen - Type: Chat / text generation - Released: 2026-08-25 - Context window: 1,131,072 tokens / Max output: 131,072 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, video input - Pricing (as of 2026-10-03): Input $0.11 per 1M tokens; Output $0.39 per 1M tokens; Cache read $0.011 per 1M tokens; Cache write $0.14 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: anthropic, openai - Model page: https://ofox.run/models/qwen/qwen3.8-flash ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="qwen/qwen3.8-flash", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "qwen/qwen3.8-flash", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## qwen/qwen3.8-max Qwen: Qwen3.8 Max — Qwen3.8 Max is Alibaba Cloud Bailian's flagship model in the Qwen3.8 series, released on 2026-08-03. It supports deep reasoning (reasoning_content) with longer chains of thought, plus strong coding ability and long-horizon autonomous execution. Context window: 1M tokens, output: 131K. Available via OpenAI and Anthropic protocols through Ofox. - Model ID: `qwen/qwen3.8-max` - Provider: Qwen - Type: Chat / text generation - Released: 2026-08-03 - Context window: 1,131,072 tokens / Max output: 131,072 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, video input - Pricing (as of 2026-10-03): Input $1.71 per 1M tokens; Output $5.14 per 1M tokens; Cache read $0.17 per 1M tokens; Cache write $2.14 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/qwen/qwen3.8-max ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="qwen/qwen3.8-max", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "qwen/qwen3.8-max", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## qwen/qwen3.8-max-0902 Qwen: Qwen3.8 Max 0902 — Qwen3.8 Max 0902 is the 2026-09-02 snapshot upgrade of Alibaba Cloud Bailian's flagship Qwen3.8 Max. It notably strengthens coding and collaborative agent capability and improves visual understanding, while keeping the 1M-token context, deep reasoning (thinking mode), and image and video input of the base model. It also supports tool use, prompt caching, and web search. Context window: 1M tokens, output: 131K. Pinning this dated snapshot keeps behaviour stable for production workloads that cannot absorb changes to the rolling alias. Available via OpenAI and Anthropic protocols through Ofox. - Model ID: `qwen/qwen3.8-max-0902` - Provider: Qwen - Type: Chat / text generation - Released: 2026-09-02 - Context window: 1,000,000 tokens / Max output: 131,072 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, video input - Pricing (as of 2026-10-03): Input $1.71 per 1M tokens; Output $5.14 per 1M tokens; Cache read $0.17 per 1M tokens; Cache write $2.14 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/qwen/qwen3.8-max-0902 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="qwen/qwen3.8-max-0902", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "qwen/qwen3.8-max-0902", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## qwen/text-embedding-v4 Qwen Text Embedding V4 — Qwen Text Embedding V4 is Alibaba's latest text embedding model, served through Dashscope and producing 1024-dimensional vectors with improved accuracy. It is designed for semantic search and retrieval, RAG pipelines, clustering, classification, and similarity comparison. Input context: up to 8K tokens per request, which takes in long passages without aggressive chunking. Accessible via the OpenAI-compatible protocol. - Model ID: `qwen/text-embedding-v4` - Provider: Qwen - Type: Text embedding - Released: 2025-01-01 - Context window: 8,192 tokens / Max output: not published - Capabilities: none declared - Pricing (as of 2026-10-03): Input $0.072 per 1M tokens - Endpoints: openai: /v1/embeddings - Protocols: openai - Model page: https://ofox.run/models/qwen/text-embedding-v4 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") result = client.embeddings.create(model="qwen/text-embedding-v4", input="The quick brown fox") print(len(result.data[0].embedding)) ``` --- # Doubao Models Source: https://ofox.run/models/volcengine 11 Doubao models available through OfoxAI. All of them are called with the same API key and billed to the same balance as every other model on the platform. --- ## volcengine/doubao-seed-2.0-code Doubao Seed 2.0 Code — Doubao-Seed-2.0-Code is ByteDance's enterprise-grade coding model, extending Seed 2.0's strong agentic and vision-language capabilities with dedicated code enhancements — including advanced frontend skills and repository-level understanding. Context: 256K tokens, output: 128K. Available via OpenAI and Anthropic protocols. - Model ID: `volcengine/doubao-seed-2.0-code` - Provider: Doubao - Type: Chat / text generation - Released: 2026-02-14 - Context window: 256,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, video input - Pricing (as of 2026-10-03): Input $0.67 per 1M tokens; Output $3.36 per 1M tokens; Cache read $0.14 per 1M tokens; Cache write $0.0024 per 1M tokens - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: anthropic, openai - Model page: https://ofox.run/models/volcengine/doubao-seed-2.0-code ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="volcengine/doubao-seed-2.0-code", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "volcengine/doubao-seed-2.0-code", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## volcengine/doubao-seed-2.0-lite Doubao Seed 2.0 Lite — Doubao-Seed-2.0-Lite is ByteDance's balanced model for high-frequency enterprise workloads, offering strong performance beyond Doubao-Seed-1.8 at a competitive cost. It handles unstructured information processing, structured generation, and general-purpose chat. Context: 256K tokens, output: 32K. Available via OpenAI and Anthropic protocols. - Model ID: `volcengine/doubao-seed-2.0-lite` - Provider: Doubao - Type: Chat / text generation - Released: 2026-02-14 - Context window: 256,000 tokens / Max output: 32,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, video input - Pricing (as of 2026-10-03): Input $0.13 per 1M tokens; Output $0.76 per 1M tokens; Cache read $0.03 per 1M tokens; Cache write $0.0024 per 1M tokens - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: anthropic, openai - Model page: https://ofox.run/models/volcengine/doubao-seed-2.0-lite ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="volcengine/doubao-seed-2.0-lite", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "volcengine/doubao-seed-2.0-lite", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## volcengine/doubao-seed-2.0-mini Doubao Seed 2.0 Mini — Doubao-Seed-2.0-Mini is ByteDance's ultra-low-latency, high-concurrency model for cost-sensitive scenarios, with flexible reasoning deployment and performance comparable to Doubao-Seed-1.6. Context: 256K tokens, output: 32K. Available via OpenAI and Anthropic protocols. - Model ID: `volcengine/doubao-seed-2.0-mini` - Provider: Doubao - Type: Chat / text generation - Released: 2026-02-14 - Context window: 256,000 tokens / Max output: 32,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, video input - Pricing (as of 2026-10-03): Input $0.06 per 1M tokens; Output $0.56 per 1M tokens; Cache read $0.02 per 1M tokens; Cache write $0.0024 per 1M tokens - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: anthropic, openai - Model page: https://ofox.run/models/volcengine/doubao-seed-2.0-mini ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="volcengine/doubao-seed-2.0-mini", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "volcengine/doubao-seed-2.0-mini", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## volcengine/doubao-seed-2.0-pro Doubao Seed 2.0 Pro — Doubao-Seed-2.0-Pro is ByteDance's flagship general-purpose model for the agentic era, designed for complex reasoning and long-horizon task execution. It excels at multimodal understanding, long-context reasoning, structured generation, and tool use. Context: 256K tokens, output: 128K. Available via OpenAI and Anthropic protocols. - Model ID: `volcengine/doubao-seed-2.0-pro` - Provider: Doubao - Type: Chat / text generation - Released: 2026-02-14 - Context window: 256,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, video input - Pricing (as of 2026-10-03): Input $0.67 per 1M tokens; Output $3.36 per 1M tokens; Cache read $0.14 per 1M tokens; Cache write $0.0024 per 1M tokens - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: anthropic, openai - Model page: https://ofox.run/models/volcengine/doubao-seed-2.0-pro ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="volcengine/doubao-seed-2.0-pro", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "volcengine/doubao-seed-2.0-pro", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## volcengine/doubao-seed-2.1-pro Doubao Seed 2.1 Pro — Doubao-Seed-2.1-Pro is ByteDance Volcengine's flagship reasoning model with embedded thinking, delivering strong performance on complex tasks. Context: 256K tokens, output: 256K. Available via OpenAI and Anthropic protocols through Ofox. - Model ID: `volcengine/doubao-seed-2.1-pro` - Provider: Doubao - Type: Chat / text generation - Released: 2026-06-23 - Context window: 256,000 tokens / Max output: 256,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, video input - Pricing (as of 2026-10-03): Input $0.7072 per 1M tokens (list price $0.884); Output $3.536 per 1M tokens (list price $4.42); Cache read $0.1416 per 1M tokens (list price $0.177); Cache write $0.002 per 1M tokens (list price $0.0025); Web search $0.0012 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/volcengine/doubao-seed-2.1-pro ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="volcengine/doubao-seed-2.1-pro", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "volcengine/doubao-seed-2.1-pro", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## volcengine/doubao-seed-2.1-turbo Doubao Seed 2.1 Turbo — Doubao-Seed-2.1-Turbo is ByteDance Volcengine's high-speed reasoning model with embedded thinking, optimized for lower latency while retaining strong reasoning capability. Context: 256K tokens, output: 256K. Available via OpenAI and Anthropic protocols through Ofox. - Model ID: `volcengine/doubao-seed-2.1-turbo` - Provider: Doubao - Type: Chat / text generation - Released: 2026-06-23 - Context window: 256,000 tokens / Max output: 256,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, video input - Pricing (as of 2026-10-03): Input $0.3536 per 1M tokens (list price $0.442); Output $1.7696 per 1M tokens (list price $2.212); Cache read $0.068 per 1M tokens (list price $0.085); Cache write $0.0019 per 1M tokens (list price $0.0024); Web search $0.0012 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: anthropic, openai - Model page: https://ofox.run/models/volcengine/doubao-seed-2.1-turbo ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="volcengine/doubao-seed-2.1-turbo", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "volcengine/doubao-seed-2.1-turbo", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## volcengine/doubao-seed-character Doubao Seed Character — Doubao-Seed-Character is ByteDance Volcengine's role-play and character dialogue model, specialized for immersive conversational scenarios with stable persona adherence. Context: 256K tokens, output: 256K. Accessible via the OpenAI-compatible protocol through Ofox. - Model ID: `volcengine/doubao-seed-character` - Provider: Doubao - Type: Chat / text generation - Released: 2026-06-23 - Context window: 256,000 tokens / Max output: 256,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, video input - Pricing (as of 2026-10-03): Input $0.177 per 1M tokens; Output $0.884 per 1M tokens; Cache read $0.024 per 1M tokens; Cache write $0.0025 per 1M tokens; Web search $0.0012 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai - Model page: https://ofox.run/models/volcengine/doubao-seed-character ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="volcengine/doubao-seed-character", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "volcengine/doubao-seed-character", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## volcengine/doubao-seed-evolving Doubao Seed Evolving — Doubao-Seed-Evolving is ByteDance Volcengine's continuously updated reasoning model, tracking the latest-version weights with embedded thinking. It is designed for users who need access to the most recent improvements as they ship. Context: 256K tokens, output: 256K. Available via OpenAI and Anthropic protocols through Ofox. - Model ID: `volcengine/doubao-seed-evolving` - Provider: Doubao - Type: Chat / text generation - Released: 2026-06-23 - Context window: 256,000 tokens / Max output: 256,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, video input - Pricing (as of 2026-10-03): Input $0.884 per 1M tokens; Output $4.42 per 1M tokens; Cache read $0.177 per 1M tokens; Cache write $0.0025 per 1M tokens; Web search $0.0012 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: anthropic, openai - Model page: https://ofox.run/models/volcengine/doubao-seed-evolving ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="volcengine/doubao-seed-evolving", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "volcengine/doubao-seed-evolving", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## volcengine/doubao-seedream-4.5 Doubao Seedream 4.5 — Seedream 4.5 is ByteDance's multimodal image generation model, integrating text-to-image, image-to-image, and multi-image output with commonsense and reasoning capabilities. It delivers a significant visual quality improvement over Seedream 4.0, with strong instruction following and creative composition. Accessible via the OpenAI-compatible protocol. - Model ID: `volcengine/doubao-seedream-4.5` - Provider: Doubao - Type: Image generation - Released: 2025-11-28 - Context window: 100,000 tokens / Max output: 100,000 tokens - Capabilities: vision - Pricing (as of 2026-10-03): Image output $0.04 per image - Endpoints: openai: /v1/images/generations - Protocols: openai - Model page: https://ofox.run/models/volcengine/doubao-seedream-4.5 ```bash curl https://api.ofox.run/v1/images/generations \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "volcengine/doubao-seedream-4.5", "prompt": "A watercolor fox reading a map", "n": 1}' ``` --- ## volcengine/doubao-seedream-5.0-lite Doubao Seedream 5.0 Lite — Doubao-Seedream-5.0-Lite is ByteDance's latest image creation model, introducing real-time web search to incorporate current information into generated images. It features improved semantic comprehension and visual aesthetics at a lightweight scale. Accessible via the OpenAI-compatible protocol. - Model ID: `volcengine/doubao-seedream-5.0-lite` - Provider: Doubao - Type: Image generation - Released: 2026-01-28 - Context window: 100,000 tokens / Max output: 100,000 tokens - Capabilities: vision - Pricing (as of 2026-10-03): Image output $0.035 per image - Endpoints: openai: /v1/images/generations - Protocols: openai - Model page: https://ofox.run/models/volcengine/doubao-seedream-5.0-lite ```bash curl https://api.ofox.run/v1/images/generations \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "volcengine/doubao-seedream-5.0-lite", "prompt": "A watercolor fox reading a map", "n": 1}' ``` --- ## volcengine/doubao-seedream-5.0-pro Doubao Seedream 5.0 Pro — Doubao-Seedream-5.0-Pro is ByteDance's flagship image creation model released in July 2026, with comprehensive improvements over the previous generation in text-image alignment, structural coherence, text rendering, and visual aesthetics. Accessible via the OpenAI-compatible protocol. - Model ID: `volcengine/doubao-seedream-5.0-pro` - Provider: Doubao - Type: Image generation - Released: 2026-07-08 - Context window: 100,000 tokens / Max output: 100,000 tokens - Capabilities: vision - Pricing (as of 2026-10-03): Image output $0.05 per image - Endpoints: openai: /v1/images/generations - Protocols: openai - Model page: https://ofox.run/models/volcengine/doubao-seedream-5.0-pro ```bash curl https://api.ofox.run/v1/images/generations \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "volcengine/doubao-seedream-5.0-pro", "prompt": "A watercolor fox reading a map", "n": 1}' ``` --- # xAI Models Source: https://ofox.run/models/x-ai 6 xAI models available through OfoxAI. All of them are called with the same API key and billed to the same balance as every other model on the platform. --- ## x-ai/grok-4.1-fast xAI: Grok 4.1 Fast — Grok 4.1 Fast is xAI's speed-focused Grok 4.1 release, published on 2025-11-19, and delivers xAI's strongest agentic performance without explicit reasoning, together with improved tool calling. It supports vision (image input), tool use (function calling), prompt caching, and web search, making it a good fit for tool-driven assistants that need to move quickly. Context window: 2M tokens, output: 30K. Prompt caching keeps long-context agent loops affordable. Accessible via the OpenAI-compatible protocol through Ofox. - Model ID: `x-ai/grok-4.1-fast` - Provider: xAI - Type: Chat / text generation - Released: 2025-11-19 - Context window: 2,000,000 tokens / Max output: 30,000 tokens - Capabilities: vision, function calling, prompt caching, web search - Pricing (as of 2026-10-03): Input $0.20 per 1M tokens; Output $0.50 per 1M tokens; Cache read $0.05 per 1M tokens; Web search $0.014 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai - Model page: https://ofox.run/models/x-ai/grok-4.1-fast ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="x-ai/grok-4.1-fast", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "x-ai/grok-4.1-fast", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## x-ai/grok-4.20 xAI: Grok 4.20 — Grok 4.20 is xAI's newest flagship model, released on 2026-03-31, combining industry-leading speed with agentic tool calling. It pairs a very low hallucination rate with strict prompt adherence, delivering consistently precise and truthful responses. Capabilities include vision (image input), tool use (function calling), prompt caching, and web search, covering the full range of agentic and multimodal workloads. Context window: 2M tokens, output: 128K, wide enough to hold large document sets in a single session. Accessible via the OpenAI-compatible protocol through Ofox. - Model ID: `x-ai/grok-4.20` - Provider: xAI - Type: Chat / text generation - Released: 2026-03-31 - Context window: 2,000,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, prompt caching, web search - Pricing (as of 2026-10-03): Input $4.00 per 1M tokens; Output $12.00 per 1M tokens; Cache read $0.40 per 1M tokens; Web search $0.014 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai - Model page: https://ofox.run/models/x-ai/grok-4.20 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="x-ai/grok-4.20", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "x-ai/grok-4.20", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## x-ai/grok-4.3 xAI: Grok 4.3 — Grok 4.3 is a reasoning model from xAI, released on 2026-04-30. It accepts text and image inputs with text output and is suited to agentic workflows, instruction-following tasks, and applications that require high factual accuracy. Capabilities include reasoning, vision, tool use (function calling), and prompt caching, so multi-step tool loops can reuse a shared context cheaply. Context window: 1M tokens, output: 128K. Available via OpenAI and Anthropic protocols through Ofox. - Model ID: `x-ai/grok-4.3` - Provider: xAI - Type: Chat / text generation - Released: 2026-04-30 - Context window: 1,000,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching - Pricing (as of 2026-10-03): Input $1.25 per 1M tokens; Output $2.50 per 1M tokens; Cache read $0.20 per 1M tokens; Web search $0.014 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/x-ai/grok-4.3 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="x-ai/grok-4.3", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "x-ai/grok-4.3", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## x-ai/grok-4.5 xAI: Grok 4.5 — Grok 4.5 is xAI's frontier model, released on 2026-07-08, delivering strong performance on coding, knowledge work, and STEM tasks. It is a reasoning model with tool use (function calling) and prompt caching support, letting multi-step agent loops reuse a shared context cheaply. Context window: 500K tokens, output: 64K. Accessible via the OpenAI-compatible protocol. - Model ID: `x-ai/grok-4.5` - Provider: xAI - Type: Chat / text generation - Released: 2026-07-08 - Context window: 500,000 tokens / Max output: 65,536 tokens - Capabilities: function calling, reasoning, prompt caching - Pricing (as of 2026-10-03): Input $2.00 per 1M tokens; Output $6.00 per 1M tokens; Cache read $0.30 per 1M tokens - Endpoints: openai: /v1/chat/completions - Protocols: openai - Model page: https://ofox.run/models/x-ai/grok-4.5 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="x-ai/grok-4.5", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "x-ai/grok-4.5", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## x-ai/grok-4.6 xAI: Grok 4.6 — Grok 4.6 is xAI's most intelligent and fastest model to date, released on 2026-08-12, with frontier performance on coding, knowledge work, and STEM tasks. It is a reasoning model with tool use (function calling) and prompt caching support, well suited to fast-moving agentic workflows. Context window: 500K tokens, output: 64K. Accessible via the OpenAI-compatible protocol. - Model ID: `x-ai/grok-4.6` - Provider: xAI - Type: Chat / text generation - Released: 2026-08-12 - Context window: 500,000 tokens / Max output: 65,536 tokens - Capabilities: function calling, reasoning, prompt caching - Pricing (as of 2026-10-03): Input $2.00 per 1M tokens; Output $6.00 per 1M tokens; Cache read $0.50 per 1M tokens - Endpoints: openai: /v1/chat/completions - Protocols: openai - Model page: https://ofox.run/models/x-ai/grok-4.6 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="x-ai/grok-4.6", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "x-ai/grok-4.6", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## x-ai/grok-4.7 xAI: Grok 4.7 — Grok 4.7 is xAI's reasoning model succeeding Grok 4.6, intended for coding, agentic tool use, and general chat. It supports always-on reasoning, function calling, structured outputs, and prompt caching. Context window: 500K tokens, output: 65,536 tokens. Requests with at least 200K prompt tokens use the long-context billing rate. The Ofox catalog exposes text input and text output for this model. Accessible through Ofox via the OpenAI-compatible Chat Completions endpoint, with model capabilities and pricing shown in the live catalog. - Model ID: `x-ai/grok-4.7` - Provider: xAI - Type: Chat / text generation - Released: 2026-09-21 - Context window: 500,000 tokens / Max output: 65,536 tokens - Capabilities: function calling, reasoning, prompt caching - Pricing (as of 2026-10-03): Input $2.00 per 1M tokens; Output $6.00 per 1M tokens; Cache read $0.50 per 1M tokens - Endpoints: openai: /v1/chat/completions - Protocols: openai - Model page: https://ofox.run/models/x-ai/grok-4.7 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="x-ai/grok-4.7", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "x-ai/grok-4.7", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- # GLM Models Source: https://ofox.run/models/z-ai 11 GLM models available through OfoxAI. All of them are called with the same API key and billed to the same balance as every other model on the platform. --- ## z-ai/glm-4.6 Z.ai: GLM-4.6 — GLM-4.6 is Z.ai's GLM-series chat model released on 2025-09-30, offering strong instruction following and an excellent balance of performance and cost, with particular strength on Chinese-language tasks. It supports reasoning, function calling (tool use), prompt caching, and web search, so multi-turn sessions can reuse a shared context cheaply while still reaching out for current information. Context: 200K tokens, output: 128K. Available via OpenAI and Anthropic protocols. - Model ID: `z-ai/glm-4.6` - Provider: GLM - Type: Chat / text generation - Released: 2025-09-30 - Context window: 200,000 tokens / Max output: 128,000 tokens - Capabilities: function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $0.60 per 1M tokens; Output $2.20 per 1M tokens; Cache read $0.11 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions - Protocols: openai, anthropic - Model page: https://ofox.run/models/z-ai/glm-4.6 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="z-ai/glm-4.6", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "z-ai/glm-4.6", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## z-ai/glm-4.7 Z.ai: GLM 4.7 — GLM-4.7 is Z.ai's flagship GLM model released on 2025-12-23, built around advanced reasoning and best-in-class Chinese language understanding. It supports reasoning, function calling (tool use), prompt caching, and web search, letting it plan multi-step tasks and bring in external information when a request calls for it. Context: 200K tokens, output: 128K, leaving room for long documents and extended agent traces. Available via OpenAI and Anthropic protocols through Ofox. - Model ID: `z-ai/glm-4.7` - Provider: GLM - Type: Chat / text generation - Released: 2025-12-23 - Context window: 200,000 tokens / Max output: 128,000 tokens - Capabilities: function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $0.40 per 1M tokens; Output $2.20 per 1M tokens; Cache read $0.11 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions - Protocols: anthropic, openai - Model page: https://ofox.run/models/z-ai/glm-4.7 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="z-ai/glm-4.7", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "z-ai/glm-4.7", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## z-ai/glm-4.7-flashx Z.ai: GLM-4.7 FlashX — GLM-4.7 FlashX is Z.ai's fast, low-cost variant of the 30B-class GLM-4.7-Flash model, balancing performance with efficiency and further optimized for agentic coding: stronger coding capability, long-horizon task planning, and tool collaboration. It supports reasoning, function calling (tool use), prompt caching, and web search. Context: 200K tokens, output: 128K. It is priced aggressively for its capability class. Available via OpenAI and Anthropic protocols. - Model ID: `z-ai/glm-4.7-flashx` - Provider: GLM - Type: Chat / text generation - Released: 2026-01-19 - Context window: 200,000 tokens / Max output: 128,000 tokens - Capabilities: function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $0.072 per 1M tokens; Output $0.40 per 1M tokens; Cache read $0.01 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions - Protocols: openai, anthropic - Model page: https://ofox.run/models/z-ai/glm-4.7-flashx ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="z-ai/glm-4.7-flashx", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "z-ai/glm-4.7-flashx", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## z-ai/glm-5 Z.ai: GLM-5 — GLM-5 is Z.ai's flagship open-source foundation model, engineered for complex systems design and long-horizon agent workflows. Built for expert developers, it delivers production-grade performance on large-scale programming tasks and combines advanced agentic planning, deep backend reasoning, and iterative self-correction, moving beyond code generation toward full-system construction and autonomous execution. It supports reasoning, function calling (tool use), prompt caching, and web search. Context: 200K tokens, output: 128K. Available via OpenAI and Anthropic protocols through Ofox. - Model ID: `z-ai/glm-5` - Provider: GLM - Type: Chat / text generation - Released: 2026-02-11 - Context window: 200,000 tokens / Max output: 128,000 tokens - Capabilities: function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $1.00 per 1M tokens; Output $3.20 per 1M tokens; Cache read $0.20 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions - Protocols: openai, anthropic - Model page: https://ofox.run/models/z-ai/glm-5 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="z-ai/glm-5", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "z-ai/glm-5", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## z-ai/glm-5-turbo Z.ai: GLM-5-Turbo — GLM-5-Turbo is a Z.ai foundation model deeply optimized for the OpenClaw scenario. It has been tuned from the training phase onward for the core requirements of OpenClaw tasks, strengthening key capabilities such as tool invocation, command following, timed and persistent tasks, and long-chain execution. It supports reasoning, function calling (tool use), prompt caching, and web search. Context: 200K tokens, output: 128K. Available via OpenAI and Anthropic protocols. - Model ID: `z-ai/glm-5-turbo` - Provider: GLM - Type: Chat / text generation - Released: 2026-03-16 - Context window: 200,000 tokens / Max output: 128,000 tokens - Capabilities: function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $1.20 per 1M tokens; Output $4.00 per 1M tokens; Cache read $0.24 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions - Protocols: openai, anthropic - Model page: https://ofox.run/models/z-ai/glm-5-turbo ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="z-ai/glm-5-turbo", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "z-ai/glm-5-turbo", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## z-ai/glm-5.1 Z.ai: GLM 5.1 — GLM 5.1 is Z.ai's GLM-series chat model released on 2026-03-27, a newer iteration of the line aimed at general-purpose and agentic workloads. It supports reasoning, function calling (tool use), prompt caching, and web search, covering agentic loops that mix tool calls with retrieval. Context: 200K tokens, output: 128K. Available via OpenAI and Anthropic protocols through Ofox. - Model ID: `z-ai/glm-5.1` - Provider: GLM - Type: Chat / text generation - Released: 2026-03-27 - Context window: 200,000 tokens / Max output: 128,000 tokens - Capabilities: function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $1.40 per 1M tokens; Output $4.40 per 1M tokens; Cache read $0.26 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions - Protocols: anthropic, openai - Model page: https://ofox.run/models/z-ai/glm-5.1 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="z-ai/glm-5.1", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "z-ai/glm-5.1", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## z-ai/glm-5.2 Z.ai: GLM-5.2 — GLM-5.2 is the flagship reasoning model on Z.ai's international platform (api.z.ai), built on an open-weights MoE architecture with a 1M-token context window. It features embedded thinking (reasoning_content), prompt caching, tool calling, and web search, with strong coding and long-horizon agentic execution. Available via OpenAI-compatible and Anthropic protocols. - Model ID: `z-ai/glm-5.2` - Provider: GLM - Type: Chat / text generation - Released: 2026-06-16 - Context window: 1,048,576 tokens / Max output: 128,000 tokens - Capabilities: function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $1.40 per 1M tokens; Output $4.40 per 1M tokens; Cache read $0.26 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: openai, anthropic - Model page: https://ofox.run/models/z-ai/glm-5.2 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="z-ai/glm-5.2", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "z-ai/glm-5.2", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## z-ai/glm-5.3 Z.ai: GLM-5.3 — GLM-5.3 is Z.ai's flagship reasoning model on its international platform (api.z.ai), building on GLM-5.2 with stronger coding ability and a better performance-per-token balance for complex software engineering and long-horizon agent tasks. Thinking is always on, with low/high/max reasoning effort levels to tune cost against depth. It supports prompt caching, tool use (function calling), and web search. Context window: 1M tokens, output: 128K. Available via OpenAI and Anthropic protocols. - Model ID: `z-ai/glm-5.3` - Provider: GLM - Type: Chat / text generation - Released: 2026-08-19 - Context window: 1,048,576 tokens / Max output: 128,000 tokens - Capabilities: function calling, reasoning, prompt caching, web search - Pricing (as of 2026-10-03): Input $1.40 per 1M tokens; Output $4.40 per 1M tokens; Cache read $0.26 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: anthropic, openai - Model page: https://ofox.run/models/z-ai/glm-5.3 ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="z-ai/glm-5.3", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "z-ai/glm-5.3", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## z-ai/glm-5.3-flash Z.ai: GLM-5.3-Flash — GLM-5.3-Flash is the lightweight, high-speed model on Z.ai's international site (api.z.ai), released on 2026-08-26. It takes native multimodal input including images and video and targets efficient coding and long-horizon agent tasks. Thinking is always on and cannot be disabled, but the reasoning effort is selectable across low, high, and max. It supports tool use, prompt caching, and web search. Context window: 1M tokens, output: 131K. Available via OpenAI-compatible and Anthropic protocols; note that the OpenAI-compatible surface exposes chat completions and responses endpoints through Ofox. - Model ID: `z-ai/glm-5.3-flash` - Provider: GLM - Type: Chat / text generation - Released: 2026-08-26 - Context window: 1,048,576 tokens / Max output: 131,072 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, video input - Pricing (as of 2026-10-03): Input $0.15 per 1M tokens; Output $0.50 per 1M tokens; Cache read $0.03 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions, /v1/responses - Protocols: anthropic, openai - Model page: https://ofox.run/models/z-ai/glm-5.3-flash ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="z-ai/glm-5.3-flash", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "z-ai/glm-5.3-flash", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## z-ai/glm-5.3-flashx Z.ai: GLM-5.3-FlashX — GLM-5.3-FlashX is z.ai's lightweight, high-speed model on the international site (api.z.ai), built for efficient coding and long-horizon Agent tasks, with inference speeds of up to 200 tokens/s for a faster, smoother experience. It supports native multimodal input (images, video), 1M context, always-on thinking (which cannot be disabled, with low, high, and max reasoning effort levels), prompt caching, tool calling, and web search, and connects through both the OpenAI-compatible and Anthropic protocols (the Responses endpoint is not supported). - Model ID: `z-ai/glm-5.3-flashx` - Provider: GLM - Type: Chat / text generation - Released: 2026-08-26 - Context window: 1,048,576 tokens / Max output: 131,072 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, video input - Pricing (as of 2026-10-03): Input $0.15 per 1M tokens (list price $0.37); Output $0.50 per 1M tokens (list price $1.25); Cache read $0.03 per 1M tokens (list price $0.075); Web search $0.01 per request - Endpoints: openai: /v1/chat/completions - Protocols: anthropic, openai - Model page: https://ofox.run/models/z-ai/glm-5.3-flashx ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="z-ai/glm-5.3-flashx", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "z-ai/glm-5.3-flashx", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- ## z-ai/glm-5v-turbo GLM-5V-Turbo — GLM-5V-Turbo is Z.ai's first multimodal coding foundation model, built for vision-based coding tasks. It natively processes multimodal inputs such as images, video, PDFs, and text, while also excelling at long-horizon planning, complex coding, and action execution. Capabilities include vision, video input, PDF input, reasoning, function calling (tool use), prompt caching, and web search. Context: 200K tokens, output: 128K. Available via OpenAI and Anthropic protocols through Ofox. - Model ID: `z-ai/glm-5v-turbo` - Provider: GLM - Type: Chat / text generation - Released: 2026-04-01 - Context window: 200,000 tokens / Max output: 128,000 tokens - Capabilities: vision, function calling, reasoning, prompt caching, web search, video input, PDF input - Pricing (as of 2026-10-03): Input $1.20 per 1M tokens; Output $4.00 per 1M tokens; Cache read $0.24 per 1M tokens; Web search $0.01 per request - Endpoints: openai: /v1/chat/completions - Protocols: anthropic, openai - Model page: https://ofox.run/models/z-ai/glm-5v-turbo ```python from openai import OpenAI client = OpenAI(base_url="https://api.ofox.run/v1", api_key="") response = client.chat.completions.create( model="z-ai/glm-5v-turbo", messages=[{"role": "user", "content": "Hello!"}], ) print(response.choices[0].message.content) ``` ```bash curl https://api.ofox.run/v1/chat/completions \ -H "Authorization: Bearer " \ -H "Content-Type: application/json" \ -d '{"model": "z-ai/glm-5v-turbo", "messages": [{"role": "user", "content": "Hello!"}]}' ``` --- # Pricing Tables Source: https://ofox.run/pricing OfoxAI bills pay-as-you-go with no monthly platform fee and no per-seat licence: an idle month costs nothing. Models are billed in one of three units depending on what they produce: per token for text models, per image for some image models, per second for video models. The tables below restate the catalog prices in one place so they can be compared side by side. All figures were generated on 2026-10-03 — the live rate for any model is always the one published on its model page. Prices shown are the rate actually charged. Where a model is currently discounted below its list price, the list price is given in brackets next to it. ## Anthropic token pricing Token prices for Anthropic models on OfoxAI, in US dollars per 1 million tokens, as of 2026-10-03. | Model | Input $/1M | Output $/1M | Context window | |---|---|---|---| | `anthropic/claude-fable-5` | $10.00 | $50.00 | 1,000,000 | | `anthropic/claude-fable-5.1` | $10.00 | $50.00 | 1,000,000 | | `anthropic/claude-haiku-4.5` | $1.00 | $5.00 | 200,000 | | `anthropic/claude-opus-4.6` | $5.00 | $25.00 | 1,000,000 | | `anthropic/claude-opus-4.7` | $5.00 | $25.00 | 1,000,000 | | `anthropic/claude-opus-4.8` | $5.00 | $25.00 | 1,000,000 | | `anthropic/claude-opus-5` | $5.00 | $25.00 | 1,000,000 | | `anthropic/claude-opus-5.5` | $4.00 | $20.00 | 1,000,000 | | `anthropic/claude-sonnet-4.6` | $3.00 | $15.00 | 1,000,000 | | `anthropic/claude-sonnet-5` | $2.00 | $10.00 | 1,000,000 | | `anthropic/claude-sonnet-5.5` | $2.00 | $10.00 | 1,000,000 | ## DeepSeek token pricing Token prices for DeepSeek models on OfoxAI, in US dollars per 1 million tokens, as of 2026-10-03. | Model | Input $/1M | Output $/1M | Context window | |---|---|---|---| | `deepseek/deepseek-v3.2` | $0.29 | $0.43 | 128,000 | | `deepseek/deepseek-v4-flash-0423` | $0.19 | $0.51 | 1,000,000 | | `deepseek/deepseek-v4-flash-0731` | $0.308 (list $0.44) | $0.924 (list $1.32) | 1,000,000 | | `deepseek/deepseek-v4-pro-0423` | $1.32 | $3.96 | 1,000,000 | | `deepseek/deepseek-v4-pro-0813` | $0.924 (list $1.32) | $2.772 (list $3.96) | 1,000,000 | | `deepseek/deepseek-v4.1-flash` | $0.21 (list $0.30) | $0.84 (list $1.20) | 1,000,000 | ## Google token pricing Token prices for Google models on OfoxAI, in US dollars per 1 million tokens, as of 2026-10-03. | Model | Input $/1M | Output $/1M | Context window | |---|---|---|---| | `google/gemini-2.5-flash` | $0.30 | $2.50 | 1,048,576 | | `google/gemini-2.5-flash-lite` | $0.10 | $0.40 | 1,048,576 | | `google/gemini-2.5-pro` | $1.25 | $10.00 | 1,048,576 | | `google/gemini-3-flash-preview` | $0.50 | $3.00 | 1,048,576 | | `google/gemini-3.1-flash-lite` | $0.25 | $1.50 | 1,000,000 | | `google/gemini-3.1-pro-preview` | $2.00 | $12.00 | 1,048,576 | | `google/gemini-3.5-flash` | $1.50 | $9.00 | 1,000,000 | | `google/gemini-3.5-flash-lite` | $0.30 | $2.50 | 1,000,000 | | `google/gemini-3.6-flash` | $0.75 | $3.75 | 1,000,000 | | `google/gemini-3.7-flash` | $0.75 (list $1.50) | $3.75 (list $7.50) | 1,000,000 | | `google/gemini-3.8-flash` | $0.75 (list $1.50) | $3.75 (list $7.50) | 1,000,000 | | `google/gemini-embedding-2-preview` | $0.20 | $0.00 | 8,192 | ## MiniMax token pricing Token prices for MiniMax models on OfoxAI, in US dollars per 1 million tokens, as of 2026-10-03. | Model | Input $/1M | Output $/1M | Context window | |---|---|---|---| | `minimax/m2-her` | $0.30 | $1.20 | 200,000 | | `minimax/minimax-m2` | $0.30 | $1.20 | 204,800 | | `minimax/minimax-m2.1` | $0.30 | $1.20 | 204,800 | | `minimax/minimax-m2.1-lightning` | $0.30 | $2.40 | 204,800 | | `minimax/minimax-m2.5` | $0.30 | $1.20 | 200,000 | | `minimax/minimax-m2.5-lightning` | $0.30 | $2.40 | 200,000 | | `minimax/minimax-m2.7` | $0.30 | $1.20 | 200,000 | | `minimax/minimax-m2.7-highspeed` | $0.60 | $2.40 | 200,000 | | `minimax/minimax-m3` | $0.60 | $2.40 | 1,131,000 | ## Moonshot token pricing Token prices for Moonshot models on OfoxAI, in US dollars per 1 million tokens, as of 2026-10-03. | Model | Input $/1M | Output $/1M | Context window | |---|---|---|---| | `moonshotai/kimi-k2.6` | $0.95 | $4.00 | 262,144 | | `moonshotai/kimi-k2.7-code` | $0.95 | $4.00 | 262,144 | | `moonshotai/kimi-k2.7-code-highspeed` | $1.90 | $8.00 | 262,144 | | `moonshotai/kimi-k3` | $3.00 | $15.00 | 1,048,576 | ## OpenAI token pricing Token prices for OpenAI models on OfoxAI, in US dollars per 1 million tokens, as of 2026-10-03. | Model | Input $/1M | Output $/1M | Context window | |---|---|---|---| | `openai/gpt-4.1` | $1.60 (list $2.00) | $6.40 (list $8.00) | 1,047,576 | | `openai/gpt-4.1-mini` | $0.32 (list $0.40) | $1.28 (list $1.60) | 1,047,576 | | `openai/gpt-4o` | $2.00 (list $2.50) | $8.00 (list $10.00) | 128,000 | | `openai/gpt-4o-mini` | $0.12 (list $0.15) | $0.48 (list $0.60) | 128,000 | | `openai/gpt-4o-mini-transcribe` | $1.00 (list $1.25) | $4.00 (list $5.00) | 128,000 | | `openai/gpt-4o-transcribe-diarize` | $2.00 (list $2.50) | $8.00 (list $10.00) | 128,000 | | `openai/gpt-5` | $1.00 (list $1.25) | $8.00 (list $10.00) | 256,000 | | `openai/gpt-5-mini` | $0.20 (list $0.25) | $1.60 (list $2.00) | 256,000 | | `openai/gpt-5-nano` | $0.04 (list $0.05) | $0.32 (list $0.40) | 128,000 | | `openai/gpt-5.1` | $1.00 (list $1.25) | $8.00 (list $10.00) | 256,000 | | `openai/gpt-5.1-codex-max` | $1.00 (list $1.25) | $8.00 (list $10.00) | 256,000 | | `openai/gpt-5.1-codex-mini` | $0.20 (list $0.25) | $1.60 (list $2.00) | 256,000 | | `openai/gpt-5.2` | $1.40 (list $1.75) | $11.20 (list $14.00) | 512,000 | | `openai/gpt-5.2-codex` | $1.40 (list $1.75) | $11.20 (list $14.00) | 512,000 | | `openai/gpt-5.3-codex` | $1.40 (list $1.75) | $11.20 (list $14.00) | 512,000 | | `openai/gpt-5.4` | $2.00 (list $2.50) | $12.00 (list $15.00) | 1,050,000 | | `openai/gpt-5.4-mini` | $0.60 (list $0.75) | $3.60 (list $4.50) | 400,000 | | `openai/gpt-5.4-nano` | $0.16 (list $0.20) | $1.00 (list $1.25) | 400,000 | | `openai/gpt-5.4-pro` | $24.00 (list $30.00) | $144.00 (list $180.00) | 1,050,000 | | `openai/gpt-5.5` | $4.00 (list $5.00) | $24.00 (list $30.00) | 1,050,000 | | `openai/gpt-5.6-luna` | $0.20 (list $1.00) | $1.20 (list $6.00) | 1,050,000 | | `openai/gpt-5.6-sol` | $2.50 (list $5.00) | $15.00 (list $30.00) | 1,050,000 | | `openai/gpt-5.6-terra` | $2.00 (list $2.50) | $12.00 (list $15.00) | 1,050,000 | | `openai/gpt-6-astra` | $8.00 (list $10.00) | $40.00 (list $50.00) | 1,050,000 | | `openai/gpt-6-luna` | $0.08 (list $0.10) | $0.40 (list $0.50) | 1,050,000 | | `openai/gpt-6-sol` | $1.60 (list $2.00) | $8.00 (list $10.00) | 1,050,000 | | `openai/gpt-6.1-sol` | $2.00 | $10.00 | 1,050,000 | | `openai/text-embedding-3-large` | $0.14 | $0.00 | 8,200 | | `openai/text-embedding-3-small` | $0.02 | $0.00 | 8,200 | ## Qwen token pricing Token prices for Qwen models on OfoxAI, in US dollars per 1 million tokens, as of 2026-10-03. | Model | Input $/1M | Output $/1M | Context window | |---|---|---|---| | `qwen/qwen-flash` | $0.022 | $0.22 | 1,000,000 | | `qwen/qwen-max` | $0.35 | $1.38 | 32,000 | | `qwen/qwen-plus` | $0.12 | $0.29 | 1,000,000 | | `qwen/qwen-turbo` | $0.043 | $0.09 | 128,000 | | `qwen/qwen-vl-max` | $0.23 | $0.58 | 128,000 | | `qwen/qwen3-coder-flash` | $0.50 | $2.50 | 1,000,000 | | `qwen/qwen3-coder-next` | $0.20 | $1.50 | 256,000 | | `qwen/qwen3-coder-plus` | $1.80 | $9.00 | 1,000,000 | | `qwen/qwen3-max` | $0.36 | $1.43 | 256,000 | | `qwen/qwen3.5-122b-a10b` | $0.29 | $2.29 | 256,000 | | `qwen/qwen3.5-27b` | $0.29 | $2.05 | 256,000 | | `qwen/qwen3.5-35b-a3b` | $0.29 | $1.83 | 256,000 | | `qwen/qwen3.5-397b-a17b` | $0.55 | $3.50 | 256,000 | | `qwen/qwen3.5-flash` | $0.10 | $0.40 | 1,000,000 | | `qwen/qwen3.5-plus` | $0.40 | $2.40 | 1,000,000 | | `qwen/qwen3.6-27b` | $0.43 | $2.57 | 256,000 | | `qwen/qwen3.6-flash` | $0.25 | $1.50 | 1,000,000 | | `qwen/qwen3.6-max-preview` | $2.00 | $12.00 | 256,000 | | `qwen/qwen3.6-plus` | $0.50 | $3.00 | 1,000,000 | | `qwen/qwen3.7-max` | $1.71 | $5.14 | 1,064,000 | | `qwen/qwen3.7-plus` | $0.40 | $1.60 | 1,064,000 | | `qwen/qwen3.8-27b` | $0.50 | $1.71 | 1,131,072 | | `qwen/qwen3.8-flash` | $0.11 | $0.39 | 1,131,072 | | `qwen/qwen3.8-max` | $1.71 | $5.14 | 1,131,072 | | `qwen/qwen3.8-max-0902` | $1.71 | $5.14 | 1,000,000 | | `qwen/text-embedding-v4` | $0.072 | $0.00 | 8,192 | ## Doubao token pricing Token prices for Doubao models on OfoxAI, in US dollars per 1 million tokens, as of 2026-10-03. | Model | Input $/1M | Output $/1M | Context window | |---|---|---|---| | `volcengine/doubao-seed-2.0-code` | $0.67 | $3.36 | 256,000 | | `volcengine/doubao-seed-2.0-lite` | $0.13 | $0.76 | 256,000 | | `volcengine/doubao-seed-2.0-mini` | $0.06 | $0.56 | 256,000 | | `volcengine/doubao-seed-2.0-pro` | $0.67 | $3.36 | 256,000 | | `volcengine/doubao-seed-2.1-pro` | $0.7072 (list $0.884) | $3.536 (list $4.42) | 256,000 | | `volcengine/doubao-seed-2.1-turbo` | $0.3536 (list $0.442) | $1.7696 (list $2.212) | 256,000 | | `volcengine/doubao-seed-character` | $0.177 | $0.884 | 256,000 | | `volcengine/doubao-seed-evolving` | $0.884 | $4.42 | 256,000 | ## xAI token pricing Token prices for xAI models on OfoxAI, in US dollars per 1 million tokens, as of 2026-10-03. | Model | Input $/1M | Output $/1M | Context window | |---|---|---|---| | `x-ai/grok-4.1-fast` | $0.20 | $0.50 | 2,000,000 | | `x-ai/grok-4.20` | $4.00 | $12.00 | 2,000,000 | | `x-ai/grok-4.3` | $1.25 | $2.50 | 1,000,000 | | `x-ai/grok-4.5` | $2.00 | $6.00 | 500,000 | | `x-ai/grok-4.6` | $2.00 | $6.00 | 500,000 | | `x-ai/grok-4.7` | $2.00 | $6.00 | 500,000 | ## GLM token pricing Token prices for GLM models on OfoxAI, in US dollars per 1 million tokens, as of 2026-10-03. | Model | Input $/1M | Output $/1M | Context window | |---|---|---|---| | `z-ai/glm-4.6` | $0.60 | $2.20 | 200,000 | | `z-ai/glm-4.7` | $0.40 | $2.20 | 200,000 | | `z-ai/glm-4.7-flashx` | $0.072 | $0.40 | 200,000 | | `z-ai/glm-5` | $1.00 | $3.20 | 200,000 | | `z-ai/glm-5-turbo` | $1.20 | $4.00 | 200,000 | | `z-ai/glm-5.1` | $1.40 | $4.40 | 200,000 | | `z-ai/glm-5.2` | $1.40 | $4.40 | 1,048,576 | | `z-ai/glm-5.3` | $1.40 | $4.40 | 1,048,576 | | `z-ai/glm-5.3-flash` | $0.15 | $0.50 | 1,048,576 | | `z-ai/glm-5.3-flashx` | $0.15 (list $0.37) | $0.50 (list $1.25) | 1,048,576 | | `z-ai/glm-5v-turbo` | $1.20 | $4.00 | 200,000 | ## Image model pricing Image models on OfoxAI use two different billing units, and which one applies is a property of the model rather than a setting you choose. 5 of the 16 image models are billed per generated image; the other 11 are billed per token like a text model, with the generated image charged as output tokens. Prices are as of 2026-10-03. ### Billed per image For these models the token rates are zero by design; you are charged a flat amount per image returned. | Model | $ per image | |---|---| | `qwen/qwen-image-3.0` | $0.03 | | `qwen/qwen-image-3.0-pro` | $0.07 | | `volcengine/doubao-seedream-5.0-pro` | $0.05 | | `volcengine/doubao-seedream-5.0-lite` | $0.035 | | `volcengine/doubao-seedream-4.5` | $0.04 | ### Billed per token These models are billed like text models: prompt tokens at the input rate, and the generated image billed as image output tokens. All figures are US dollars per 1 million tokens. | Model | Input $/1M | Output $/1M | Image output $/1M | |---|---|---|---| | `openai/gpt-image-2.5-flare` | $4.00 (list $5.00) | $24.00 (list $30.00) | $30.00 | | `openai/gpt-image-2.5-sunburst` | $4.00 (list $5.00) | $24.00 (list $30.00) | $30.00 | | `google/gemini-3.1-flash-lite-image` | $0.25 | $1.50 | $30.00 | | `microsoft/mai-image-2.5-pro` | $5.00 | $106.00 | $106.00 | | `google/gemini-3.1-flash-image` | $0.50 | $3.00 | $60.00 | | `google/gemini-3-pro-image` | $2.00 | $12.00 | $120.00 | | `microsoft/mai-image-2.5` | $5.00 | $47.00 | $47.00 | | `microsoft/mai-image-2.5-flash` | $5.00 | $26.00 | $19.50 | | `openai/gpt-image-2` | $4.00 (list $5.00) | $24.00 (list $30.00) | $30.00 | | `openai/gpt-image-1.5` | $4.00 (list $5.00) | $8.00 (list $10.00) | $32.00 | | `google/gemini-2.5-flash-image` | $0.30 | $2.50 | $30.00 | ## Video model pricing Video models are billed per second of generated video rather than per token. The figure below is the entry price — the cheapest tier each model offers, as of 2026-10-03. Higher resolutions and video-to-video modes cost more; the full tier table for each model is in its Model Catalog entry above. | Model | From $/second | Resolutions | |---|---|---| | `alibaba/wan-3.0` | $0.054 | 480p, 720p, 1080p | | `alibaba/wan-3.0-prime` | $0.064 | 480p, 720p, 1080p | | `bytedance/seedance-2.5` | $0.11 | 480p, 720p, 1080p | | `minimax/hailuo-3` | $0.08 | 768p, 2k | | `minimax/hailuo-3-max` | $0.05 | 480p, 768p | | `alibaba/happyhorse-1.0` | $0.13 | 720p, 1080p | | `alibaba/happyhorse-1.1` | $0.13 | 720p, 1080p | | `bytedance/seedance-2.0-mini` | $0.02 (list $0.04) | 480p, 720p | | `bytedance/seedance-2.0` | $0.063 (list $0.07) | 480p, 720p, 1080p, 4k | | `bytedance/seedance-2.0-fast` | $0.042 (list $0.06) | 480p, 720p | | `alibaba/wan-2.7` | $0.10 | 720p, 1080p | | `alibaba/wan-2.6` | $0.10 | 720p, 1080p | --- # Integration Guides Source: https://ofox.run/docs/integrations Every tool below is configured the same way: point it at an OfoxAI base URL and give it an OfoxAI API key. Anthropic-protocol clients use `https://api.ofox.run/anthropic`; OpenAI-protocol clients use `https://api.ofox.run/v1`. Do not add a trailing slash to either base URL. Nothing about your prompts, your tools or your MCP servers changes. Model IDs must include the vendor prefix — `openai/gpt-6.1-sol`, not the bare model name after the slash. A request that omits the prefix will be rejected. ## Claude Code Source: https://ofox.run/docs/integrations/claude-code Claude Code speaks the Anthropic protocol, so it needs the Anthropic base URL `https://api.ofox.run/anthropic` with no trailing slash. Set two environment variables in your shell profile (`~/.zshrc` or `~/.bashrc`) and restart the terminal: ```bash export ANTHROPIC_BASE_URL=https://api.ofox.run/anthropic export ANTHROPIC_AUTH_TOKEN= ``` Run `claude` as usual afterwards, and confirm the setup with `/status` — it should report the Anthropic base URL as `https://api.ofox.run/anthropic`. Tool use, vision, streaming and extended thinking behave identically to the direct Anthropic API, because the protocol is native rather than translated. To pin a specific model, set `ANTHROPIC_MODEL` to a catalog ID such as `anthropic/claude-sonnet-5.5`. A GUI alternative is documented as well: the CC Switch desktop app writes `~/.claude/settings.json` for you, using the same request URL and the `ANTHROPIC_AUTH_TOKEN` auth field. ## Codex CLI Source: https://ofox.run/docs/integrations/codex Codex reads its configuration from `~/.codex/config.toml`. Register OfoxAI as an OpenAI-compatible provider and select it as the active model provider: ```toml model_provider = "ofoxai" model = "openai/gpt-6.1-sol" [model_providers.ofoxai] name = "OfoxAI" base_url = "https://api.ofox.run/v1" env_key = "OFOXAI_API_KEY" ``` `env_key` names the environment variable Codex reads the key from, so export it in your shell profile: ```bash export OFOXAI_API_KEY= ``` Verify with `codex "hello"`, and pick a model per invocation with `codex --model openai/gpt-6.1-sol "Refactor this function"`. One caveat: Codex CLI uses the OpenAI Responses API (`wire_api = "responses"`), and not every model supports that format — check the Endpoints line of a model's catalog entry for `/v1/responses` before selecting it. ## Cline Source: https://ofox.run/docs/integrations/cline Cline is a VS Code extension configured through its settings panel rather than a file. Install it from the VS Code extension marketplace, click the Cline icon in the activity bar, open Settings from the top right, then fill in the API configuration: - API Provider: OpenAI Compatible - Base URL: `https://api.ofox.run/v1` - OpenAI Compatible API Key: your OfoxAI API key - Model ID: any catalog ID, for example `openai/gpt-6.1-sol` Choose the OpenAI Compatible provider specifically. The dedicated Anthropic and Google Gemini provider entries in Cline do not accept custom model names, so they cannot address the full OfoxAI catalog. ## OpenCode Source: https://ofox.run/docs/integrations/opencode OpenCode can be configured either by environment variables or by its config file. The environment-variable route uses the OpenAI-compatible endpoint: ```bash export OPENAI_API_KEY= export OPENAI_BASE_URL=https://api.ofox.run/v1 ``` The Anthropic protocol works too, if you would rather drive Claude models natively: ```bash export ANTHROPIC_API_KEY= export ANTHROPIC_BASE_URL=https://api.ofox.run/anthropic ``` The config file lives at `~/.config/opencode/config.toml` and is TOML, not JSON: ```toml [providers.ofoxai] api_key = "" base_url = "https://api.ofox.run/v1" [models.default] provider = "ofoxai" model = "anthropic/claude-sonnet-5.5" ``` Verify with `opencode "Hello, how are you?"`. ## Zed Source: https://ofox.run/docs/integrations/zed Zed stores editor configuration in `~/.config/zed/settings.json`. Register OfoxAI under `language_models` as an OpenAI-compatible provider: ```json { "language_models": { "openai_compatible": { "OfoxAI": { "api_url": "https://api.ofox.run/v1", "available_models": [ { "name": "openai/gpt-6.1-sol", "display_name": "OpenAI: GPT-6.1 Sol", "max_tokens": 1050000, "max_output_tokens": 128000, "capabilities": { "tools": true, "images": true } } ] } } } } ``` Add one object per model you want in the picker. `name` and `max_tokens` are required; `display_name`, `max_output_tokens` and `capabilities` are optional. The API key is not stored in this file — Zed keeps it in the system keychain and prompts for it when you add the provider through the agent settings panel (`agent: open settings` from the command palette). ## CherryStudio Source: https://ofox.run/docs/integrations/cherry-studio CherryStudio is a desktop client configured entirely through its UI. Open Settings, choose Model Services in the sidebar, scroll to the bottom and click Add, then name the provider and pick its type. All three protocols are available: | Provider type | API address | |---|---| | OpenAI | `https://api.ofox.run/v1` | | Anthropic | `https://api.ofox.run/anthropic` | | Gemini | `https://api.ofox.run/gemini` | Paste the OfoxAI API key and the matching API address, then use the Manage button to fetch the model list and add the models you want in the picker. Finally, make sure the toggle in the top-right corner of the provider page is on — a configured provider that is toggled off does not appear in the model selector at all, which is the most common reason a correct configuration seems not to work. ## Any other tool Any client that exposes a custom base URL works without a dedicated guide. Use `https://api.ofox.run/v1` for OpenAI-compatible clients and `https://api.ofox.run/anthropic` for Anthropic-protocol clients, supply the OfoxAI API key, and enter a vendor-prefixed model ID from the Model Catalog section. This covers Chatbox, OpenClaw, Kilo, Aider, Gemini CLI and the OpenAI and Anthropic SDKs in every language they ship in. --- # Blog Index Source: https://ofox.run/blog The OfoxAI blog publishes integration walkthroughs, model comparisons and API tutorials. Below is a URL index of the 346 English articles currently published, derived from the blog sitemap; each label is reconstructed from the article's URL slug and is not necessarily its exact headline. Translations of many of these exist under the /zh/, /ja/ and /ru/ prefixes. - [30 Dollar AI Coding Stack Setup Guide 2026](https://ofox.run/blog/30-dollar-ai-coding-stack-setup-guide-2026/) - [429 Too Many Requests Rate Limit Exceeded When to Retry 2026](https://ofox.run/blog/429-too-many-requests-rate-limit-exceeded-when-to-retry-2026/) - [AI API Aggregation Access Every Model One Endpoint](https://ofox.run/blog/ai-api-aggregation-access-every-model-one-endpoint/) - [AI API Error Handling Troubleshooting Guide 2026](https://ofox.run/blog/ai-api-error-handling-troubleshooting-guide-2026/) - [AI API Pricing Comparison 2026](https://ofox.run/blog/ai-api-pricing-comparison-2026/) - [AI Ad Campaign Brief to Assets Ofox](https://ofox.run/blog/ai-ad-campaign-brief-to-assets-ofox/) - [AI Agent Development Python Guide 2026](https://ofox.run/blog/ai-agent-development-python-guide-2026/) - [AI Customer Feedback Classification Template](https://ofox.run/blog/ai-customer-feedback-classification-template/) - [AI Meeting Notes Action Items Template](https://ofox.run/blog/ai-meeting-notes-action-items-template/) - [AI Model Rankings May 2026](https://ofox.run/blog/ai-model-rankings-may-2026/) - [AI Product Photo Background](https://ofox.run/blog/ai-product-photo-background/) - [AI Tools API Configuration Guide 2026](https://ofox.run/blog/ai-tools-api-configuration-guide-2026/) - [AI Video Generation API Cost Per Usable Clip 2026](https://ofox.run/blog/ai-video-generation-api-cost-per-usable-clip-2026/) - [AI Video Generation APIs Sora Veo Kling Compared 2026](https://ofox.run/blog/ai-video-generation-apis-sora-veo-kling-compared-2026/) - [AI Weekly Report From Work Notes](https://ofox.run/blog/ai-weekly-report-from-work-notes/) - [AI Youtube Thumbnail Editable Text](https://ofox.run/blog/ai-youtube-thumbnail-editable-text/) - [Agentic Coding Claude Codex Gemini Cursor 2026](https://ofox.run/blog/agentic-coding-claude-codex-gemini-cursor-2026/) - [Apple Foundation Models 3 WWDC 2026 Developer Read](https://ofox.run/blog/apple-foundation-models-3-wwdc-2026-developer-read/) - [Best AI Coding Agent Harness Model Pairing 2026](https://ofox.run/blog/best-ai-coding-agent-harness-model-pairing-2026/) - [Best AI Model for Agents 2026](https://ofox.run/blog/best-ai-model-for-agents-2026/) - [Best AI Model for Coding 2026](https://ofox.run/blog/best-ai-model-for-coding-2026/) - [Best AI Model for OCR 2026](https://ofox.run/blog/best-ai-model-for-ocr-2026/) - [Best AI Models Complete Guide 2026](https://ofox.run/blog/best-ai-models-complete-guide-2026/) - [Best LLM API Providers 2026](https://ofox.run/blog/best-llm-api-providers-2026/) - [Best LLM Coding By Task Decision Matrix](https://ofox.run/blog/best-llm-coding-by-task-decision-matrix/) - [Best LLM for Coding Ranked Real Use 2026](https://ofox.run/blog/best-llm-for-coding-ranked-real-use-2026/) - [Cc Switch Install Multi CLI Setup 2026](https://ofox.run/blog/cc-switch-install-multi-cli-setup-2026/) - [Cc Switch vs Ofox Desktop](https://ofox.run/blog/cc-switch-vs-ofox-desktop/) - [ChatGPT Devicecheck Registration Failed Fix 2026](https://ofox.run/blog/chatgpt-devicecheck-registration-failed-fix-2026/) - [Cherry Studio API Configuration Guide 2026](https://ofox.run/blog/cherry-studio-api-configuration-guide-2026/) - [Claude API Error 529 Overloaded Fix 2026](https://ofox.run/blog/claude-api-error-529-overloaded-fix-2026/) - [Claude API Pricing Complete Breakdown 2026](https://ofox.run/blog/claude-api-pricing-complete-breakdown-2026/) - [Claude Agent SDK Migration Import Errors 2026](https://ofox.run/blog/claude-agent-sdk-migration-import-errors-2026/) - [Claude Code Artifact Schema 400 Fix](https://ofox.run/blog/claude-code-artifact-schema-400-fix/) - [Claude Code Cd Command Prompt Cache Edge Cases](https://ofox.run/blog/claude-code-cd-command-prompt-cache-edge-cases/) - [Claude Code Fallbackmodel 3 Tier Failover 2026](https://ofox.run/blog/claude-code-fallbackmodel-3-tier-failover-2026/) - [Claude Code Hooks Subagents Skills Complete Guide 2026](https://ofox.run/blog/claude-code-hooks-subagents-skills-complete-guide-2026/) - [Claude Code Hybrid Routing Pattern 2026](https://ofox.run/blog/claude-code-hybrid-routing-pattern-2026/) - [Claude Code Nested Subagents 2026](https://ofox.run/blog/claude-code-nested-subagents-2026/) - [Claude Code OfoxAI Configuration Guide 2026](https://ofox.run/blog/claude-code-ofoxai-configuration-guide-2026/) - [Claude Code Rate Limit Reached Error Fix 2026](https://ofox.run/blog/claude-code-rate-limit-reached-error-fix-2026/) - [Claude Code Safe Mode Guide 2026](https://ofox.run/blog/claude-code-safe-mode-guide-2026/) - [Claude Code Safety Prevent Accidental File Deletion](https://ofox.run/blog/claude-code-safety-prevent-accidental-file-deletion/) - [Claude Code Source Leak What 512k Lines Reveal 2026](https://ofox.run/blog/claude-code-source-leak-what-512k-lines-reveal-2026/) - [Claude Code Ssl Certificate Error 2026](https://ofox.run/blog/claude-code-ssl-certificate-error-2026/) - [Claude Code Switch Tutorial 2026](https://ofox.run/blog/claude-code-switch-tutorial-2026/) - [Claude Code Token Optimization 2026](https://ofox.run/blog/claude-code-token-optimization-2026/) - [Claude Code Usage Limit Hit Too Fast 2026](https://ofox.run/blog/claude-code-usage-limit-hit-too-fast-2026/) - [Claude Code vs Codex CLI vs Cursor vs DeepSeek Tui 2026](https://ofox.run/blog/claude-code-vs-codex-cli-vs-cursor-vs-deepseek-tui-2026/) - [Claude Fable 5 1 API Guide 2026](https://ofox.run/blog/claude-fable-5-1-api-guide-2026/) - [Claude Fable 5 1 vs Fable 5 vs Opus 5 2026](https://ofox.run/blog/claude-fable-5-1-vs-fable-5-vs-opus-5-2026/) - [Claude Fable 5 vs Opus 4 8 vs GPT 5 5 Swe Bench 2026](https://ofox.run/blog/claude-fable-5-vs-opus-4-8-vs-gpt-5-5-swe-bench-2026/) - [Claude Fable 5 vs Sonnet 5 2026](https://ofox.run/blog/claude-fable-5-vs-sonnet-5-2026/) - [Claude Go to Sleep Bug Explained 2026](https://ofox.run/blog/claude-go-to-sleep-bug-explained-2026/) - [Claude Haiku 4 API Budget Developer Guide 2026](https://ofox.run/blog/claude-haiku-4-api-budget-developer-guide-2026/) - [Claude Haiku 4 vs GPT 5 4 Mini Budget Models English 2026](https://ofox.run/blog/claude-haiku-4-vs-gpt-5-4-mini-budget-models-english-2026/) - [Claude Invalid Signature Thinking Block](https://ofox.run/blog/claude-invalid-signature-thinking-block/) - [Claude Max Throttling May 2026](https://ofox.run/blog/claude-max-throttling-may-2026/) - [Claude Opus 4 6 API Pricing Review 2026](https://ofox.run/blog/claude-opus-4-6-api-pricing-review-2026/) - [Claude Opus 4 6 vs GPT 5 5 vs Gemini 3 1 Pro Reasoning 2026](https://ofox.run/blog/claude-opus-4-6-vs-gpt-5-5-vs-gemini-3-1-pro-reasoning-2026/) - [Claude Opus 4 7 API Review Upgrade Guide 2026](https://ofox.run/blog/claude-opus-4-7-api-review-upgrade-guide-2026/) - [Claude Opus 4 7 Production Reliability Fix 2026](https://ofox.run/blog/claude-opus-4-7-production-reliability-fix-2026/) - [Claude Opus 4 8 Release Review 2026](https://ofox.run/blog/claude-opus-4-8-release-review-2026/) - [Claude Opus 5 5 Claude Code Access Limits](https://ofox.run/blog/claude-opus-5-5-claude-code-access-limits/) - [Claude Opus 5 5 Pricing Cache Fast Mode](https://ofox.run/blog/claude-opus-5-5-pricing-cache-fast-mode/) - [Claude Opus 5 5 Repository Review Workflow](https://ofox.run/blog/claude-opus-5-5-repository-review-workflow/) - [Claude Opus 5 5 vs GPT 6 Astra Complex Coding](https://ofox.run/blog/claude-opus-5-5-vs-gpt-6-astra-complex-coding/) - [Claude Opus 5 API Guide 2026](https://ofox.run/blog/claude-opus-5-api-guide-2026/) - [Claude Opus 5 Alternatives Switch or Pin 2026](https://ofox.run/blog/claude-opus-5-alternatives-switch-or-pin-2026/) - [Claude Opus 5 vs GPT 5 6 Sol 2026](https://ofox.run/blog/claude-opus-5-vs-gpt-5-6-sol-2026/) - [Claude Opus 5 vs Grok 4 6 Cost 2026](https://ofox.run/blog/claude-opus-5-vs-grok-4-6-cost-2026/) - [Claude Sonnet 5 5 API Migration 400](https://ofox.run/blog/claude-sonnet-5-5-api-migration-400/) - [Claude Sonnet 5 5 API Pricing Task Cost](https://ofox.run/blog/claude-sonnet-5-5-api-pricing-task-cost/) - [Claude Sonnet 5 5 Claude Code Setup](https://ofox.run/blog/claude-sonnet-5-5-claude-code-setup/) - [Claude Sonnet 5 5 Document Slide Spreadsheet Prompts](https://ofox.run/blog/claude-sonnet-5-5-document-slide-spreadsheet-prompts/) - [Claude Sonnet 5 5 Effort High Xhigh Max](https://ofox.run/blog/claude-sonnet-5-5-effort-high-xhigh-max/) - [Claude Sonnet 5 5 vs GPT 6 Sol Coding Cost](https://ofox.run/blog/claude-sonnet-5-5-vs-gpt-6-sol-coding-cost/) - [Claude Sonnet 5 5 vs Opus 5 5 Coding](https://ofox.run/blog/claude-sonnet-5-5-vs-opus-5-5-coding/) - [Claude Sonnet 5 Cline Setup 2026](https://ofox.run/blog/claude-sonnet-5-cline-setup-2026/) - [Claude Sonnet 5 vs 5 5 Upgrade](https://ofox.run/blog/claude-sonnet-5-vs-5-5-upgrade/) - [Claude Sonnet 5 vs Opus 4 8 2026](https://ofox.run/blog/claude-sonnet-5-vs-opus-4-8-2026/) - [Claude Tag Slack Setup Guide 2026](https://ofox.run/blog/claude-tag-slack-setup-guide-2026/) - [Claude Tool Use Missing Tool Result 400](https://ofox.run/blog/claude-tool-use-missing-tool-result-400/) - [Claude vs GPT vs Gemini Model Comparison Guide 2026](https://ofox.run/blog/claude-vs-gpt-vs-gemini-model-comparison-guide-2026/) - [Cline Vscode AI API Configuration Guide 2026](https://ofox.run/blog/cline-vscode-ai-api-configuration-guide-2026/) - [Codex Agents Md Not Loading Symlinked Workspaces 2026](https://ofox.run/blog/codex-agents-md-not-loading-symlinked-workspaces-2026/) - [Codex CLI 401 Unauthorized Fix 2026](https://ofox.run/blog/codex-cli-401-unauthorized-fix-2026/) - [Codex CLI API Configuration Guide 2026](https://ofox.run/blog/codex-cli-api-configuration-guide-2026/) - [Codex CLI Cc Switch Fable 5 1 Setup 2026](https://ofox.run/blog/codex-cli-cc-switch-fable-5-1-setup-2026/) - [Codex CLI Config Toml Deep Dive](https://ofox.run/blog/codex-cli-config-toml-deep-dive/) - [Codex CLI Corporate Proxy Pac Wpad](https://ofox.run/blog/codex-cli-corporate-proxy-pac-wpad/) - [Codex CLI Custom Model Providers Byo Setup](https://ofox.run/blog/codex-cli-custom-model-providers-byo-setup/) - [Codex CLI Real World Coding Workflow](https://ofox.run/blog/codex-cli-real-world-coding-workflow/) - [Codex Chrome Extension Codex App 2026](https://ofox.run/blog/codex-chrome-extension-codex-app-2026/) - [Codex Command Failed Retry Without Sandbox Fix 2026](https://ofox.run/blog/codex-command-failed-retry-without-sandbox-fix-2026/) - [Codex Command Not Found Fix npm Install 2026](https://ofox.run/blog/codex-command-not-found-fix-npm-install-2026/) - [Codex Computer Use Permissions Not Working](https://ofox.run/blog/codex-computer-use-permissions-not-working/) - [Codex Couldnt Load Its Resources Fix 2026](https://ofox.run/blog/codex-couldnt-load-its-resources-fix-2026/) - [Codex Desktop Not Showing Custom Models 2026](https://ofox.run/blog/codex-desktop-not-showing-custom-models-2026/) - [Codex Errors Fixes Index 2026](https://ofox.run/blog/codex-errors-fixes-index-2026/) - [Codex Failed to Start App Server Windows 2026](https://ofox.run/blog/codex-failed-to-start-app-server-windows-2026/) - [Codex GPT 5 5 Model Not Found 404 Fix 2026](https://ofox.run/blog/codex-gpt-5-5-model-not-found-404-fix-2026/) - [Codex Goal Mode Remote Computer Use 2026](https://ofox.run/blog/codex-goal-mode-remote-computer-use-2026/) - [Codex Keeps Testing Never Finishes](https://ofox.run/blog/codex-keeps-testing-never-finishes/) - [Codex Mobile App Iphone Android 2026](https://ofox.run/blog/codex-mobile-app-iphone-android-2026/) - [Codex Official Installation Complete](https://ofox.run/blog/codex-official-installation-complete/) - [Codex Stream Disconnected Before Completion](https://ofox.run/blog/codex-stream-disconnected-before-completion/) - [Codex Weekly Limit Cap Spend API 2026](https://ofox.run/blog/codex-weekly-limit-cap-spend-api-2026/) - [Codex Weekly Limit Drained 2026](https://ofox.run/blog/codex-weekly-limit-drained-2026/) - [Codex Windows Wsl Installation](https://ofox.run/blog/codex-windows-wsl-installation/) - [Computer Use API Offline Controller](https://ofox.run/blog/computer-use-api-offline-controller/) - [Computer Use Competitor Research Table](https://ofox.run/blog/computer-use-competitor-research-table/) - [Computer Use Landing Page Qa Lab](https://ofox.run/blog/computer-use-landing-page-qa-lab/) - [Cursor Claude Code Cline Custom API Setup 2026](https://ofox.run/blog/cursor-claude-code-cline-custom-api-setup-2026/) - [Cursor Composer 2 5 Setup Guide 2026](https://ofox.run/blog/cursor-composer-2-5-setup-guide-2026/) - [DeepSeek API Price Increase New Rates Peak Hours Cache Cost 2026](https://ofox.run/blog/deepseek-api-price-increase-new-rates-peak-hours-cache-cost-2026/) - [DeepSeek API Pricing Guide 2026](https://ofox.run/blog/deepseek-api-pricing-guide-2026/) - [DeepSeek Harness Dsh Setup Custom Providers 2026](https://ofox.run/blog/deepseek-harness-dsh-setup-custom-providers-2026/) - [DeepSeek Harness Dsh Version Updates Stability Production 2026](https://ofox.run/blog/deepseek-harness-dsh-version-updates-stability-production-2026/) - [DeepSeek R1 Reasoning API Production English 2026](https://ofox.run/blog/deepseek-r1-reasoning-api-production-english-2026/) - [DeepSeek V3 2 Prompt Caching Setup 2026](https://ofox.run/blog/deepseek-v3-2-prompt-caching-setup-2026/) - [DeepSeek V4 1 Flash API Setup](https://ofox.run/blog/deepseek-v4-1-flash-api-setup/) - [DeepSeek V4 1 Flash Claude Code Setup](https://ofox.run/blog/deepseek-v4-1-flash-claude-code-setup/) - [DeepSeek V4 1 Flash Codex Setup](https://ofox.run/blog/deepseek-v4-1-flash-codex-setup/) - [DeepSeek V4 1 Flash Different Results Clients](https://ofox.run/blog/deepseek-v4-1-flash-different-results-clients/) - [DeepSeek V4 1 Flash Preview](https://ofox.run/blog/deepseek-v4-1-flash-preview/) - [DeepSeek V4 1 Flash Total Active Parameters Memory](https://ofox.run/blog/deepseek-v4-1-flash-total-active-parameters-memory/) - [DeepSeek V4 1 Flash vs Gemini 3 8 Flash vs Qwen3 8 Flash](https://ofox.run/blog/deepseek-v4-1-flash-vs-gemini-3-8-flash-vs-qwen3-8-flash/) - [DeepSeek V4 Flash Free Zero Cost Paths 2026](https://ofox.run/blog/deepseek-v4-flash-free-zero-cost-paths-2026/) - [DeepSeek V4 Flash Pay Less Price Increase 2026](https://ofox.run/blog/deepseek-v4-flash-pay-less-price-increase-2026/) - [DeepSeek V4 Flash Vision Exp Image Tokens 2026](https://ofox.run/blog/deepseek-v4-flash-vision-exp-image-tokens-2026/) - [DeepSeek V4 Flash vs Gemini 3 6 Flash 2026](https://ofox.run/blog/deepseek-v4-flash-vs-gemini-3-6-flash-2026/) - [DeepSeek V4 Pro 0813 Price Weights Benchmarks API Access 2026](https://ofox.run/blog/deepseek-v4-pro-0813-price-weights-benchmarks-api-access-2026/) - [DeepSeek V4 Pro Real Cost Cache Miss Thinking 2026](https://ofox.run/blog/deepseek-v4-pro-real-cost-cache-miss-thinking-2026/) - [DeepSeek V4 Pro to V4 1 Flash Migration](https://ofox.run/blog/deepseek-v4-pro-to-v4-1-flash-migration/) - [DeepSeek V4 Pro vs Flash](https://ofox.run/blog/deepseek-v4-pro-vs-flash/) - [DeepSeek V4 Release Guide 2026](https://ofox.run/blog/deepseek-v4-release-guide-2026/) - [DeepSeek V4 in Claude Code Real Cost](https://ofox.run/blog/deepseek-v4-in-claude-code-real-cost/) - [Delegate Claude Code Tasks to Mistral Vibe Save Tokens 2026](https://ofox.run/blog/delegate-claude-code-tasks-to-mistral-vibe-save-tokens-2026/) - [Dify GPT Image 2 5 Setup](https://ofox.run/blog/dify-gpt-image-2-5-setup/) - [Discounted LLM API Pricing 2026](https://ofox.run/blog/discounted-llm-api-pricing-2026/) - [Doubao Seed 2 0 API Guide Bytedance Budget LLM 2026](https://ofox.run/blog/doubao-seed-2-0-api-guide-bytedance-budget-llm-2026/) - [Doubao Seed 2 1 API Pro Turbo No Volcano Signup 2026](https://ofox.run/blog/doubao-seed-2-1-api-pro-turbo-no-volcano-signup-2026/) - [Embedding API RAG Complete Guide 2026](https://ofox.run/blog/embedding-api-rag-complete-guide-2026/) - [Extract Text JSON CSV LLM API](https://ofox.run/blog/extract-text-json-csv-llm-api/) - [Faceless Youtube Channel Automation Video API 2026](https://ofox.run/blog/faceless-youtube-channel-automation-video-api-2026/) - [Fal AI Alternatives Video Generation API 2026](https://ofox.run/blog/fal-ai-alternatives-video-generation-api-2026/) - [Fal vs Replicate vs Ofox Video API Pricing 2026](https://ofox.run/blog/fal-vs-replicate-vs-ofox-video-api-pricing-2026/) - [Fal vs Wavespeed vs Atlascloud Video API 2026](https://ofox.run/blog/fal-vs-wavespeed-vs-atlascloud-video-api-2026/) - [Flux 2 Max Image API Developer English 2026](https://ofox.run/blog/flux-2-max-image-api-developer-english-2026/) - [Free LLM API Tiers Ranked Coding 2026](https://ofox.run/blog/free-llm-api-tiers-ranked-coding-2026/) - [Function Calling Tool Use Complete Guide 2026](https://ofox.run/blog/function-calling-tool-use-complete-guide-2026/) - [GPT 5 4 Pro API Flagship Deep Review 2026](https://ofox.run/blog/gpt-5-4-pro-api-flagship-deep-review-2026/) - [GPT 5 4 Pro API Guide Pricing Setup 2026](https://ofox.run/blog/gpt-5-4-pro-api-guide-pricing-setup-2026/) - [GPT 5 4 Pro vs Claude Sonnet Gemini Production API 2026](https://ofox.run/blog/gpt-5-4-pro-vs-claude-sonnet-gemini-production-api-2026/) - [GPT 5 5 API vs Claude Opus Gemini 3 1 Flagship 2026](https://ofox.run/blog/gpt-5-5-api-vs-claude-opus-gemini-3-1-flagship-2026/) - [GPT 5 5 Instant OpenAI Default Model Overview 2026](https://ofox.run/blog/gpt-5-5-instant-openai-default-model-overview-2026/) - [GPT 5 5 Release Guide 2026](https://ofox.run/blog/gpt-5-5-release-guide-2026/) - [GPT 5 6 Model Not Available Error Fix 2026](https://ofox.run/blog/gpt-5-6-model-not-available-error-fix-2026/) - [GPT 5 6 Sol Terra Luna Which Tier 2026](https://ofox.run/blog/gpt-5-6-sol-terra-luna-which-tier-2026/) - [GPT 5 6 Terra vs GPT 5 5 Coding Cost 2026](https://ofox.run/blog/gpt-5-6-terra-vs-gpt-5-5-coding-cost-2026/) - [GPT 6 Astra API Error Model Not Found Fix 2026](https://ofox.run/blog/gpt-6-astra-api-error-model-not-found-fix-2026/) - [GPT 6 Astra API Pricing 2026](https://ofox.run/blog/gpt-6-astra-api-pricing-2026/) - [GPT 6 Astra Coding Agent Setup 2026](https://ofox.run/blog/gpt-6-astra-coding-agent-setup-2026/) - [GPT 6 Astra Plus Not Showing Chat Work Codex 2026](https://ofox.run/blog/gpt-6-astra-plus-not-showing-chat-work-codex-2026/) - [GPT 6 Astra Review 2026](https://ofox.run/blog/gpt-6-astra-review-2026/) - [GPT 6 Astra vs Claude Fable 5 1 2026](https://ofox.run/blog/gpt-6-astra-vs-claude-fable-5-1-2026/) - [GPT 6 Astra vs GPT 5 6 Sol 2026](https://ofox.run/blog/gpt-6-astra-vs-gpt-5-6-sol-2026/) - [GPT 6 Luna API Pricing Batch Cost](https://ofox.run/blog/gpt-6-luna-api-pricing-batch-cost/) - [GPT 6 Luna Sol Agent Routing Thresholds](https://ofox.run/blog/gpt-6-luna-sol-agent-routing-thresholds/) - [GPT 6 Luna Structured Outputs Data Extraction](https://ofox.run/blog/gpt-6-luna-structured-outputs-data-extraction/) - [GPT 6 Luna vs Gemini 3 5 Flash Lite Extraction](https://ofox.run/blog/gpt-6-luna-vs-gemini-3-5-flash-lite-extraction/) - [GPT 6 Pro Codex Usage Limits Reset 2026](https://ofox.run/blog/gpt-6-pro-codex-usage-limits-reset-2026/) - [GPT 6 Sol API Pricing Cache Task Cost](https://ofox.run/blog/gpt-6-sol-api-pricing-cache-task-cost/) - [GPT 6 Sol Codex Subscription vs API Cost](https://ofox.run/blog/gpt-6-sol-codex-subscription-vs-api-cost/) - [GPT 6 Sol High vs Xhigh Coding](https://ofox.run/blog/gpt-6-sol-high-vs-xhigh-coding/) - [GPT 6 Sol Luna Access Codex ChatGPT Work](https://ofox.run/blog/gpt-6-sol-luna-access-codex-chatgpt-work/) - [GPT 6 Sol Luna Image Understanding Fix Retest](https://ofox.run/blog/gpt-6-sol-luna-image-understanding-fix-retest/) - [GPT 6 Sol Luna Opus 5 5 Creator Brief Workflow](https://ofox.run/blog/gpt-6-sol-luna-opus-5-5-creator-brief-workflow/) - [GPT 6 Sol Luna Tool Calling Responses Migration](https://ofox.run/blog/gpt-6-sol-luna-tool-calling-responses-migration/) - [GPT 6 Sol vs Claude Opus 5 5 Coding Cost](https://ofox.run/blog/gpt-6-sol-vs-claude-opus-5-5-coding-cost/) - [GPT 6 Sol vs Luna vs Astra Task Selection](https://ofox.run/blog/gpt-6-sol-vs-luna-vs-astra-task-selection/) - [GPT 61 Sol API Pricing Cache Cost](https://ofox.run/blog/gpt-61-sol-api-pricing-cache-cost/) - [GPT 61 Sol Codex Setup Access](https://ofox.run/blog/gpt-61-sol-codex-setup-access/) - [GPT 61 Sol Tool Calling Responses Fix](https://ofox.run/blog/gpt-61-sol-tool-calling-responses-fix/) - [GPT 61 Sol Upgrade Guide](https://ofox.run/blog/gpt-61-sol-upgrade-guide/) - [GPT 61 Sol vs Claude Sonnet 5 5](https://ofox.run/blog/gpt-61-sol-vs-claude-sonnet-5-5/) - [GPT 61 Sol vs GPT 6 Astra](https://ofox.run/blog/gpt-61-sol-vs-gpt-6-astra/) - [GPT Image 2 5 API Guide](https://ofox.run/blog/gpt-image-2-5-api-guide/) - [GPT Image 2 5 API Pricing](https://ofox.run/blog/gpt-image-2-5-api-pricing/) - [GPT Image 2 5 Flare vs Sunburst](https://ofox.run/blog/gpt-image-2-5-flare-vs-sunburst/) - [GPT Image 2 5 Node SDK Errors](https://ofox.run/blog/gpt-image-2-5-node-sdk-errors/) - [GPT Image 2 5 Size Guide](https://ofox.run/blog/gpt-image-2-5-size-guide/) - [GPT Image 2 5 Transparent Background](https://ofox.run/blog/gpt-image-2-5-transparent-background/) - [GPT Image 2 5 UI Mockup Workflow](https://ofox.run/blog/gpt-image-2-5-ui-mockup-workflow/) - [GPT Image 2 Generation Failures Fixes](https://ofox.run/blog/gpt-image-2-generation-failures-fixes/) - [GPT Image 2 Release Guide 2026](https://ofox.run/blog/gpt-image-2-release-guide-2026/) - [GPT Image 2 Slow Quality Parameter 2026](https://ofox.run/blog/gpt-image-2-slow-quality-parameter-2026/) - [GPT Image 2 to 2 5 Migration](https://ofox.run/blog/gpt-image-2-to-2-5-migration/) - [Gemini 2 5 Pro Missing AI Studio Access](https://ofox.run/blog/gemini-2-5-pro-missing-ai-studio-access/) - [Gemini 3 1 Flash Lite vs DeepSeek V4 Flash Budget Agents 2026](https://ofox.run/blog/gemini-3-1-flash-lite-vs-deepseek-v4-flash-budget-agents-2026/) - [Gemini 3 1 Pro API Pricing Performance Guide 2026](https://ofox.run/blog/gemini-3-1-pro-api-pricing-performance-guide-2026/) - [Gemini 3 1 Pro vs Claude Opus 4 6 Comparison 2026](https://ofox.run/blog/gemini-3-1-pro-vs-claude-opus-4-6-comparison-2026/) - [Gemini 3 5 Flash Coding Agents Guide 2026](https://ofox.run/blog/gemini-3-5-flash-coding-agents-guide-2026/) - [Gemini 3 5 Pro Release Date Expected Specs 2026](https://ofox.run/blog/gemini-3-5-pro-release-date-expected-specs-2026/) - [Gemini 3 7 Flash API Guide 2026](https://ofox.run/blog/gemini-3-7-flash-api-guide-2026/) - [Gemini 3 8 Flash API Pricing 2026](https://ofox.run/blog/gemini-3-8-flash-api-pricing-2026/) - [Gemini 3 8 Flash vs DeepSeek V4 Flash 2026](https://ofox.run/blog/gemini-3-8-flash-vs-deepseek-v4-flash-2026/) - [Gemini 3 8 Flash vs Gemini 3 7 Flash 2026](https://ofox.run/blog/gemini-3-8-flash-vs-gemini-3-7-flash-2026/) - [Gemini 3 8 TTS Speech Metadata Migration](https://ofox.run/blog/gemini-3-8-tts-speech-metadata-migration/) - [Gemini CLI API Configuration Guide 2026](https://ofox.run/blog/gemini-cli-api-configuration-guide-2026/) - [Gemini CLI Free Tier Shutdown Fix 2026](https://ofox.run/blog/gemini-cli-free-tier-shutdown-fix-2026/) - [Gemini Image 429 Free Tier Limit Zero](https://ofox.run/blog/gemini-image-429-free-tier-limit-zero/) - [Gemini Missing Thought Signature Tool Call](https://ofox.run/blog/gemini-missing-thought-signature-tool-call/) - [Github Copilot Byok Oai Compatible API Setup](https://ofox.run/blog/github-copilot-byok-oai-compatible-api-setup/) - [Glm 5 2 Access Guide 2026](https://ofox.run/blog/glm-5-2-access-guide-2026/) - [Glm 5 2 Free Zero Cost Paths 2026](https://ofox.run/blog/glm-5-2-free-zero-cost-paths-2026/) - [Glm 5 2 Run Locally Gguf 2026](https://ofox.run/blog/glm-5-2-run-locally-gguf-2026/) - [Glm 5 2 Self Host Vllm Hardware Cost 2026](https://ofox.run/blog/glm-5-2-self-host-vllm-hardware-cost-2026/) - [Glm 5 2 in Cline 2026](https://ofox.run/blog/glm-5-2-in-cline-2026/) - [Glm 5 2 vs GPT 5 5 Cost 2026](https://ofox.run/blog/glm-5-2-vs-gpt-5-5-cost-2026/) - [Glm 5 3 API Pricing Endpoints Reasoning Effort 2026](https://ofox.run/blog/glm-5-3-api-pricing-endpoints-reasoning-effort-2026/) - [Glm 5 3 Benchmarks Access 2026](https://ofox.run/blog/glm-5-3-benchmarks-access-2026/) - [Glm 5 3 Flash Three Parameter Counts 2026](https://ofox.run/blog/glm-5-3-flash-three-parameter-counts-2026/) - [Glm 5 3 Flash vs DeepSeek V4 Flash 2026](https://ofox.run/blog/glm-5-3-flash-vs-deepseek-v4-flash-2026/) - [Glm 5 API Pricing Pony Alpha Zhipu AI Guide 2026](https://ofox.run/blog/glm-5-api-pricing-pony-alpha-zhipu-ai-guide-2026/) - [Google Antigravity 2 Explained Gemini Desktop Agent Platform 2026](https://ofox.run/blog/google-antigravity-2-explained-gemini-desktop-agent-platform-2026/) - [Grok 4 6 API Pricing vs Grok 4 5 2026](https://ofox.run/blog/grok-4-6-api-pricing-vs-grok-4-5-2026/) - [Grok 4 7 API Setup Responses Migration](https://ofox.run/blog/grok-4-7-api-setup-responses-migration/) - [Grok 4 7 Access Cursor Build Web App](https://ofox.run/blog/grok-4-7-access-cursor-build-web-app/) - [Grok 4 7 Pricing vs 4 6 Task Cost](https://ofox.run/blog/grok-4-7-pricing-vs-4-6-task-cost/) - [Grok 4 7 vs Mimo 2 6 Pro Coding Cost](https://ofox.run/blog/grok-4-7-vs-mimo-2-6-pro-coding-cost/) - [Grok API Pricing Setup Access Guide 2026](https://ofox.run/blog/grok-api-pricing-setup-access-guide-2026/) - [Grok Build 300 Month vs API Cost 2026](https://ofox.run/blog/grok-build-300-month-vs-api-cost-2026/) - [Grok Imagine Image 2 0 API Pricing By Quality](https://ofox.run/blog/grok-imagine-image-2-0-api-pricing-by-quality/) - [Grok Imagine Image API Generate Your First Image](https://ofox.run/blog/grok-imagine-image-api-generate-your-first-image/) - [Grok Imagine Image Quality Retirement What to Change](https://ofox.run/blog/grok-imagine-image-quality-retirement-what-to-change/) - [Grok Imagine Video API Pricing By Resolution](https://ofox.run/blog/grok-imagine-video-api-pricing-by-resolution/) - [Hermes Agent Self Improving AI Complete](https://ofox.run/blog/hermes-agent-self-improving-ai-complete/) - [Hermes Agent vs Openclaw Migration](https://ofox.run/blog/hermes-agent-vs-openclaw-migration/) - [How to Choose AI Model Not the Biggest](https://ofox.run/blog/how-to-choose-ai-model-not-the-biggest/) - [How to Choose a Video Generation API By Use Case](https://ofox.run/blog/how-to-choose-a-video-generation-api-by-use-case/) - [How to Reduce AI API Costs 2026](https://ofox.run/blog/how-to-reduce-ai-api-costs-2026/) - [How to Use Kimi K3 2026](https://ofox.run/blog/how-to-use-kimi-k3-2026/) - [How to Use Seedance 2 0 Prompt Guide Fixes Free Access 2026](https://ofox.run/blog/how-to-use-seedance-2-0-prompt-guide-fixes-free-access-2026/) - [How to Use Seedance 2 5 Prompts Timestamps Cost 2026](https://ofox.run/blog/how-to-use-seedance-2-5-prompts-timestamps-cost-2026/) - [Image API Errors Troubleshooting Guide](https://ofox.run/blog/image-api-errors-troubleshooting-guide/) - [Is OpenRouter Reliable Honest Review 2026](https://ofox.run/blog/is-openrouter-reliable-honest-review-2026/) - [Is Seedance 2 0 Free Trial Credits Where to Get Access Cost 2026](https://ofox.run/blog/is-seedance-2-0-free-trial-credits-where-to-get-access-cost-2026/) - [Jev AI Guide Pricing Use Cases](https://ofox.run/blog/jev-ai-guide-pricing-use-cases/) - [Jev Benchmark Results Limitations](https://ofox.run/blog/jev-benchmark-results-limitations/) - [Jev Confidence Threshold Evaluation](https://ofox.run/blog/jev-confidence-threshold-evaluation/) - [Jev LLM Routing Cost Break Even](https://ofox.run/blog/jev-llm-routing-cost-break-even/) - [K2 Horizon 36b A4b Local Deployment Memory](https://ofox.run/blog/k2-horizon-36b-a4b-local-deployment-memory/) - [Kimi K2 5 API Pricing Access Guide 2026](https://ofox.run/blog/kimi-k2-5-api-pricing-access-guide-2026/) - [Kimi K2 6 Release Guide 2026](https://ofox.run/blog/kimi-k2-6-release-guide-2026/) - [Kimi K2 6 vs Claude Opus 4 6 Coding 2026](https://ofox.run/blog/kimi-k2-6-vs-claude-opus-4-6-coding-2026/) - [Kimi K2 7 Code Free Zero Cost Paths 2026](https://ofox.run/blog/kimi-k2-7-code-free-zero-cost-paths-2026/) - [Kimi K2 7 Code Token Cut Lower Bill 2026](https://ofox.run/blog/kimi-k2-7-code-token-cut-lower-bill-2026/) - [Kimi K2 7 Code vs Glm 5 2 Cost Per Run 2026](https://ofox.run/blog/kimi-k2-7-code-vs-glm-5-2-cost-per-run-2026/) - [Kimi K3 vs GPT 5 5 Opus 4 8 2026](https://ofox.run/blog/kimi-k3-vs-gpt-5-5-opus-4-8-2026/) - [Kling 2 6 Pro Video API English Complete 2026](https://ofox.run/blog/kling-2-6-pro-video-api-english-complete-2026/) - [LLM API Cache Hit Math Real Bills 2026](https://ofox.run/blog/llm-api-cache-hit-math-real-bills-2026/) - [LLM API Error Codes Reference 2026](https://ofox.run/blog/llm-api-error-codes-reference-2026/) - [LLM API Pricing Hidden Costs Index 2026](https://ofox.run/blog/llm-api-pricing-hidden-costs-index-2026/) - [LLM API Rate Limits Compared 2026](https://ofox.run/blog/llm-api-rate-limits-compared-2026/) - [LLM API Selection Decision Matrix Mid 2026 English](https://ofox.run/blog/llm-api-selection-decision-matrix-mid-2026-english/) - [LLM Benchmarks Text Extraction Summarization 2026](https://ofox.run/blog/llm-benchmarks-text-extraction-summarization-2026/) - [LLM Leaderboard Best AI Models Ranked 2026](https://ofox.run/blog/llm-leaderboard-best-ai-models-ranked-2026/) - [Landscape Photo to Portrait AI](https://ofox.run/blog/landscape-photo-to-portrait-ai/) - [Llama 4 API Access Complete English Guide 2026](https://ofox.run/blog/llama-4-api-access-complete-english-guide-2026/) - [Long Context LLM Benchmarks 200k Tokens 2026](https://ofox.run/blog/long-context-llm-benchmarks-200k-tokens-2026/) - [Migrate Claude Code to Codex 2026](https://ofox.run/blog/migrate-claude-code-to-codex-2026/) - [Mimo 2 6 API Python Opencode Setup](https://ofox.run/blog/mimo-2-6-api-python-opencode-setup/) - [Mimo 2 6 Distill Qwen 9b Local Vram](https://ofox.run/blog/mimo-2-6-distill-qwen-9b-local-vram/) - [Mimo 2 6 Pricing Pro Flash Ultraspeed](https://ofox.run/blog/mimo-2-6-pricing-pro-flash-ultraspeed/) - [Mimo 2 6 vs DeepSeek V4 1 Flash Coding](https://ofox.run/blog/mimo-2-6-vs-deepseek-v4-1-flash-coding/) - [Minimax API Guide M2 5 M2 7 Access Via Ofox 2026](https://ofox.run/blog/minimax-api-guide-m2-5-m2-7-access-via-ofox-2026/) - [Minimax H3 API Pricing 2026](https://ofox.run/blog/minimax-h3-api-pricing-2026/) - [Minimax M2 API Pricing Comparison 2026](https://ofox.run/blog/minimax-m2-api-pricing-comparison-2026/) - [Minimax M3 vs Claude Opus 4 8 Coding 2026](https://ofox.run/blog/minimax-m3-vs-claude-opus-4-8-coding-2026/) - [Minimax M3 vs GPT 5 5 Coding Benchmark 2026](https://ofox.run/blog/minimax-m3-vs-gpt-5-5-coding-benchmark-2026/) - [Multi Model Router One API 2026](https://ofox.run/blog/multi-model-router-one-api-2026/) - [Multimodal AI API Vision TTS Stt Guide 2026](https://ofox.run/blog/multimodal-ai-api-vision-tts-stt-guide-2026/) - [N8n Custom API Base Url 2026](https://ofox.run/blog/n8n-custom-api-base-url-2026/) - [Nano Banana 2 Token Billing Reasoning Tokens 2026](https://ofox.run/blog/nano-banana-2-token-billing-reasoning-tokens-2026/) - [Ofox Desktop Multi Tool Setup Guide](https://ofox.run/blog/ofox-desktop-multi-tool-setup-guide/) - [OpenAI API Model Not Found Errors Troubleshooting](https://ofox.run/blog/openai-api-model-not-found-errors-troubleshooting/) - [OpenAI Agents API Codex Harness Hosted Sandboxes](https://ofox.run/blog/openai-agents-api-codex-harness-hosted-sandboxes/) - [OpenAI SDK Migration to OfoxAI Guide 2026](https://ofox.run/blog/openai-sdk-migration-to-ofoxai-guide-2026/) - [OpenAI Workspace Agents Open Source Lark 2026](https://ofox.run/blog/openai-workspace-agents-open-source-lark-2026/) - [OpenRouter Alternatives 2026](https://ofox.run/blog/openrouter-alternatives-2026/) - [OpenRouter Kimi K3 429 Rate Limit Failover 2026](https://ofox.run/blog/openrouter-kimi-k3-429-rate-limit-failover-2026/) - [OpenRouter Pricing Hidden Markup Breakdown 2026](https://ofox.run/blog/openrouter-pricing-hidden-markup-breakdown-2026/) - [OpenRouter Request Timeout Troubleshooting](https://ofox.run/blog/openrouter-request-timeout-troubleshooting/) - [Openclaw API Model Configuration Guide 2026](https://ofox.run/blog/openclaw-api-model-configuration-guide-2026/) - [Opencode API Configuration Guide 2026](https://ofox.run/blog/opencode-api-configuration-guide-2026/) - [Opencode API Provider Comparison](https://ofox.run/blog/opencode-api-provider-comparison/) - [Opencode vs Codex CLI Terminal Coding Agent 2026](https://ofox.run/blog/opencode-vs-codex-cli-terminal-coding-agent-2026/) - [Opus 5 5 Product Demo Video From Screenshots](https://ofox.run/blog/opus-5-5-product-demo-video-from-screenshots/) - [Opus 5 5 Video Landscape to Vertical](https://ofox.run/blog/opus-5-5-video-landscape-to-vertical/) - [Opus 5 5 Video Prompts Mp4 Guide](https://ofox.run/blog/opus-5-5-video-prompts-mp4-guide/) - [Opus 5 5 Video Voiceover Subtitles Sync](https://ofox.run/blog/opus-5-5-video-voiceover-subtitles-sync/) - [Pi Coding Agent Custom Provider Setup 2026](https://ofox.run/blog/pi-coding-agent-custom-provider-setup-2026/) - [Product Photo Short Video Storyboard](https://ofox.run/blog/product-photo-short-video-storyboard/) - [Prompt Caching Cost Math Anthropic vs OpenAI 2026](https://ofox.run/blog/prompt-caching-cost-math-anthropic-vs-openai-2026/) - [Qwen 3 6 27b vs Claude Opus 4 6 Coding 2026](https://ofox.run/blog/qwen-3-6-27b-vs-claude-opus-4-6-coding-2026/) - [Qwen 3 6 Plus API Complete Guide 2026](https://ofox.run/blog/qwen-3-6-plus-api-complete-guide-2026/) - [Qwen 3 6 Plus vs DeepSeek V4 Pro Coding 2026](https://ofox.run/blog/qwen-3-6-plus-vs-deepseek-v4-pro-coding-2026/) - [Qwen 3 7 Max Coding Arena Rank 4 vs Claude Opus 2026](https://ofox.run/blog/qwen-3-7-max-coding-arena-rank-4-vs-claude-opus-2026/) - [Qwen 3 7 Max vs Kimi K3 vs DeepSeek V4 2026](https://ofox.run/blog/qwen-3-7-max-vs-kimi-k3-vs-deepseek-v4-2026/) - [Qwen 3 7 Plus vs Qwen 3 7 Max Real Benchmark 2026](https://ofox.run/blog/qwen-3-7-plus-vs-qwen-3-7-max-real-benchmark-2026/) - [Qwen 3 8 27b Run Locally Vram Gguf 2026](https://ofox.run/blog/qwen-3-8-27b-run-locally-vram-gguf-2026/) - [Qwen 3 8 Flash API Pricing Guide 2026](https://ofox.run/blog/qwen-3-8-flash-api-pricing-guide-2026/) - [Qwen 3 8 Max 0902 What Changed 2026](https://ofox.run/blog/qwen-3-8-max-0902-what-changed-2026/) - [Qwen 3 8 Max Codex CLI Config Cost 2026](https://ofox.run/blog/qwen-3-8-max-codex-cli-config-cost-2026/) - [Qwen 3 8 Max Price Context Window API Access Open Weights 2026](https://ofox.run/blog/qwen-3-8-max-price-context-window-api-access-open-weights-2026/) - [Qwen 3 8 Max vs DeepSeek V4 Flash 2026](https://ofox.run/blog/qwen-3-8-max-vs-deepseek-v4-flash-2026/) - [Qwen Image 3 0 Pro Free API 2026](https://ofox.run/blog/qwen-image-3-0-pro-free-api-2026/) - [Qwen3 7 Max Developer Guide 2026](https://ofox.run/blog/qwen3-7-max-developer-guide-2026/) - [Qwen3 API Guide Access Qwen3 Max Coder Via Ofox 2026](https://ofox.run/blog/qwen3-api-guide-access-qwen3-max-coder-via-ofox-2026/) - [Repurpose Webinar Blog Email Social AI](https://ofox.run/blog/repurpose-webinar-blog-email-social-ai/) - [Seedance 2 0 Fast Mini Tier Comparison 2026](https://ofox.run/blog/seedance-2-0-fast-mini-tier-comparison-2026/) - [Seedance 2 0 Video API Access 2026](https://ofox.run/blog/seedance-2-0-video-api-access-2026/) - [Seedance 2 0 vs Wan Video API 2026](https://ofox.run/blog/seedance-2-0-vs-wan-video-api-2026/) - [Seedance 2 5 First Last Frame Image to Video 2026](https://ofox.run/blog/seedance-2-5-first-last-frame-image-to-video-2026/) - [Seedance 2 5 Price Audit Bytedance Formula 2026](https://ofox.run/blog/seedance-2-5-price-audit-bytedance-formula-2026/) - [Seedance 2 5 vs 2 0 What Changed 2026](https://ofox.run/blog/seedance-2-5-vs-2-0-what-changed-2026/) - [Seedance API Guide Tiers Pricing 2026](https://ofox.run/blog/seedance-api-guide-tiers-pricing-2026/) - [Seedream 4 5 Doubao Image API English 2026](https://ofox.run/blog/seedream-4-5-doubao-image-api-english-2026/) - [Sora 2 Pro API Developer Complete Guide 2026](https://ofox.run/blog/sora-2-pro-api-developer-complete-guide-2026/) - [Text Embedding Models Compared 2026](https://ofox.run/blog/text-embedding-models-compared-2026/) - [Translate CSV Preserve Sku Glossary](https://ofox.run/blog/translate-csv-preserve-sku-glossary/) - [Translate Image Text Keep Layout](https://ofox.run/blog/translate-image-text-keep-layout/) - [Transparent Background Not Supported for This Model Fix 2026](https://ofox.run/blog/transparent-background-not-supported-for-this-model-fix-2026/) - [Veo 3 1 Google Video API English Tutorial 2026](https://ofox.run/blog/veo-3-1-google-video-api-english-tutorial-2026/) - [Verify LLM API Provider Model Routing Billing](https://ofox.run/blog/verify-llm-api-provider-model-routing-billing/) - [Video Generation API Polling 202 Wait Times 2026](https://ofox.run/blog/video-generation-api-polling-202-wait-times-2026/) - [Wan 3 0 API Error Fix 2026](https://ofox.run/blog/wan-3-0-api-error-fix-2026/) - [Wan 3 0 API Pricing 2026](https://ofox.run/blog/wan-3-0-api-pricing-2026/) - [Wan 3 0 vs Seedance 2 5 2026](https://ofox.run/blog/wan-3-0-vs-seedance-2-5-2026/) - [Wan 3 0 vs Wan 2 7 2026](https://ofox.run/blog/wan-3-0-vs-wan-2-7-2026/) - [What Is a Context Window Token Limits By Model 2026](https://ofox.run/blog/what-is-a-context-window-token-limits-by-model-2026/) - [Why LLM API Gateway How to Choose 2026](https://ofox.run/blog/why-llm-api-gateway-how-to-choose-2026/) - [Xiaomi Mimo 2 6 Access Opencode Desktop](https://ofox.run/blog/xiaomi-mimo-2-6-access-opencode-desktop/) - [Zed Editor AI Configuration Guide 2026](https://ofox.run/blog/zed-editor-ai-configuration-guide-2026/) --- # FAQ Source: https://ofox.run ## What is OfoxAI? OfoxAI is a unified AI gateway that gives you 151 models — across the GPT, Claude, Gemini, DeepSeek, Kimi, Qwen and GLM families, plus video and image generation models — through one API key and one balance. It speaks the OpenAI, Anthropic and Gemini protocols natively, bills pay-as-you-go with no platform fee, and publishes a 99.9% SLA on the platform layer it operates. ## How do I start using the OfoxAI API? Create an account, generate an API key in the console, and replace your existing OpenAI, Anthropic or Google base URL with the matching OfoxAI endpoint. For OpenAI-compatible requests that is `https://api.ofox.run/v1`. Your existing code keeps working without an SDK change, because the request and response formats are the ones your SDK already sends and parses. ## Which protocols does OfoxAI support? Three, natively: OpenAI at `https://api.ofox.run/v1`, Anthropic at `https://api.ofox.run/anthropic`, and Gemini at `https://api.ofox.run/gemini`. Native means the gateway speaks each protocol's own wire format rather than translating everything into one canonical shape, so protocol-specific features — Anthropic extended thinking, Gemini multimodal input, OpenAI structured output — behave as documented by the SDK vendor. ## Do I need a different API key for each model or protocol? No. One OfoxAI key authenticates every model and every protocol, and all usage lands in a single balance and a single billing history. The same credential can drive a terminal coding agent, an editor plugin and a production service at the same time. ## How much does OfoxAI cost? There is no markup on the model price and no monthly platform fee. Text models are billed per token, video models per second of generated video, and image models either per generated image or per token, depending on the model. Every model's rate is published in the Model Catalog and Pricing Tables sections of this document and on its model page, before you send a request. Where a model is currently discounted, the rate charged is the discounted one and the list price is shown alongside it. An account that sends no traffic is billed nothing. ## Which models are cheapest? The lowest input price in the current catalog is `qwen/qwen-flash` at $0.022 per 1M input tokens. Every model's rate is listed in the Pricing Tables section of this document, sorted by vendor, so the cheapest option for a given task can be read off directly rather than estimated. Because billing is per token with no monthly fee, a low-volume workload on a cheap model costs very little in absolute terms. ## What does the 99.9% SLA actually cover? It covers the platform layer OfoxAI operates: the API gateway, the console, billing and account services. It does not cover upstream model availability, which is set by the model providers themselves. When an upstream provider degrades, the gateway returns the provider's real error rather than masking it, so your retry logic can react to what actually happened. ## What happens when an upstream provider goes down? The gateway surfaces the upstream error with its original status code. For models served by more than one upstream provider, traffic can be routed to a healthy provider instead. Because the error is passed through rather than rewritten, a 429 from a provider still reads as a 429 and a 500 still reads as a 500. ## How do I choose between models? Every model entry in the Model Catalog section lists its context window, its declared capabilities, and its exact price, which is usually enough to narrow the field. For a guided choice, the model finder at https://ofox.run/model-finder ranks models against a described task by quality, cost and speed. Model pages also allow side-by-side comparison. ## What is the difference between context window and max output? The context window is the total number of tokens a model can consider in one request, counting your prompt and its own reply. Max output is the ceiling on the reply alone. A model with a 1,000,000-token context window and a 128,000-token max output can read a very large input but will still stop generating at 128,000 tokens. ## Which coding tools work with OfoxAI? Claude Code, Codex, OpenCode, Zed, Cline, Chatbox, CherryStudio, OpenClaw, Kilo and Aider all work, as does any tool that lets you set a custom API base URL. Configuration is two values — a base URL and an API key — usually one environment variable or a few lines in a config file. Per-tool setup is in the Integration Guides section of this document. ## Do I have to change my prompts or tool definitions when I switch gateways? No. Because the protocols are native, your prompts, your tool and function definitions, and your MCP server configuration are sent unchanged. The only value that changes is the base URL, and the only behaviour that changes is which account is billed. ## Can I generate video and images with the same key? Yes. Video generation uses the asynchronous `POST /v1/videos` endpoint — create a job, poll it until it reaches a terminal state, then download the result — and image generation uses `POST /v1/images/generations`. Both authenticate with the same API key as text models and draw down the same balance. ## How is video generation billed? Per second of generated video, not per token. Each video model publishes a tier table where the price depends on resolution and on the input mode (text-to-video, image-to-video, video-to-video) and, for some models, on whether audio is generated. The entry price in the Pricing Tables section is the cheapest tier; the full tier table for each model is in its Model Catalog entry. ## Does prompt caching work through OfoxAI? For models whose Capabilities line lists prompt caching, yes, and cache read and cache write prices are published as separate pricing fields on those models. Caching behaviour differs by protocol: the OpenAI protocol matches prefixes automatically, while the Anthropic protocol requires explicit cache breakpoints in the request. Send the cache directives your SDK documents and the gateway passes them through. ## Is OfoxAI available worldwide? Yes. Requests are served over a globally accelerated network with edge nodes in several regions, which keeps latency predictable regardless of where your service runs. Support is available in English and Chinese. ## How quickly are new models added? New models are onboarded continuously and appear in the catalog — and therefore in this document, which is generated from the catalog — as soon as they are enabled. The model list on this page is not hand-maintained, so it cannot drift out of sync with what the API actually serves. ## Where is the machine-readable model list? `GET /v1/models` returns every model with live pricing and capability data, and the catalog is browsable at https://ofox.run/models. This document is regenerated from the same source, so the IDs quoted here are the IDs the API accepts. ## Does OfoxAI support enterprise requirements like SSO and private deployment? The Enterprise plan covers dedicated capacity, SSO, audit logs, invoicing, purchase orders and custom SLA terms, along with private and hybrid cloud deployment options. Details and contact routes are at https://ofox.run/enterprise. ## Where do I report a problem or ask a question? Documentation is at https://ofox.run/docs, the model catalog at https://ofox.run/models, and support is reachable from the console. For questions about a specific model's behaviour, the model page for that ID is the most precise starting point.