20% off all GPT ๐ŸŽ‰ 30% off DeepSeekLearn more

OpenAI: GPT-5.6 Luna

Chat-80%
openai/gpt-5.6-luna

GPT-5.6 Luna is the fast, cost-efficient tier of OpenAI's GPT-5.6 series, released on 2026-07-09. It is built for high-volume, latency-sensitive workloads such as chat, classification, and lightweight agentic workflows, providing capable reasoning for its price tier. It supports reasoning, vision, tool use (function calling), prompt caching, and web search. Context window: 1M tokens, output: 128K. Prompt caching keeps repeated context inexpensive across high-traffic sessions. Available via OpenAI and Anthropic protocols through Ofox.

Context Window
1M
Max Output Tokens
128K
Released
2026-07-09
Capabilities
VisionFunction CallingReasoningPrompt CachingWeb Search
Available Providers
OpenAIAzure
Supported Protocols
openaianthropic

Providers

Azureprovider.type: "azure_foundry"
-80%
Input Tokens
$0.2/M
$1/M
Output Tokens
$1.2/M
$6/M
Cache Read
$0.02/M
$0.1/M
Cache Write
$0.25/M
$1.25/M
Output Image
$32/M
Web Search
$0.035/R
Protocols
openai/v1/chat/completions/v1/responses
anthropic
OpenAIprovider.type: "openai"
-80%
Input Tokens
$0.2/M
$1/M
Output Tokens
$1.2/M
$6/M
Cache Read
$0.02/M
$0.1/M
Cache Write
$0.25/M
$1.25/M
Output Image
$32/M
Web Search
$0.01/R
Protocols
openai/v1/chat/completions/v1/responses

Code Examples

from openai import OpenAI
client = OpenAI(
base_url="https://api.ofox.run/v1",
api_key="YOUR_OFOX_API_KEY",
)
response = client.chat.completions.create(
model="openai/gpt-5.6-luna",
messages=[
{"role": "user", "content": "Hello!"}
],
)
print(response.choices[0].message.content)

Uptime & Status

Further Reading on GPT-5.6 Luna

Frequently Asked Questions

OpenAI: GPT-5.6 Luna on Ofox.ai costs $1/M per million input tokens and $6/M per million output tokens. Pay-as-you-go, no monthly fees.