← ChatGPT & Codex changelog
OpenAI API — September, 2026
This release introduces major advancements across models and APIs, enhancing AI capabilities for a wide range of users. GPT-6 Astra, our most capable model, offers advanced controls for long-running work, while the new Agents API enables AI with computer use and managed sessions. New GPT-6.1 Sol, Sol, and Luna models provide enhanced reasoning, and GPT Image 2.5 Sunburst and Flare improve image generation. API key governance, general availability for GPT-Live 1, and a fix for image understanding in Sol/Luna ensure a more robust and versatile platform for all users.
Added
- Added computer use to the Agents API. Agents can complete tasks in an OpenAI-hosted browser, with website access approvals and sign-in handled by your application.
- Released GPT-6.1 Sol (gpt-6.1-sol) for complex coding and professional work at a lower cost than GPT-6 Astra.
- Added Ultrafast mode for GPT-6 Astra in the Responses API. Use gpt-6-astra with service_tier: "ultrafast" to reduce the time between generated output tokens. It is available to API customers, subject to rate limits, with global processing and US data residency. EU and other regional inference residency aren't supported. See Ultrafast pricing.
- Released GPT-6 Sol (gpt-6-sol) and GPT-6 Luna (gpt-6-luna).
- Added API key creation governance controls at the organization and project levels. Administrators can allow only service-account keys, allow only user-owned project keys, or disable all new API key creation. Organization restrictions take precedence over project settings, and existing API keys are unaffected. See production best practices for details.
- Released the Agents API in public beta. Build agents with a managed Codex harness while OpenAI handles session orchestration, context compaction, and recovery.
- Released GPT Image 2.5 Sunburst and GPT Image 2.5 Flare for image generation and editing through the Image API and the Responses API image generation tool.
- Use Sunburst for workflows where editing precision matters most, or Flare for fast, high-quality everyday image generation. Both models support the new xhigh and max quality settings and use GPT Image 2 token rates. See the image generation guide and pricing.
- Released GPT-6 Astra, our most capable model, built for the hardest end-to-end work.
- Added new controls for long-running work with GPT-6 Astra in the Responses API:
Changed
- Standard pricing per 1M tokens for prompts with up to 272K input tokens is $2 input, $0.10 cached input, $2.50 cache write, and $10 output.
- GPT-6.1 Sol also supports Multi-agent in beta. Let the model delegate work to subagents in a Responses API request.
- Use the Responses API for tool calling. See GPT-6 model guidance for reasoning settings and pricing for available processing tiers.
- If your use cases involve image inputs, we recommend rerunning your evaluations and retrying workflows affected by the issue.
- These reasoning models accept text and image inputs and generate text through the Responses and Chat Completions APIs.
- Standard pricing per 1M tokens for prompts with up to 272K input tokens:
- GPT-6 Sol: $2 input, $0.20 cached input, and $10 output.
- GPT-6 Luna: $0.10 input, $0.01 cached input, and $0.50 output.
- Compare capabilities in the model catalog, and see pricing for cache writes, longer prompts, and other processing tiers.
- You can now set expiration dates when creating project API keys. Administrators can also enforce a maximum key lifetime at the organization or project level in Platform settings, requiring newly created keys to expire within the configured limit. See production best practices for guidance on key expiration and rotation.
- Use durable sessions to continue work across turns, stream progress, and connect your own tools and MCP servers. Run agents in OpenAI-hosted sandboxes or connect a sandbox from your own infrastructure or a supported provider.
- Start with the Agents API quickstart.
- GPT-Live 1 is now generally available in the API. Build full-duplex voice conversations that can continue while a backend model or agent handles reasoning and tools.
- Use Responses delegation with an OpenAI model, or client delegation to connect your own backend. Voice sessions cost $0.05 per minute, billed per second; backend model and tool usage is charged separately.
- Start with GPT-Live, prompting, and migration guidance. See pricing for details.
- Prompt Cache Diagnostics is now generally available in the Responses API for GPT-5.6 and later supported models.
- Compare cache reuse against a previous response, identify reasons for cache misses, and follow troubleshooting guidance to improve cache reuse.
- GPT-Rosalind (gpt-rosalind-research) is now generally available through the trusted-access program for approved internal life sciences research.
- Standard pricing is $5 per 1M input tokens, $0.50 per 1M cached input tokens, and $25 per 1M output tokens. Billing begins on October 5, 2026. See pricing for details.
- Use GPT-6 Astra for reasoning, coding, computer use, research, and document creation. It combines these capabilities to carry complex tasks from an initial request to a finished result, using the context and tools you provide.
- Key changes to consider when migrating:
- GPT-6 Astra does not support the none reasoning effort level.
- GPT-6 Astra does not support custom temperature or top_p values or log probabilities (logprobs).
- Tool calling requires the Responses API. If you use tools with Chat Completions, follow the Responses migration guide.
- Misalignment monitoring asynchronously checks for potential issues during agent work in supported Responses API requests. Checks can trigger safety alerts or stop a conversation for review.
- Start with Using GPT-6 Astra for capabilities, prompting, and migration guidance. Explore computer use for browser and desktop workflows, and see pricing for available inference tiers.
- Async tool calling: Let the model continue working while your application runs function or custom tools, then return results as they become available.
- Mid-turn steering: Send additional instructions while a response is in progress over WebSockets, so the model can incorporate corrections or changing requirements.
- Change reasoning effort mid-conversation: Increase effort for difficult work or reduce it for routine follow-ups while preserving the cached prompt prefix.
- Updated API errors so applications can distinguish traffic that increases too quickly from temporary model overload.
- Traffic that increases too quickly can return a 429 error with the slow_down code. Temporary model overload returns a 503 error with the server_is_overloaded code. Both responses may include Retry-After. When the header is present, wait at least as long as it specifies before retrying. If it's missing, use exponential backoff. See the error codes guide and rate limits guide.
- Connections to api.openai.com can now use IPv6.
Fixed
- Fixed a bug in image encoding that degraded image understanding in GPT-6 Sol and GPT-6 Luna. This update improves results on visual tasks in the API and Codex, including computer use.