OpenAI
Refined GPT-5 family release focused on chat quality, coding reliability, and better instruction-following in production use.
OpenRouter
Price record: 2026-10-03. Source: openrouter.
Confirm this specific model, usage limits, and billing terms with the provider. A subscription does not automatically include API credits.
No login needed to compare. Prices are in USD; provider charges are separate from AI Market Cap plans. Context length, caching, tools, taxes, and regional terms can change the final cost. Open weights do not mean free hosting.
23.9
Quality Score
---
Arena ELO
Undisclosed
Parameters
256K
Context
This measures the amount of verifiable public evidence we have, not how capable the model is. A missing field means it has not been verified yet, not that its value is zero.
15 of 22 public signals
Sign in to join the discussion
0
Downloads
0
Likes
Feb 2026
Released
4/5 signals
3/4 signals
5/5 signals
1/4 signals
2/4 signals
Gaps we are still tracking
Launches
1
Benchmarks
3
API
1
Safety
2
Research
2
General
4
Recent launch, pricing, benchmark, and API signals linked to this model or its provider.
Introducing GPT‑5 for developers | OpenAI Skip to main content Research Products Business Developers Company Foundation (opens in a new window) Log in Try ChatGPT (opens in a new window) Research Products Business Developers Company Foundation (opens in a new window) Try ChatGPT (opens in a new window) Login OpenAI August 7, 2025 Product Introducing GPT‑5 for developers The best model for coding and agentic tasks. Loading… Share Introduction Introduction Coding Frontend engin
Introducing GPT‑5 for developers | OpenAI Skip to main content Research Products Business Developers Company Foundation (opens in a new window) Log in Try ChatGPT (opens in a new window) Research Products Business Developers Company Foundation (opens in a new window) Try ChatGPT (opens in a new window) Login OpenAI August 7, 2025 Product Introducing GPT‑5 for developers The best model for coding and agentic tasks. Loading… Share Introduction Introduction Coding Frontend engin
OpenAI published official provider-reported benchmark evidence for GPT-5.3, GPT-5.3 Instant, gpt-5.3-instant, GPT-5.3 Chat.
View sourceLoading… Share Frontier agentic capabilities Frontier agentic capabilities Coding Web development Beyond coding An interactive collaborator How we used Codex to train and deploy GPT-5.3-Codex Securing the cyber frontier Availability & details What’s next Appendix Frontier agentic capabilities Coding Web development Beyond coding An interactive collaborator How we used Codex to train and deploy GPT-5.3-Codex Securing the cyber frontier Availability & details What’s nex
View sourceCodex Security Cloud is getting a major upgrade, with access to cyber-capable models through Daybreak Blue included by default. It scans entire GitHub repos, continuously reviews new commits, investigates and deduplicates findings, and prepares fixes for review – even when your https://t.co/up1hpkiCAK
We’ve shared details on how AI agents in our research environment sent training and evaluation data to third-party services when they shouldn’t have. Most of that data did not come from users. We have discovered 53 cases where images that people had uploaded were posted to
After the Hugging Face incident, we committed to conducting a much broader review of actions taken by our models during training and evaluation and to being transparent about our findings. This is an extensive review that is ongoing. The vast majority of actions we’ve reviewed
AI agents powered by large language models exhibit strong reasoning and problem-solving capabilities, enabling them to assist scientific research tasks such as formula derivation and code generation. However, whether these agents can reliably perform end-to-end reproduction from real scientific papers remains an open question. We introduce PRBench, a benchmark of 30 expert-curated tasks spanning 11 subfields of physics. Each task requires an agent to comprehend the methodology of a published paper, implement the corresponding algorithms from scratch, and produce quantitative results matching the original publication. Agents are provided only with the task instruction and paper content, and operate in a sandboxed execution environment. All tasks are contributed by domain experts from over 20 research groups at the School of Physics, Peking University, each grounded in a real published paper and validated through end-to-end reproduction with verified ground-truth results and detailed scoring rubrics. Using an agentified assessment pipeline, we evaluate a set of coding agents on PRBench and analyze their capabilities across key dimensions of scientific reasoning and execution. The best-performing agent, OpenAI Codex powered by GPT-5.3-Codex, achieves a mean overall score of 34%. All agents exhibit a zero end-to-end callback success rate, with particularly poor performance in data accuracy and code correctness. We further identify systematic failure modes, including errors in formula implementation, inability to debug numerical simulations, and fabrication of output data. Overall, PRBench provides a rigorous benchmark for evaluating progress toward autonomous scientific research.
OpenAI published official provider-reported benchmark evidence for GPT-5.3, GPT-5.3 Instant, gpt-5.3-instant, GPT-5.3 Chat.
Loading… Share Frontier agentic capabilities Frontier agentic capabilities Coding Web development Beyond coding An interactive collaborator How we used Codex to train and deploy GPT-5.3-Codex Securing the cyber frontier Availability & details What’s next Appendix Frontier agentic capabilities Coding Web development Beyond coding An interactive collaborator How we used Codex to train and deploy GPT-5.3-Codex Securing the cyber frontier Availability & details What’s nex