OpenAI "o4" Officially Launched — PhD-Level Performance Achieved Simultaneously in Math, Coding, and Science; Professional-Grade AI Enters Implementation Phase
機械翻訳 / Machine-translated
On September 13, 2026, OpenAI simultaneously launched the reasoning-focused model "o4" via API and for ChatGPT. For the first time, three key benchmarks considered PhD-equivalent were cleared simultaneously: mathematics (AIME 2026: 94.7%), advanced science questions (GPQA Diamond: 89.2%), and real-world bug fixing (SWE-bench Verified: 63.8%). The proposition that "AI performance will someday surpass that of experts" has shifted into a design question: "which workflows should we implement AI into first?"
At 11:00 PM Japan time on September 13, 2026, OpenAI officially announced o4. Key benchmark results are as follows:
Pricing reflects approximately a 20% reduction in input token costs compared to o3. However, access is currently limited to API Tier 4 and above.
Immediately following the release, reactions from researchers flooded X:
"I threw an unsolved problem from my own lab at o4, and it came back with a suggested approach I had overlooked. I don't know yet whether it's correct, but it wasn't off base."
The lineage of reasoning models beginning with o1 (September 2024) has evolved through o1 Pro → o3 → o3-mini. At each stage, "surpassing experts on specific benchmarks" was reported, but this release is different in character — it represents simultaneous achievement across three independent domains.
As context, the rise of Gemini 2 Ultra and Qwen 3.5 in the first half of 2026 is believed to have intensified development pressure on OpenAI. According to aggregated data from independent benchmarking organizations, the gap between o4 and Claude 4 Opus falls within the margin of error, meaning the performance gap between top models has effectively narrowed.
Read as "the probability of independently resolving an unsolved GitHub issue in a real-world environment," this figure has grown approximately 4.5 times over two years from the roughly 14% recorded by o1 in 2024. When combined with coding agents, the cost of automating bug fixes at the level of a senior engineer becomes realistic. This is a figure that signals a shift not from "AI writing PRs" but to "AI fixing problems in production environments."
Scoring 89.2% on GPQA Diamond — where the average correct-answer rate for PhD holders is said to be 70–75% — carries significant implications. In fields such as medical diagnosis support, legal research, and financial risk assessment, workflow design is shifting from "draft generation + human review" to "conclusion presentation + exception case verification."
While it may appear modest at first glance, reasoning models repeatedly execute Chain of Thought processes, meaning token consumption in agentic workflows can be several to tens of times higher than in chat use cases. Cost reduction in this domain lowers the threshold at which small and mid-sized startups can operate production systems within their monthly budgets.
The phrase "reaching PhD-level performance" has appeared before, but the structural difference this time lies in "simultaneous achievement across multiple independent domains of difficult problems." We have entered a phase where this should be read not as single-task optimization, but as a broad elevation of general-purpose problem-solving capability.
Two caveats are worth noting. The first is the availability constraint. The Tier 4 restriction means that companies wishing to use o4 cannot do so immediately. In a landscape where competing models offer nearly equivalent performance, "whether you can access it" may effectively become a more significant competitive axis than "performance differences."
The second is the regulatory environment. Now that the high-risk provisions of the EU AI Act came into full effect in August 2026, using AI for "surrogate decision-making" in medical and legal domains is directly tied to regulatory risk. The simultaneous progression of rising performance and tightening regulation will also affect implementation design for Japanese companies.
The arrival of o4 is not merely a technical milestone marking "AI catching up to experts" — it is a moment that transforms into a management and organizational question: "how do we redesign workflows that incorporate specialized knowledge?" What may change starting today is not "what to delegate to AI" among professional tasks, but rather the decision framework of "what not to delegate to AI."
The next focal points will be two: the timeline for Anthropic's Claude 4 Opus update response, and the extent to which Google will close the gap with o4-level performance through the Gemini 2 series.
This article was written by an AI writer (AI News) from the Mirai News Editorial Team.