Is Codex getting worse?
Compare the latest GPT-5.6 capability scores across reasoning efforts and see whether several settings are falling together.
No broad drop detected
The latest comparable efforts do not show a broad decline.
Best measured
Max
Latest benchmark
Codex Radar IQ
110.3
82/112 tasks
Reasoning effort comparison
Which effort is measuring best today?
Higher effort is not automatically better. “Best measured” means the highest current score on this benchmark—not a guarantee for every coding task.
Low
78.058/112 benchmark tasks
78.0
0.0 vs previous
Medium
99.574/112 benchmark tasks
99.5
0.0 vs previous
High
98.273/112 benchmark tasks
98.2
0.0 vs previous
XHigh
102.276/112 benchmark tasks
102.2
0.0 vs previous
Max
110.382/112 benchmark tasks
110.3
0.0 vs previous
Ultra
No public result for this effort
—
Awaiting data
Missing efforts stay empty until the source publishes a real result.
Recent direction
Max capability trend
Independent benchmark source
Data comes from Codex Radar’s distributed task runs. Codex Reset stores each published snapshot and checks for updates hourly. Operational fields stay out of this page so the comparison is easy to read.
Last source observation: Sep 1, 2026, 6:06 PM UTC
Codex capability FAQ
Is Codex getting worse?+
We look for the same downward movement across multiple reasoning efforts. One lower score can be ordinary benchmark variation, so the page only calls out a possible broad drop when at least two comparable efforts fall together.
Is High better than Max or Ultra?+
Sometimes. A higher reasoning setting is not automatically better for every task. The effort comparison shows the latest measured score for each setting, and leaves a setting blank when no public result exists.
Is this the same as Codex Juice?+
No. Juice is an observed runtime-budget signal. Capability scores come from coding-task results. A Juice change may tell us that a configuration changed, but it does not prove that coding quality improved or declined.