Codex capability radar

Is Codex getting worse?

Compare the latest GPT-5.6 capability scores across reasoning efforts and see whether several settings are falling together.

No broad drop detected

The latest comparable efforts do not show a broad decline.

Best measured

Max

Latest benchmark

Codex Radar IQ

110.3

82/112 tasks

Reasoning effort comparison

Which effort is measuring best today?

Higher effort is not automatically better. “Best measured” means the highest current score on this benchmark—not a guarantee for every coding task.

Low

78.0

58/112 benchmark tasks

Medium

99.5

74/112 benchmark tasks

High

98.2

73/112 benchmark tasks

XHigh

102.2

76/112 benchmark tasks

Max

110.3

82/112 benchmark tasks

Ultra

No public result for this effort

Missing efforts stay empty until the source publishes a real result.

Recent direction

Max capability trend

110.3110.3110.3110.3110.3110.3110.3110.3110.3110.3Aug 31Sep 1

Independent benchmark source

Data comes from Codex Radar’s distributed task runs. Codex Reset stores each published snapshot and checks for updates hourly. Operational fields stay out of this page so the comparison is easy to read.

Last source observation: Sep 1, 2026, 6:06 PM UTC

Data from Codex Radar

Codex capability FAQ

Is Codex getting worse?+

We look for the same downward movement across multiple reasoning efforts. One lower score can be ordinary benchmark variation, so the page only calls out a possible broad drop when at least two comparable efforts fall together.

Is High better than Max or Ultra?+

Sometimes. A higher reasoning setting is not automatically better for every task. The effort comparison shows the latest measured score for each setting, and leaves a setting blank when no public result exists.

Is this the same as Codex Juice?+

No. Juice is an observed runtime-budget signal. Capability scores come from coding-task results. A Juice change may tell us that a configuration changed, but it does not prove that coding quality improved or declined.