Skip to content

[Windows][Chat] Pro / Extra High stay selected but responses are unusually fast and low quality; web, mobile and Work behave normally #48939

Description

@FrederickYoung1010

What version of the Codex App are you using (From “About Codex” dialog)?

Windows package version: 26.924.2738.0, verified from the currently running ChatGPT.exe package path. The About dialog's display version/build was not separately verified.

Affected surface: regular Chat in the unified ChatGPT/Codex desktop app, not a Codex CLI run and not a Work task.

What subscription do you have?

ChatGPT Pro 20x (user-reported). The user reports approximately 80% remaining in the displayed usage allowance and recalls a broadly announced quota refresh/reset in the preceding two days.

Those allowance/reset details are user reports, not an independent audit of backend entitlements. The exact allowance bucket represented by the 80% display was not established; this report does not equate a Work/Codex allowance with a separate Chat Pro allowance. Nevertheless, mobile Chat, web Chat, and Work continue to behave normally for this user, and the affected desktop Chat does not show a quota-exhaustion explanation. Please check the actual Chat entitlement rather than assuming quota exhaustion.

What platform is your computer?

Windows, x64. This report concerns the official desktop application.

What issue are you seeing?

Since approximately September 27–28, 2026, regular desktop Chat has shown a persistent, substantial regression in response quality and reasoning behavior across the settings tested, explicitly including Extra High (XH) and Pro.

The user describes answers as starting almost immediately, lacking the usual visible deliberation, and performing markedly worse on both everyday planning/instruction-following tasks and a small SVG animation test. Selecting a higher setting does not restore the previous behavior.

Important distinction: this is NOT the older selector-reset problem. This user previously experienced the issue where restarting the app or selecting a reasoning level caused it to revert to the lowest/Instant setting. That older problem was resolved for them. In the current incident, the selected setting can remain visible—specifically Pro in the retained reproduction screenshot—while the response behavior remains abnormal. Please do not close this as a duplicate of the old slider-reset defect without checking whether the selected settings are actually applied at submission.

Cross-surface comparison (user observations)

Surface Current reported behavior
Windows desktop — regular Chat, XH Abnormally rapid responses and substantially worse results
Windows desktop — regular Chat, Pro Same abnormal behavior; Pro remains displayed
ChatGPT web on the same computer/account Normal reasoning behavior and expected quality
Official mobile ChatGPT app Normal behavior
Work mode Normal reasoning/quality; the current quality regression is isolated to Chat

These are repeated observations by an established user of this workflow, not a controlled benchmark proving a particular backend model substitution. The user has used Chat for strategy/task instructions and Work for execution for roughly three months; the abrupt difference is disruptive to that previously usable workflow.

Retained simple reproduction

A new regular desktop Chat was created, Pro was selected, and the following prompt was submitted:

创建一个HTML,内容是:SVG绘制一个鹈鹕骑自行车的2D动画。

English translation:

Create an HTML page containing a 2D SVG animation of a pelican riding a bicycle.

The desktop returned a single-file HTML/SVG/CSS answer very quickly. The user judges its result substantially below their usual Pro/XH outputs. In the provided rendered frame, the two wheel centers were vertically misaligned and the bird's legs did not visually connect to the pedals. These are observations of one captured frame, not a complete animation audit.

A read-only conversation inspection returned a start/end timestamp difference of approximately 5.103 seconds for this retained sample. This is the difference between the record fields exposed by the inspection tool, not independently measured first-token latency, reasoning duration, or total rendering time.

The user also compared a simple model-identity question across surfaces. A web screenshot showed Pro selected and a “Thought for 30s” indicator, whereas desktop Chat answered rapidly. A model's self-description is not reliable proof of its actual identity, and this simple question is not a capability benchmark; the more important concern is the repeated difference on practical tasks and the animation reproduction.

Troubleshooting already completed

None of the following restored desktop Chat's previous behavior:

  • Multiple brand-new regular Chat conversations.
  • Testing both Extra High and Pro.
  • Restarting the desktop application.
  • Restarting Windows.
  • Signing out of the desktop application, restarting Windows, signing back in through the browser authorization/confirmation redirect, and returning to the app.
  • Repeating the test in new desktop Chat conversations after that fresh login.

The browser authorization round-trip completed, but the symptom persisted. Some earlier test conversations were deleted; the retained reproduction above was created afterward. No claim is made that a current reinstall or profile reset has been performed.

What steps can reproduce the bug?

These reproduce the symptom on the affected installation; a clean-install or other-account reproduction has not been established.

  1. Sign in to the official Windows desktop application with the affected Pro account.
  2. Start a new regular Chat, not Work or Codex.
  3. Select Pro and confirm it remains displayed.
  4. Submit the pelican SVG/HTML prompt above.
  5. Observe the unusually rapid reply and the reported loss of output quality compared with this user's normal Pro behavior.
  6. Repeat in another new regular Chat with Extra High; the user reports the same quality/latency symptom.
  7. Compare the account's behavior in web Chat and the official mobile app, and compare reasoning/quality in Work. The user reports that these remain normal.
  8. Repeat after the completed sign-out/restart/sign-in cycle: desktop Chat remains affected.

A formal paired statistical evaluation using identical histories and repeated trials has not been performed. Maintainer-side inspection of the selected versus submitted versus executed settings would be more useful than asking the user to repeat sign-in cycles again.

What is the expected behavior?

  • The desktop Chat model/reasoning setting should be honored when a message is submitted.
  • If a selected mode cannot be used because of entitlement, allowance, availability, or another restriction, the app should explain that state rather than leave Pro/XH displayed while behaving unexpectedly.
  • Regular desktop Chat should not exhibit a persistent, large quality/behavior discrepancy from the same account's other working surfaces without an actionable explanation.
  • This is not a request to artificially delay answers or display a fixed thinking timer. The issue is whether the chosen mode is actually applied and whether desktop-specific request/state handling is faulty.

Additional information

Narrow read-only diagnostic findings

The main-process and renderer logs were inspected for the two-minute window containing the retained reproduction.

  • No connection-closed, reconnect, or explicit model-fallback event was found in that inspected window. Absence from these files is not proof that every request/transport was healthy.
  • Four 404 conversation_deleted errors concerned other, previously deleted conversations. They were not attributed to the retained new reproduction.
  • After the sample's recorded answer timestamp, the client made a queue-list request involving that regular Chat's identifier and logged the following sanitized failure:
method=thread/queue/list
errorCode=-32603
failureReason=rollout_not_found
failed to read thread: invalid thread-store request: no rollout found for thread id <redacted>

This may be an unrelated desktop Chat/local-thread adaptation problem. It happened after the answer timestamp and is not established as the cause of the response-quality regression. Please assess it separately if appropriate.

The inspected logs did not expose the actual Chat generation payload's model/reasoning parameters or authoritative backend model identity. Background suggestion/title-generation model names, IPC fallback fields, and Work model settings must not be interpreted as proof of which model answered the affected Chat.

Earlier intermittent Work reconnections were investigated separately. The user's latest report is that Work reasoning/quality is normal. There is no established causal chain from those reconnections to the current Chat-only issue.

Requested investigation

Please investigate:

  1. Whether selecting Pro/XH in regular Windows Chat correctly populates the generation request, rather than only updating the visible selector.
  2. Whether desktop-specific initialization, per-message overrides, intelligence presets, persisted state, or feature flags can override the selected mode.
  3. Whether Chat entitlement/allowance state is synchronized and interpreted consistently across desktop, web and mobile; in particular, whether an unrelated Work/Codex allowance affects desktop Chat incorrectly.
  4. Whether the affected requests were executed using the intended backend model and reasoning settings, and whether any fallback occurred.
  5. Whether thread/queue/list → rollout_not_found on a regular Chat identifier is expected, a separate bug, or relevant to this incident.
  6. What supported, narrowly scoped diagnostic can expose requested/effective mode without requiring credentials, a full profile export, or repeated sign-out/reset/reinstall attempts.

The report describes a user-visible regression and suspected mode-application/routing problem. It does not assert confirmed silent model substitution, deliberate throttling, or a proven cause.

Related reports reviewed, not asserted duplicates

Please link this to an existing tracker if the underlying defect is confirmed to be the same, while preserving the distinction between selector reset and selected mode not apparently reflected in generation behavior.

Privacy

Submitted with the user's explicit authorization. No account email, private conversation/request identifiers, local usernames or project paths, IP addresses, network endpoint details, screenshots, credentials, or raw logs are included. Any further private diagnostic correlation would require an appropriate private support channel and the user's approval.

Activity

  1. added
    bugSomething isn't working
    appIssues related to the Codex desktop app
    windows-osIssues related to Codex on Windows systems
    model-behaviorIssues related to behaviors exhibited by the model
    on Sep 28, 2026
  2. github-actions commented on Sep 28, 2026

    @github-actions
    Contributor

    Potential duplicates detected. Please review them and close your issue if it is a duplicate.

    Powered by Codex Action

  3. XAN9xXx commented on Oct 1, 2026

    @XAN9xXx

    Same issue on Windows 10, ChatGPT Desktop v26.928.3736.0. Extra High remains selected in the UI, but responses consistently start almost immediately, while the same account on ChatGPT Web with Extra High consistently spends much longer reasoning.

  4. Pizno555 commented on Oct 8, 2026

    @Pizno555

    work or chat

    Same issue on Windows 10, ChatGPT Desktop v26.928.3736.0. Extra High remains selected in the UI, but responses consistently start almost immediately, while the same account on ChatGPT Web with Extra High consistently spends much longer reasoning.

  5. XAN9xXx commented on Oct 8, 2026

    @XAN9xXx

    work or chat

    chat

  6. lostlight530 commented on Oct 11, 2026

    @lostlight530

    Related observation — ChatGPT Web / GPT-6 High (different affected surface)

    Adding a related observation rather than claiming this is the same bug. In contrast to this issue's Windows desktop Chat affected / Web normal pattern, I observed unusually fast responses and uneven analytical depth in regular ChatGPT Web with GPT-6 High selected, beginning around October 7, 2026. My Android ChatGPT experience remained generally normal, though I have not performed a strictly matched cross-client evaluation.

    I used a synthetic, self-contained Agent Runtime forensic audit with explicit contracts for event ordering, at-least-once delivery, idempotency-ledger semantics, task completion, and a later linearizable state snapshot.

    • Initial audit: the response arrived in approximately 1 second (user-observed, not instrumented). It correctly handled the main event counts and contract violation, but had concrete second-order flaws: it presented an already ledger-confirmed commit as an open hypothesis and proposed a reconciliation status without adequately preserving the irreversible terminal-state contract.
    • Guided self-audit: a follow-up response arrived in approximately 10 seconds (also user-observed). After the prompt explicitly identified the suspected flaws, it corrected several issues. This demonstrates prompted correction, not independent error discovery and not a measured improvement attributable to longer thinking.
    • Other observation: an incognito Web session produced a long, structured response with about 9 seconds of visible thinking. Length and UI thinking time are not reliable measures of effective reasoning effort.

    Important confounder / scope correction: I independently identified a proxy/network-path problem affecting Intelligent UI availability. I am not reporting the UI refusal as a model bug here, and I have not established that the earlier response-quality observations persist under a fully controlled post-network-fix setup.

    Evidence boundary: These are limited qualitative samples, not a validated before/after regression benchmark. They do not demonstrate that High was silently downgraded, a backend model was substituted, or the reasoning budget was reduced. I also cannot attribute the symptoms to the same root cause as this Windows desktop report.

    If maintainers investigate selected-versus-effective model/effort behavior across clients, it may be worth checking whether Web Chat can exhibit a related mismatch or an early-finalization pattern. Please treat this as a separate cross-surface datapoint, not a confirmed duplicate.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    appIssues related to the Codex desktop appbugSomething isn't workingmodel-behaviorIssues related to behaviors exhibited by the modelwindows-osIssues related to Codex on Windows systems

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions