Models, reasoning, and usage
On this page
Choose a model, reasoning effort, and speed with the controls below the reply box. The available options depend on the connected agent and machine.
Select a model and reasoning effort
Open the model picker and select a model. On desktop, older supported
generations appear under More models when available, and Claude threads
also accept /model to open the picker or /model <model> to select a known
model directly.
If the model offers reasoning controls, choose how much effort it should spend on the task. More reasoning can help with difficult work, such as investigating a bug with several plausible causes, but may take longer or use more of your allowance. Supported Codex models also offer Standard and Fast speed; Fast will increase usage.
Claude's model and reasoning choices stay with the thread when you reconnect or reopen it on desktop or mobile, and changing them will not change another existing thread. The last-used model, reasoning, and speed choices on a machine become the defaults for new drafts there.
Understand the context window
The context window is how much of the conversation the model can work with at once. As a Codex conversation fills that window, you may see an auto-compaction warning or Compact now. Compaction reduces the retained context so the conversation can continue; it does not reset your usage allowance.
Check your remaining allowance
Usage allowances come from the connected provider account. On desktop, the agent's rate-limit indicator sits with the model controls below the reply box; open it to check the reported allowance and reset times.
Codex shows session, weekly, and available model-specific quotas, highlighting limits for the selected model. If your ChatGPT account reports credits for resetting a usage limit, the indicator also shows their count and expiry.
Claude reports session and weekly quota shared by machines using the same Anthropic organization and account. After you sign in with a different account, an already-running thread may still use the previous credentials and show that account's quota.
Inspect recorded activity
On desktop, choose View usage from the rate-limit indicator to see the current thread's recorded activity: tokens (the units of model input and output), model calls, average and largest input context, and the balance of new input, cached input, and output.
Choose the 24-hour view for hourly bars or the seven-day view for hourly or daily bars, and refresh to load newer activity. The chart compares this thread with one combined Other activity series; it does not name or link to other threads.
These measurements come from saved conversations and may differ from the provider's quota calculation. The dialog identifies activity that cannot be matched to the selected account and warns when missing history makes the totals incomplete.
Machine usage is measured separately in boxes.dev box-hours; see Plans, seats, and box-hours.
When Claude changes the model
Claude can switch a Fable conversation to Opus when Fable's safeguards decline part of the conversation and Claude retries it with another model. The thread will show a switch notice, and the picker will show the model now in use. You can select Fable again, though another switch may occur. A refusal without a retry on another model will leave the selected model unchanged.