CLI Quota Exhausted: How Do You Order the Fallback?
When Cursor CLI / Agent hits a wall, requests fail and banners talk about usage limits or ask you to switch to Auto or another model. Reinstalling does not fix it. The useful order is: see which pool is empty, explicitly pick a model in a remaining pool, or cut the work cleanly and resume later.
This post covers only symptoms · model-switch criteria · how to stop cleanly. Grounded in Usage and limits, CLI overview, CLI parameters, and Slash commands. No plan prices or dollar amounts—only quota UI concepts (pools, remaining allowance, reset date, on-demand toggle). No invented personal reviews.
What are the symptoms?
One-line answer: If the CLI or editor shows a usage / rate-limit notice and the same model keeps refusing, treat it as an account-side pool first—not a local install bug or a silent logout.
Common signals:
| Signal | Likely meaning | Do first |
|---|---|---|
| “You’ve hit your usage limit” / “You’re out of usage” | Included usage (a pool) exhausted, or that model is blocked | Check Spending or the CLI usage UI |
| Banner urging Auto or another model | Current model/pool is blocked | Identify which pool, then switch explicitly |
| Same prompt fails on every retry | Server refusal more likely than flaky Wi‑Fi | Prefer pool/model checks over reinstall |
agent starts but the first reply is refused | Session alive; requests gated | /model, --model, agent models |
Official docs: most paid plans expose two monthly usage pools:
- Cursor Models — first-party Cursor models (docs cite Grok-family and Composer-family)
- Other Models — third-party models, metered at provider rates
In the editor you get a notification; the dashboard Spending tab shows live usage per pool, remaining allowance, and on-demand state. CLI builds have grown usage meters (see CLI changelog / current /help); list models with agent models or --list-models, and switch with /model or --model.
Do not confuse auth with quota:
- Confirm the account with
agent status/agent about. - On Spending (or CLI usage UI), see which pool is empty.
- If a pool remains but one model fails, it may be selection / Auto routing drawing the other pool.
- On Teams/Enterprise, admin spend limits or Impose Auto can also appear in dashboard/team settings.
Restart and reinstall do not refill a server-side quota. Symptom triage ends when you can name the blocked axis: Cursor Models, Other Models, or an on-demand cap.
What are the model-switch criteria?
One-line answer: Avoid the empty pool; pin an explicit model from a remaining pool. Do not trust banner auto-switch alone—use /model or agent --model … with an intentional ID.
Decision tree (concepts only):
| UI situation | Prefer | Avoid |
|---|---|---|
| Other Models empty only | Cursor Models pool (Composer / Grok family per docs) | Retry loops on the same third-party ID |
| Cursor Models empty only | Other Models if allowance remains; else wait / on-demand UI / shrink scope | Spamming random models |
| One frontier model limited | Auto or another model still in a live pool | Assuming Auto never touches the empty pool |
| Banner auto-switch also limited | Manual picker choice of a remaining model | Clicking the banner button in a loop |
| Read / design only | CLI Ask / Plan (--mode=ask, --plan) | Full Agent + force-run while gated |
Practical order:
- Read Spending / usage UI for the empty pool. Docs: model choice affects how fast included usage drains. Router (Auto) requests bill at the routed model and may draw from either pool.
- Pin a model from a remaining pool. In-session
/model, oragent --model <id> "…". Discover IDs withagent models/--list-models. - Do not rely on auto-fallback alone. Support/forum threads have reported banners switching to a model other than the suggested Auto, or Auto being temporarily unavailable and creating a loop. Operating rule: remaining pool → explicit
--model//model→ one verification send. - Downgrade the mode. Ask for questions, Plan for design. Agent tool/patch loops spend more round-trips for the same goal (scope/request count—not price figures).
- On-demand, upgrade, or wait are the documented UI options when included usage is gone. Usage resets on the billing cycle shown on Spending; unused allowance does not roll over. This post does not quote dollar amounts.
- BYOK is documented as a separate axis on individual plans (Teams/Enterprise have additional Token Rate rules). No key values or unit prices here.
Fallback one-liner:
1) See which pool is empty (Spending / usage UI)
2) Explicitly switch to a remaining-pool model (/model or --model)
3) If still blocked: Ask/Plan, shrink scope, or reset/on-demand UI
4) For long work: checkpoint and stop (next section)
How do you cut the work cleanly?
One-line answer: When the limit appears, do not start a new large patch. Persist resumable state on disk, then stop. Resume later with agent resume / --continue.
Clean-stop checklist:
- Stop mid-edit. Do not push another half-applied multi-file change. Prefer the last successful unit (one build/test chunk).
- Write progress in the repo (or ticket). Goal / Done / Next / Blocked / Paths—same axis as context-budget posts. Closing the CLI chat alone makes resume expensive after reset.
- Freeze git state. Commit, stash, or worktree per team rules so the agent’s reach is visible as a diff.
- Record the session id. CLI supports
agent ls,agent resume,agent --continue, and--resume=<chat-id>(CLI overview). After the limit clears, continue the same thread instead of re-briefing from zero. - Pre-write the next narrow prompt. Example: “Do only Next in
progress.md. Out of scope: …” A tight Done-when saves round-trips on the fallback model. - Cloud handoff is a different axis. Prepending
&to send work to Cloud Agent can differ in billing and runtime from local CLI quota. Do not mix it accidentally with the local fallback order (pool check → explicit model → stop/resume); choose cloud deliberately.
Resume sequence:
limit banner
→ progress + git snapshot
→ stop new requests
→ (reset / other pool / on-demand UI)
→ agent resume or --model <remaining> on a narrow Next only
Under a quota gate, prefer one verified slice + the next slice over handing an entire unfinished feature to another model in one shot.
FAQ
Spending and the CLI banner disagree?
Trust Spending’s pools, remaining allowance, and reset date. CLI banners are short and may urge Auto even when the remaining pool is elsewhere. Confirm usable IDs with the dashboard and agent models, then switch explicitly.
Does reinstalling restore quota?
No. Docs place usage on the account / billing cycle. Clearing local caches does not refill server pools.
Is switching to Ask/Plan always enough?
For read/design work it can finish the same goal in fewer tool round-trips. Implementation still needs Agent (or equivalent). Near a hard limit, Ask/Plan is safest as a way to leave a plan and stop.
Where are prices?
Out of scope here. Use Spending’s pools / on-demand / reset date, and official Pricing docs for amounts.
Sources
- Usage and limits — two pools, Spending, options at limit, monthly reset
- Models & Pricing — pool definitions and limit options (no dollar quotes in this post)
- CLI overview — agent / resume / continue / modes
- CLI parameters —
--model,--mode,--list-models,--resume - Slash commands —
/model,/ask,/plan,/resume