CLI Quota Exhausted: How Do You Order the Fallback?

When Cursor CLI / Agent hits a wall, requests fail and banners talk about usage limits or ask you to switch to Auto or another model. Reinstalling does not fix it. The useful order is: see which pool is empty, explicitly pick a model in a remaining pool, or cut the work cleanly and resume later.

This post covers only symptoms · model-switch criteria · how to stop cleanly. Grounded in Usage and limits, CLI overview, CLI parameters, and Slash commands. No plan prices or dollar amounts—only quota UI concepts (pools, remaining allowance, reset date, on-demand toggle). No invented personal reviews.

What are the symptoms?

One-line answer: If the CLI or editor shows a usage / rate-limit notice and the same model keeps refusing, treat it as an account-side pool first—not a local install bug or a silent logout.

Common signals:

SignalLikely meaningDo first
“You’ve hit your usage limit” / “You’re out of usage”Included usage (a pool) exhausted, or that model is blockedCheck Spending or the CLI usage UI
Banner urging Auto or another modelCurrent model/pool is blockedIdentify which pool, then switch explicitly
Same prompt fails on every retryServer refusal more likely than flaky Wi‑FiPrefer pool/model checks over reinstall
agent starts but the first reply is refusedSession alive; requests gated/model, --model, agent models

Official docs: most paid plans expose two monthly usage pools:

  • Cursor Models — first-party Cursor models (docs cite Grok-family and Composer-family)
  • Other Models — third-party models, metered at provider rates

In the editor you get a notification; the dashboard Spending tab shows live usage per pool, remaining allowance, and on-demand state. CLI builds have grown usage meters (see CLI changelog / current /help); list models with agent models or --list-models, and switch with /model or --model.

Do not confuse auth with quota:

  1. Confirm the account with agent status / agent about.
  2. On Spending (or CLI usage UI), see which pool is empty.
  3. If a pool remains but one model fails, it may be selection / Auto routing drawing the other pool.
  4. On Teams/Enterprise, admin spend limits or Impose Auto can also appear in dashboard/team settings.

Restart and reinstall do not refill a server-side quota. Symptom triage ends when you can name the blocked axis: Cursor Models, Other Models, or an on-demand cap.

What are the model-switch criteria?

One-line answer: Avoid the empty pool; pin an explicit model from a remaining pool. Do not trust banner auto-switch alone—use /model or agent --model … with an intentional ID.

Decision tree (concepts only):

UI situationPreferAvoid
Other Models empty onlyCursor Models pool (Composer / Grok family per docs)Retry loops on the same third-party ID
Cursor Models empty onlyOther Models if allowance remains; else wait / on-demand UI / shrink scopeSpamming random models
One frontier model limitedAuto or another model still in a live poolAssuming Auto never touches the empty pool
Banner auto-switch also limitedManual picker choice of a remaining modelClicking the banner button in a loop
Read / design onlyCLI Ask / Plan (--mode=ask, --plan)Full Agent + force-run while gated

Practical order:

  1. Read Spending / usage UI for the empty pool. Docs: model choice affects how fast included usage drains. Router (Auto) requests bill at the routed model and may draw from either pool.
  2. Pin a model from a remaining pool. In-session /model, or agent --model <id> "…". Discover IDs with agent models / --list-models.
  3. Do not rely on auto-fallback alone. Support/forum threads have reported banners switching to a model other than the suggested Auto, or Auto being temporarily unavailable and creating a loop. Operating rule: remaining pool → explicit --model//model → one verification send.
  4. Downgrade the mode. Ask for questions, Plan for design. Agent tool/patch loops spend more round-trips for the same goal (scope/request count—not price figures).
  5. On-demand, upgrade, or wait are the documented UI options when included usage is gone. Usage resets on the billing cycle shown on Spending; unused allowance does not roll over. This post does not quote dollar amounts.
  6. BYOK is documented as a separate axis on individual plans (Teams/Enterprise have additional Token Rate rules). No key values or unit prices here.

Fallback one-liner:

1) See which pool is empty (Spending / usage UI)
2) Explicitly switch to a remaining-pool model (/model or --model)
3) If still blocked: Ask/Plan, shrink scope, or reset/on-demand UI
4) For long work: checkpoint and stop (next section)

How do you cut the work cleanly?

One-line answer: When the limit appears, do not start a new large patch. Persist resumable state on disk, then stop. Resume later with agent resume / --continue.

Clean-stop checklist:

  1. Stop mid-edit. Do not push another half-applied multi-file change. Prefer the last successful unit (one build/test chunk).
  2. Write progress in the repo (or ticket). Goal / Done / Next / Blocked / Paths—same axis as context-budget posts. Closing the CLI chat alone makes resume expensive after reset.
  3. Freeze git state. Commit, stash, or worktree per team rules so the agent’s reach is visible as a diff.
  4. Record the session id. CLI supports agent ls, agent resume, agent --continue, and --resume=<chat-id> (CLI overview). After the limit clears, continue the same thread instead of re-briefing from zero.
  5. Pre-write the next narrow prompt. Example: “Do only Next in progress.md. Out of scope: …” A tight Done-when saves round-trips on the fallback model.
  6. Cloud handoff is a different axis. Prepending & to send work to Cloud Agent can differ in billing and runtime from local CLI quota. Do not mix it accidentally with the local fallback order (pool check → explicit model → stop/resume); choose cloud deliberately.

Resume sequence:

limit banner
  → progress + git snapshot
  → stop new requests
  → (reset / other pool / on-demand UI)
  → agent resume or --model <remaining> on a narrow Next only

Under a quota gate, prefer one verified slice + the next slice over handing an entire unfinished feature to another model in one shot.

FAQ

Spending and the CLI banner disagree?

Trust Spending’s pools, remaining allowance, and reset date. CLI banners are short and may urge Auto even when the remaining pool is elsewhere. Confirm usable IDs with the dashboard and agent models, then switch explicitly.

Does reinstalling restore quota?

No. Docs place usage on the account / billing cycle. Clearing local caches does not refill server pools.

Is switching to Ask/Plan always enough?

For read/design work it can finish the same goal in fewer tool round-trips. Implementation still needs Agent (or equivalent). Near a hard limit, Ask/Plan is safest as a way to leave a plan and stop.

Where are prices?

Out of scope here. Use Spending’s pools / on-demand / reset date, and official Pricing docs for amounts.

Sources