Code Kit 5.7 is out now, rebuilt for the Claude 5 family. Includes access to our MCP: serving up our entire blog for your Claude to analyze.
Claude FastClaude Fast
Performance

Claude Code Fast Mode: 2.5x Faster Opus 5.5 at $8/$40

Claude Code fast mode runs Opus 5.5 up to 2.5x faster at $8/$40 per MTok. Supported models, /fast setup, real costs, and when it pays off.

Stop configuring. Start shipping.Everything you're reading about and more..
Agentic Orchestration Kit for Claude Code.

Claude Code fast mode makes Opus generate output up to 2.5x faster, and on Claude Opus 5.5 it costs $8 per million input tokens and $40 per million output, exactly twice the standard $4/$20 rate. Type /fast in Claude Code, press Space, then Enter. Since Claude Code v2.1.280, fast mode runs on Opus 5.5 by default. It is the same model with the same quality. You pay more per token for a faster serving path.

Fast mode is supported on three models today: Opus 5.5, Opus 5, and Opus 4.8. It is not available on Sonnet, Haiku, Fable 5.1, Opus 4.7, or Opus 4.6. The premium also changed shape since launch. Opus 4.6 fast mode settled at six times the standard rate in February 2026. Opus 4.8 cut it to two times the standard rate in May, and Opus 5.5 keeps 2x on a base price 20% lower than Opus 5, which moves the break-even point for turning it on.

Claude Code Fast Mode at a Glance

QuestionAnswer
SpeedupUp to 2.5x higher output tokens per second. Time to first token does not improve
SupportedOpus 5.5 (default since v2.1.280), Opus 5, Opus 4.8
Price, Opus 5.5$8 input / $40 output per MTok, 2x standard, flat across the 1M context window
Price, Opus 5$10 / $50 per MTok (Opus 4.8 is the same)
Toggle/fast, then Space and Enter. Or "fastMode": true in settings
SurfacesClaude API and Claude subscriptions only. Not Bedrock, Google Cloud, Foundry, or Claude Platform on AWS
BillingSubscriptions pay from usage credits, never from included plan usage. Console bills per token
StatusResearch preview. Pricing and availability can change

Sources for every row: the Claude Code fast mode docs, the API fast mode docs, and the pricing page.

What Fast Mode Actually Is

Fast mode does not switch you to a different model. Anthropic's API docs describe it as "the same model with a faster inference configuration. There is no change to intelligence or capabilities." You keep the same weights, the same reasoning, and the same output. The request is served on a path tuned for speed instead of cost efficiency.

The speedup is specific. Fast mode raises output tokens per second. It does not shorten time to first token, and it does nothing for time spent running tools. A turn that streams 3,000 tokens of code gets most of the gain. A turn that spends 40 seconds in npm test and writes two lines gets almost none.

This is different from lowering your effort level, which makes Claude think less. Lower effort buys speed with reasoning depth. Fast mode buys speed with money and leaves the reasoning alone.

Which Models Support Fast Mode

ModelFast mode price (per MTok)Standard priceFast mode status
Opus 5.5$8 / $40$4 / $20Supported, Claude Code default since v2.1.280
Opus 5$10 / $50$5 / $25Supported
Opus 4.8$10 / $50$5 / $25Supported
Opus 4.7none$5 / $25Removed July 24, 2026. API requests with speed: "fast" return an error
Opus 4.6none$5 / $25Not available. API requests silently run at standard speed and price
Fable 5.1none$10 / $50Not offered
Sonnet, HaikunonevariesNot offered

All supported models share one fast mode rate limit pool, so heavy use on Opus 5 draws down the same budget as Opus 5.5.

The default fast mode model has moved with each Opus release. Claude Code v2.1.142 through v2.1.153 used Opus 4.7, v2.1.154 through v2.1.218 used Opus 4.8, v2.1.219 moved to Opus 5, and v2.1.280 moved to Opus 5.5. If you pinned an older Opus in /model, fast mode runs on that model as long as it is on the supported list.

How to Enable Fast Mode in Claude Code

In the CLI, run:

/fast

Press Space to toggle, then Enter to confirm. Claude Code shows "Fast mode ON" and a small ↯ icon next to the prompt. Run /fast again to check the state or turn it off.

To keep it on by default, set it in ~/.claude/settings.json:

{
  "fastMode": true
}

The settings reference lists the related keys. Six behaviors are worth knowing before you rely on it:

  • It persists. Fast mode you turn on in an interactive session stays on in later sessions unless an admin requires per-session opt-in.
  • It switches you to Opus. If your current model does not support fast mode (Sonnet, Haiku, Fable, Opus 4.7), turning it on moves you to Opus.
  • It follows model switches. Switching with /model to an unsupported model turns fast mode off. Switching back to a supported Opus turns it on again if your saved preference is on.
  • Turning it off keeps Opus. /fast off leaves you on Opus. Use /model to change models.
  • It toggles mid-turn. You can run /fast while Claude is working. The running turn finishes at its original speed and the change applies from your next turn.
  • Other surfaces. The VS Code extension has a Toggle fast mode command. In cloud sessions (Claude Code on the web and self-hosted runners, v2.1.271 or later), type /fast on, which applies to that session only.

Fast Mode in -p Mode and the Agent SDK

In non-interactive mode, /fast only works in a session launched with fast mode in its --settings value. Anywhere else, the command reports that fast mode is not available. The same applies to Agent SDK sessions, which run non-interactively. Requires Claude Code v2.1.205 or later:

claude -p --settings '{"fastMode": true}' "Refactor src/auth to use dependency injection"

On the API, fast mode is a request parameter. Set speed: "fast" with the fast-mode-2026-02-01 beta header, and read usage.speed in the response to confirm which speed you got:

response = client.beta.messages.create(
    model="claude-opus-5-5",
    max_tokens=4096,
    speed="fast",
    betas=["fast-mode-2026-02-01"],
    messages=[{"role": "user", "content": "Refactor this module to use dependency injection"}],
)
print(response.usage.speed)  # "fast" or "standard"

API access is a research preview that must be provisioned for your organization first, through your account manager or the fast mode waitlist.

Fast Mode Pricing on Opus 5.5

Every line of the bill doubles. Fast mode prices are a flat multiplier on the standard rate, and prompt caching multipliers apply on top of them. On Opus 5.5, cache reads cost 0.05x base input, so by our arithmetic a fast mode cache read is $0.40 per MTok against $0.20 at standard speed. There is no long-context surcharge: fast mode pricing is flat across the full 1M window. Fast mode does not work with the Batch API or a Priority Tier commitment.

Two comparisons put the price in context:

  • Against launch pricing. Opus 4.6 fast mode launched in February 2026 with a 50% introductory discount through February 16, then settled at $30/$150 per MTok for prompts under 200K tokens and $60/$225 above that, six times the standard Opus rate. Anthropic's Opus 4.8 release on May 28, 2026 made it "three times cheaper" at $10/$50. Opus 5.5 fast mode is $8/$40 with no context tier. The same interactive session costs about a quarter of what it did in February 2026 at matched token counts.
  • Against Fable 5.1. Opus 5.5 in fast mode ($8/$40) costs less per token than Fable 5.1 at standard speed ($10/$50), and Opus 5.5 leads Fable 5.1 on all nine rows of Anthropic's launch table.

The Mid-Session Switch Cost

The first time you enable fast mode in a conversation, you pay the full fast mode uncached input price for the entire existing context. Requests at different speeds do not share cached prefixes, so the cache you built at standard speed does not carry over.

Worked example on Opus 5.5 with 150K tokens already in context:

ScenarioCost of re-reading that context
Next turn at standard speed (cache read, $0.20)$0.03
First fast mode turn (uncached fast input, $8)$1.20

The $1.20 is paid once per conversation. Toggling off and on later does not repeat it. It still means the cheapest time to turn fast mode on is the first message, before the context window fills.

Where Fast Mode Spend Shows Up

Run /status to see how you are billed. On Pro and Max, fast mode draws from usage credits, even when your plan still has included usage left, and the Usage credits section of claude.ai Settings > Usage shows the monthly total. On Team and Enterprise, run /usage to see your own usage-credits spend. On Claude Console, group the Usage and Cost pages by Speed (Research Preview) to split fast mode from standard usage.

When Fast Mode Pays Off

At 2.5x output speed, output generation takes 40% as long, so a long streamed answer finishes up to 60% sooner. At a 2x price, the question is whether that saved wait is worth one more copy of the turn's cost. Our rule: turn it on when you are watching the output and the turn is mostly generation.

Fast mode pays during:

  • Rapid iteration. Change, run, ask for an adjustment, repeat. The saving applies on every exchange, and at 2x instead of 6x, a full afternoon of iteration no longer needs a deadline to justify it.
  • Live debugging. Reading stack traces and proposed fixes as they stream keeps you on the problem instead of waiting on the next paragraph.
  • Long generated outputs you review immediately. A new component, a migration script, a test file. These turns are nearly all output tokens, which is exactly what fast mode accelerates.
  • Interactive planning. Back-and-forth design discussion where each reply sets up your next question.

Skip fast mode for:

  • Long autonomous runs. If you start a refactor and walk away, you do not see the speed. Standard pricing gets the same result.
  • Tool-heavy turns. When most of a turn is test runs, builds, or searches, fast mode speeds up only the small generated part.
  • CI, headless batch jobs, and scheduled runs. Nobody is waiting. On the API, the Batch API is half price and fast mode is unavailable there anyway.
  • Large analysis passes on a budget. Output quality is identical, so paying double buys nothing when latency does not matter.

Fast Mode vs Effort Level vs Model

Claude Code gives you three independent speed controls. They combine freely.

LeverWhat it changesCost effect
Fast modeOutput tokens per second on the same Opus model2x per token on Opus 5.5
EffortHow much Claude thinks before and between actionsFewer tokens at lower effort
ModelCapability tier (Opus, Sonnet, Haiku, Fable)Set by the model's rate card

Opus 5.5 changes how these combine. Its default effort on the API is medium, and at medium it beats Opus 5 at max on Terminal-Bench 4.0 for about a fifth of the cost per attempt, per Anthropic's launch charts. Anthropic also reports Opus 5.5 output is more than 30% faster than Opus 5 before fast mode is involved. Our default for interactive work is therefore Opus 5.5 at medium with fast mode on. Our Opus 5.5 best practices guide covers which effort level to use per task, and model versus effort covers when to change the dial instead of the model.

For straightforward work such as formatting, boilerplate, or small refactors, fast mode plus low effort is the quickest combination. For architecture and hard debugging, keep effort at high or above and let fast mode handle the latency. The speed optimization guide covers the non-billing ways to cut wait time.

Rate Limits and Fallback

Fast mode has its own rate limits, separate from standard Opus. When you hit them in Claude Code:

  1. Fast mode falls back to standard speed automatically.
  2. The ↯ icon turns gray to show the cooldown.
  3. You keep working at standard speed and standard pricing.
  4. Fast mode re-enables on its own when the cooldown ends.

Run /fast to turn it off instead of waiting. If you run out of usage credits mid-session, Claude Code retries each rejected fast request at standard speed, shows "Fast mode disabled · usage credits exhausted", and turns fast mode off for the rest of that session without changing your saved preference.

Requirements and Availability

  • Anthropic direct only. Fast mode runs on the Claude API and Claude subscription plans. It is not available on Amazon Bedrock, Google Cloud, Microsoft Foundry, or Claude Platform on AWS.
  • Usage credits on subscriptions. Pro, Max, Team, and Enterprise accounts need usage credits turned on. On Pro and Max, enable them under claude.ai Settings > Usage or run /usage-credits. On Team and Enterprise, a member with billing access enables them in Admin settings > Usage.
  • Owner enablement on Team and Enterprise. Fast mode is off by default for these organizations. An Owner enables it at Admin Settings > Claude Code.
  • Provisioned access on Console. An admin enables fast mode in Claude Code preferences, and the organization needs fast mode access provisioned. Without it, every fast request returns a 429 that Claude Code treats as a rate limit.

Fast Mode Error Messages and Fixes

Message from /fastCauseFix
"Fast mode has been disabled by your organization."Not enabled for the org, or managed settings set fastMode: false or require per-session opt-inAsk a Team/Enterprise Owner to enable it. Per-session opt-in allows /fast on only in an interactive terminal
"Fast mode requires usage credits"Usage credits are off on a subscription planTurn them on in Settings > Usage, or run /usage-credits
"Fast mode unavailable during evaluation. Please purchase credits."Console account on the free Evaluation planPurchase credits in Console billing settings
"Fast mode unavailable due to network connectivity issues"The availability check to api.anthropic.com failed, common behind LLM gatewaysAllowlist api.anthropic.com, or set CLAUDE_CODE_SKIP_FAST_MODE_NETWORK_ERRORS=1
"is not in your organization's allowed models"The availableModels allowlist excludes the fast mode Opus modelSwitch to an allowed Opus model that supports fast mode, then run /fast
Fast mode "not available" in -p or the Agent SDKThe session was not launched with fast mode in --settingsLaunch with --settings '{"fastMode": true}' on v2.1.205 or later

Admin Controls

Two settings cover cost control. fastModePerSessionOptIn makes every session start with fast mode off, so nobody leaves it running across five parallel terminals by accident:

{
  "fastModePerSessionOptIn": true
}

Owners on Team or Enterprise can deploy that key organization-wide through server-managed settings. To remove fast mode entirely on a machine, set CLAUDE_CODE_DISABLE_FAST_MODE=1.

Fast Mode with Agent Teams and Subagents

A teammate's model and fast mode are fixed when it spawns. /fast and /model only change the lead's settings, and typing either while viewing a teammate shows a notice saying so. In practice, turn fast mode on for the lead before you spawn a team if you want the whole team fast, or leave teammates at standard speed and keep fast mode for the lead's interactive coordination. Teammates run long autonomous tasks where the speedup is least visible, so the second pattern usually costs less. The agent teams guide covers the rest of the setup.

Frequently Asked Questions

Does fast mode make Claude less accurate? No. It runs the same model weights with a faster inference configuration. Anthropic states there is no change to intelligence or capabilities.

Which models support Claude fast mode? Opus 5.5, Opus 5, and Opus 4.8. Sonnet, Haiku, Fable 5.1, Opus 4.7, and Opus 4.6 do not support it.

How much does fast mode cost? $8/$40 per million input/output tokens on Opus 5.5 and $10/$50 on Opus 5 and Opus 4.8. Both are 2x the standard rate for that model.

Is fast mode included in Pro or Max? It is available on Pro, Max, Team, and Enterprise, but it always bills to usage credits, never to your plan's included usage.

Does fast mode work on Bedrock or Google Cloud? No. It is limited to the Claude API and Claude subscriptions.

Does fast mode reduce time to first token? No. It raises output tokens per second. The first token arrives no sooner.

The Short Version

Fast mode used to be a deadline tool at six times the price. Opus 4.8 brought the premium down to 2x, and on Opus 5.5 it is that same 2x on a model that already costs 20% less than Opus 5, which makes it a reasonable default for any session where you sit and watch the output. Turn it on at the start of the session, keep it off for runs you walk away from, and watch your usage credits for the first week.

If you want model, effort, and speed choices made per task instead of per session, ClaudeFast's Code Kit ships complexity routing that sends mechanical work to cheaper tiers and saves full Opus reasoning for the work that needs it.

Last updated on