Anthropic Explains How to Cut Claude Code Costs with Prompt Caching

Anthropic breaks down how prefill, decode, and prompt caching drive Claude Code costs, with practical tips to avoid cache misses and context bloat.

Last Updated: August 14, 2026 Editorial Process
Editorial Process
See more of Inside AI's trusted news by adding us as a preferred source on Google.
AI neural network visualization
Published on: August 14, 2026

August 15, 2026, (Inside AI) — Anthropic has published a detailed technical guide on how users can cut costs and extend session limits when working with Claude Code, its terminal-based coding assistant. The company explains that every request passes through a GPU in two distinct phases, and understanding this split is the key to controlling spending.

The first phase, called prefill, involves the model reading the entire conversation context: the system prompt, user instructions, files Claude has opened, and outputs from commands it has executed. Those are input tokens. The second phase, decode, generates output tokens one at a time. A 200-token response means the model runs 200 separate decode steps. Because decode keeps the GPU busy much longer per token, output is priced at roughly five times input.

This asymmetry matters because a large share of output tokens in any Claude Code session are thinking tokens, which the model uses to plan before responding. The effort level setting controls how much thinking occurs per turn, and like the model choice, the effort level persists as a default across sessions. Anthropic advises users to run /model and /effort once in a fresh session to see current settings. For low-stakes grunt work, the command claude can disable thinking entirely for one session, except on Fable 5, where it is not available.

The guide’s central cost mechanism is prompt caching. When a request begins with the exact same tokens as a previous one, the server reuses the computed state for that shared prefix instead of recomputing it. Cache reads cost 0.1 times the input price, while cache writes cost up to 2 times normal input. But the write happens once per token, and every subsequent turn benefits from the 0.1x read rate.

Claude Code manages prompt caching automatically, but users can inadvertently break it. For example, typing “fix the failing test in” triggers five separate requests, each containing the entire conversation history. Only new tokens in each turn are prefilled at full price; the rest are cache reads. Anthropic notes this applies to subscription users too, since the same requests draw down usage limits.

Cache invalidation occurs when anything changes toward the front of the request prefix. Switching models forces a full re-prefill at full price, including transitions in and out of plan mode with opusplan. Changing the effort level has the same effect. Fast mode is also part of the cache key, so users should enable it at the start of a session. Compacting the conversation replaces it with a shorter one, invalidating the cache, though writing the summary is cheap if the old conversation is still cached. Time is another factor: the cache expires after one hour on a subscription or five minutes on an API key, though the --cache-ttl flag can extend it to one hour. Resuming an old session almost always triggers a full re-prefill.

Anthropic recommends rewinding to just before unwanted turns instead of compacting, because rewinding only cuts turns off the end and preserves the cache for everything before them. Compacting always costs something.

Context bloat is the real budget killer

The guide emphasizes that nothing gets sent just once. Every file Claude reads or command output it sees is resent on every subsequent turn for the rest of the session. Cache reads make this cheap, but cheap is not free, and bloated context forces the model to think around irrelevant data.

Several factors determine how many tokens end up in context. Tool definitions, the system prompt, and CLAUDE.md load at startup. Users should keep CLAUDE.md specific and move workflow instructions into skills, which load only when used. Unneeded MCP servers can be disabled with /mcp.

Most context growth comes from tool results. If a user says “the tests are failing,” Claude must grep for relevant files and read several candidates, all of which stay in context. A more specific instruction like “Fix the failing test in ” skips the search. Using @-mentions attaches a file to the message before any request is sent, eliminating a Read call. However, mentioning the same file twice generally attaches a second copy.

Command output is another major source. Outputs under 30,000 characters are appended to the conversation in full. A test runner printing 400 passing lines stays under the limit and becomes part of every remaining turn. Anthropic suggests adding quiet flags to frequently used commands or using a hook to rewrite noisy commands before they run.

Subagents and session hygiene limit token drag

Long sessions cost more than the same work spread over several short ones, because turn 40 re-reads the previous 39 turns. Anthropic advises starting a new session when beginning a new task and compacting when the earlier part of a task is done. Users should run /export before clearing if they want to restore the session later, and can add a “Compact instructions” section to CLAUDE.md. For 1M-token models, the --auto-compact flag restores the auto-compact safety net in Claude Code v2.1.221 or later.

Background tasks also consume turns. A cron job fires as a full turn in the session where it was set up, carrying the entire conversation each time. If more than an hour has passed, it is also a cache miss. Anthropic recommends running loops from a fresh session in another terminal.

Subagents offer a way to keep noisy work out of the main context. Each subagent gets its own context window, system prompt, tools, and permissions, but not the main conversation. It runs its own turns and returns only its final answer. This is ideal for jobs that produce lots of disposable output, like log analysis. Users can assign a subagent definition with model: haiku or sonnet for repetitive noisy tasks, keeping the main session lean.

Anthropic’s guide concludes with four monitoring priorities, roughly ordered by cost impact: model choice, effort level, context size, and session length. The company frames the entire cost model as a function of how many tokens enter the context, how many turns they persist, and how many contexts run simultaneously.

More from Inside AI

  • AI Policy & Regulation

    US-China AI Trade Decoupling ‘Manageable’ but Tech Curbs a ‘Wild Card’: Citi

    August 14, 2026
  • AI In Business

    Bumble Abandons Women-First Rule as Dating Apps Bet on AI

    August 14, 2026
  • AI In Business

    IT Stocks Recover Most Losses on Valuation Bet and Anti-AI Trade

    August 14, 2026
  • AI In Business

    Nvidia Partners with Six Financial Giants to Raise $500 Billion for AI Infrastructure

    August 14, 2026
  • Agentic AI

    DeepSeek Launches Harness Framework for Autonomous AI Agents

    August 14, 2026
  • AI In Business

    Research: The Innovation Problems AI Can’t Solve

    August 14, 2026
  • AI In Business

    Not even the AI candyfloss economy can defy reality forever

    August 14, 2026
  • Cybersecurity AI

    Z.ai’s GLM-5.3 Nears Anthropic’s Mythos 5 in Cyber-Defence Tests

    August 14, 2026

Never Miss a Breakthrough

Join 50,000+ readers who get our daily AI intelligence briefing. No fluff, just what matters.

Inside AI is an independent publication covering artificial intelligence news, machine learning research, and the tools shaping the future of technology. No hype. Just what's happening in the AI world.

Topics

  • Artificial Intelligence
  • Machine Learning
  • Generative AI
  • Agentic AI
  • Vibe Coding
  • Prompt Engineering
  • AI Policy & Regulation
  • AI Hardware & Infrastructure
  • AI Tools
  • AI In Business
  • Robotics
  • Cybersecurity AI
  • AI Safety
  • AI Tools & Reviews (Coming soon)

Company

  • Editorial Standards
  • Privacy Policy
  • Terms of Service
  • Contact
  • About Us

Others

  • Press Releases
  • Features
  • Sponsored Content

© 2026 Inside AI. All rights reserved.

Designed by Blue Flare Digital