Stop Burning Money! How to Run Leaner Sessions in Claude Code

Stop Burning Money! How to Run Leaner Sessions in Claude Code

Written by Jeffrey Hebert

For modern developers, Claude Code has been an absolute game changer. Rather than simply suggesting code snippets, it takes real action directly from your terminal.

There is, however, one catch nobody mentions in the shiny launch videos: sloppy sessions are expensive.

Claude consumes your token limit with every terminal command, every error log, and every file it opens. When you leave messy sessions dragging on while you bounce between client projects, you’re not just wasting time — you’re basically paying full Parkway toll rates on every single prompt, over and over, just to re-feed the AI with the same history.

Whether you’re freelancing out of Asbury Park or running a dev shop servicing clients from Montclair down to South Jersey, keeping your budget intact means running a lean setup. You don’t have to let your API usage spin out of control like Mother Leeds’s 13th child haunting the Pine Barrens.

The good news? Most waste is caused by bad habits, and I’ll show you how to easily fix them.

Wait, What’s a “Token” and Why Should I Care?

Here’s a crash course, because everything else depends on it.

In AI models, chunks of text are called tokens, and they read and write them as chunks. Generally speaking, a token is part of a word. As such, a token is anything Claude sees or reads: every message you send, every file it accesses, every terminal output.

Here’s what trips people up. Claude reads more than just your latest message. Each time it opens a file, every error log, every back-and-forth from two hours ago, it reads the whole conversation again. This entire stack is known as the context window, and it acts as Claude’s short-term memory.

As your session gets longer and messier, you end up paying more for each message, despite the fact that 90% of it has nothing to do with what you’re doing right now. It’s like having a mechanic inspect your entire truck every time you top off the wiper fluid.

How Does Session Management Actually Affect Your Efficiency?

The biggest money pit is dragging old discussions into new projects. For instance, let’s say you spend an hour fixing an authentication bug, and then ask Claude to create a new landing page. Each message about CSS now carries along that entire hour of login debugging as well. You’re not only paying for it, but actually, it can make Claude worse at the new task since it’s distracted by junk that’s no longer important.

Thankfully, the fix is simple: type /clear when switching to unrelated work. In Anthropic’s own documentation, it says stale context wastes tokens and recommends clearing them when you move on to something else. The best part is that clearing costs nothing. Honestly, there’s no reason not to do it.

Worried about losing your place? Don’t be. Run /rename before clearing so the session is easy to find later, then use /resume to jump back into it.

There’s also /compact, which summarizes the conversation and keeps the important stuff if you don’t want to completely clear the slate (say you’re still working on the same project). You can even tell it what to preserve: “/compact Focus on code samples and API usage” tells Claude what to preserve when it summarizes.

But be aware that it is not free. Because compacting reads the whole conversation it sums up, it is a big request to do so on a huge context.

In short, if you’re working on the same task and things are slowing down? /compact. Switching gears to another task? /clear. It’s as simple as that.

Running a team? In Anthropic’s own breakdown, the cause of surprise API bills is almost always one of two things: massive sessions that nobody cleared, or leaving heavyweight models like Opus running for basic grunt work. This brings us to our next point.

Are You Using a Sledgehammer to Hang a Picture?

Claude Code lets you choose which model to use. For gnarly architecture problems and hard debugging, the bigger, smarter models are incredible. They also cost more. It’s like using a structural engineer to hang drywall when you use the top-tier model to rename variables or write basic forms.

To put it another way, choose the right model for the job. Keep the heavy hitters for the heavy lifting, and let the faster, cheaper models handle the routine tasks. As one of the two highest-impact habits to teach your team, Anthropic recommends separating unrelated tasks and matching the model to the job.

What Are the Best Ways to Keep Your Context Clean?

As soon as you’ve mastered clearing and model selection, let’s talk about preventing sessions from getting bloated.

  • Point Claude straight at the file. Use @ mentions to attach the exact file you’re talking about. Otherwise, Claude might explore your project to find it, opening every file it finds along the way. Rather than scavenger hunts, tell it exactly where to look.
  • Shut up the noisy commands. There are some terminal commands that generate mountains of text: installation logs, verbose test runs, and build output. All of that gets thrown into the conversation. To make sure Claude only sees what’s relevant, use quiet flags.
  • Let hooks do the dirty work. This one is a little more advanced, but it changes everything. With a custom hook, instead of Claude reading a 10,000-line log, he can grep for “ERROR” and return only the matching lines, greatly reducing the context.
  • Hand the messy stuff to subagents. Subagents are simply helpers Claude can assign to do a side job in their own workspace. When they’re done crunching through output or digging through files, they report back with a summary. The mess stays in their context, not yours.
  • Keep your CLAUDE.md lean. Your CLAUDE.md file contains Claude’s standing instructions: how to code, how the project is laid out, that sort of thing. It’s awesome. The problem is, it loads continuously, so each line you put in there costs you repeatedly. As such, keep it to the essentials. Additionally, you can add compaction instructions to CLAUDE.md, telling Claude what to focus on (like test output and code changes).

How Does Prompt Caching Affect What You Pay?

It’s here that many people lose money without knowing why.

There’s something called prompt caching in Claude Code. As long as your conversation starts the same way from message to message, the system can save and reuse it at a big discount instead of reprocessing it all. It saves you a lot of time when it works.

There’s just one problem: it’s fragile. When you switch models mid-conversation or bump up or down the effort setting up, the cache gets thrown out. As a result, the whole conversation is reprocessed at full price. Do that a few times during a long refactoring session and you’ll feel it.

Timing is also important. As described in this cost guide, Claude Code’s cache lasts approximately five minutes after your last message, so if you’ve stepped away longer than that, you can use /compact to reprocess everything at full price, and /clear to start over.

Here are some practical takeaways:

  • Pick your model and settings at the start of a task and leave them alone. You should switch when you’re starting something new anyway, right after a /clear.
  • Coming back from lunch? Start fresh rather than compiling stale sessions.
  • If you’re building your own tools on the API, keep your system prompt stable and put anything that changes (timestamps, user data, whatever) at the end, not at the beginning. Every time something changes at the top, the cache is broken.

How Do You Keep an Eye on All This?

The only way to fix something is to measure it. Claude Code comes with built-in commands to help you see where your tokens are going: /context shows you what’s taking up space in the current session, and /usage shows how much you’ve spent. You should check them out. Make it a habit. If you’re managing a team, make sure everyone is doing it as well.

For the full rundown straight from the source, read Anthropic’s official cost management guide for Claude Code. Unlike most documentation, it’s actually readable.

The Bottom Line

Right now, Claude Code is one of the most useful tools for small dev shops and freelancers. But it’s not magic, and it’s not free. The people who benefit most from it aren’t those who throw the biggest model at every problem and leave sessions open for three days. They’re the ones who work clean: clear between jobs, pick the right tool for the task, keep junk out of the conversation, and watch the meter.

Think of your context window as a truck bed. Haul what you need for the job, then clean it out before the next one.

Claude Code Quick Reference

  • /clear: Starting a new job that has nothing to do with the previous one. It wipes the slate clean and costs nothing. Don’t be afraid to use it more than you think you should.
  • /compact: It’s the same job, but the session is getting bloated. Keeps the important details and summarizes the history.
  • /compact [notes]: Same as above, but you tell it what to hang onto. Example: “/compact focus on the checkout bug.”
  • /rename + /resume: If you name a session before clearing it, you can pick it up again without having to start from scratch.
  • /model: Make sure the model is right for the job. For hard problems, use the big model; for grunt work, use the cheaper model. You should switch at the beginning of a task, not in the middle.
  • @filename: Don’t let Claude dig through your entire project; just point it at the file you mean.
  • /context: Analyze your current session to see what’s taking up space.
  • /usage: Check how much you’ve burned. Make it a habit.

The golden rule: New job? Clear. Same job, getting heavy? Compact. Back from a break? Clear. Don’t swap models mid-task. Keep CLAUDE.md lean.