Saving tokens, Claude only: work in chunks, save, then compact

Claude only. Six habits (work in chunks, keep one model, return within 1 hour, ak save, compact, rest) that keep the earlier conversation in memory and add only what changed: fewer tokens, context kept. Does not apply to ChatGPT.

This page explains where tokens leak in a long session with Claude, and the order that shrinks the conversation while the content stays in AiAkiv. It is the written form of the "Saving tokens, Claude only" deck on the tips page.

This tip is Claude only. ChatGPT continues a conversation differently, so it does not carry over.

1. Why tokens leak

An LLM has two ways to continue a conversation.

  • The first rewrites the whole conversation into memory every time. The longer the conversation, the more tokens each answer costs.
  • The second keeps the earlier conversation in memory and adds only what changed at the end. That is the cache. When the next request starts with exactly what was kept, that part is re-used far more cheaply. Re-use runs from the very start up to the first difference, so a change in the middle means everything after it is read again.

One caveat: when the model changes, or the memory goes unused, the conversation kept there is gone. Using the second way wherever possible cuts tokens a great deal. The tips here help with that: fewer tokens, with the context kept.

The conversation kept in memory (the cache) has three properties.

Property What it means
It has a lifetime Unused, it expires. The default is 5 minutes, and 1 hour is an option. Claude Code uses 1 hour
Use extends it Each read of the cache extends its lifetime at no extra cost. While you keep talking, it stays
It is per model Switch model and the new model has no cache, so it reads the conversation again from the start

The way to shrink the conversation itself is compacting. In Claude Code, /compact turns the long conversation into a short summary, so from then on the AI is sent much less. The detail the summary leaves out is dropped.

So the principle fits in one line: send the AI less, but lose none of the content. Compacting sends less; the AiAkiv save is what keeps the content.

2. The six habits

Follow them in order. Habits 1 to 3 keep the cache alive; habits 4 to 6 shrink the conversation without losing what it held.

1) Work in chunks you can close, even inside one session

Split the work into pieces you can finish in one go (one feature, one investigation) and make it possible to stop at the boundary. The save and the compact below depend on having that boundary. If several jobs are mixed into one conversation, there is no clear place to save or to shrink.

2) Once a chunk starts, keep the same model

The cache is per model. Switch mid-chunk and the new model reads the whole conversation so far from the start. If you need a different model, finish the chunk, save, and compact first; the new model then reads only the short summary.

3) Stay until the chunk is done

Each use extends the cache's life, but unused it simply expires, and coming back after that means the long conversation is read again. If you do step away, be back within 1 hour. That is Claude Code; tools on the default setting give you 5 minutes.

4) Save each chunk to AiAkiv

When a chunk ends, ak save keeps it in memory. That save is why the next step can shrink the conversation without losing anything. Section 3 lists what to put in it.

5) Compact the session after saving

In Claude Code that is /compact. The long conversation becomes a short summary, so from now on the AI is sent far less. Compacting drops detail, which is why saving comes first and compacting second.

6) Now you can step away

After the compact, little is left to re-read even if the cache expires. When you come back, you start from the short summary plus what AiAkiv remembers. The next chunk starts again from habit 1.

3. What to put in the save that closes a chunk

After a compact, only the summary and the memory remain. The summary is short and loses detail, so whatever you need to restart goes into the save.

  • Decisions and reasons. What you decided and why, including any path you considered and dropped, with the reason.
  • What is left. The work to pick up in the next chunk, and anything blocked.
  • Exact names. File paths, function names, commit ids, and setting names as they are, not shortened or paraphrased. Later searches find the record by these.

Ask for the save like this:

ak save. Save this chunk (the login error fix) with the decisions and reasons, what is left, and the files changed and commit ids. Keep the summary to one line.

Then compact. After the compact, or when you come back the next day, restart from memory:

Find where we left off last time. What was left on the login error fix?

The AI looks the chunk up in AiAkiv and answers from it, including detail the compact summary no longer has. If a long chunk was saved in several parts, chain them with Chained saves so the whole thread comes back together.

4. Things to watch

  • Compacting before saving loses the detail. After the compact, decisions and names that are not in the summary are hard to get back from the conversation. Check that the save went through, then compact.
  • The 1 hour is Claude Code. Tools on the default setting lose the cache after 5 minutes. If you do not know your tool's cache lifetime, assume 5 minutes.
  • It does not carry over to ChatGPT. ChatGPT continues a conversation differently, so this tip does not fit it. The lifetimes and the compacting here are Claude.
  • When you split work across several AIs, the place to change model is a chunk boundary. The handing side leaves it with ak save, and the receiving side starts from that memory (Working with several AIs).

5. See also

Full index → Docs View as Markdown