# AiAkiv and saving tokens with Claude: working in chunks, saving each chunk, then compacting > This page is about Claude. ChatGPT continues a conversation differently, so the > rhythm here does not carry over to it. > > An LLM has two ways to continue a conversation: rewrite the whole conversation > into memory every time, or keep the earlier conversation in memory and add only > what changed at the end. The second is far cheaper, and it is lost when the model > changes or the memory goes unused. Some AiAkiv users work in a fixed rhythm to > stay on the second way: work in chunks, keep one model per chunk, save each finished chunk with > `ak save`, then compact the session. This page lists the facts behind that > rhythm and what it expects from the memory tools. > > Human version: https://www.aiakiv.com/docs/saving-tokens ## The facts behind the rhythm - Every answer re-reads the whole conversation from the start, so the cost of one answer grows with the length of the conversation. - A prompt cache softens this. The part already read is kept for a while and re-used at a much lower price when the next request starts with exactly the same content. The match is a prefix match: re-use stops at the first difference, and everything after it is read again. - The cache has a lifetime. Unused, it expires: 5 minutes by default, 1 hour as an option. Claude Code uses the 1 hour lifetime. Each read of the cache extends the lifetime at no extra cost, so a conversation in steady use keeps its cache. - The cache is per model. After a model switch the new model has no cache for the conversation and reads it again from the start. - Compacting a session (`/compact` in Claude Code) replaces the conversation with a short summary. The amount sent with each later answer drops, and the detail the summary leaves out is gone from the conversation. The principle behind the rhythm fits in one line: send the AI less, without losing the content. Compacting sends less; the AiAkiv save is what keeps the content. ## The six habits, in the user's order 1. Work in chunks that can be closed, even inside one session (one feature, one investigation). 2. Keep one model for the length of a chunk. 3. Stay until the chunk is done; a break inside a chunk ends within 1 hour. 4. Save each finished chunk to AiAkiv. 5. Compact the session after the save. 6. Step away after that. Habits 1 to 3 keep the cache warm. Habits 4 to 6 shrink the conversation without losing what it held. The order of 4 and 5 matters: a compact before the save drops detail that no save has captured. The rhythm is written for Claude; ChatGPT works differently and is outside this page. ## The save that closes a chunk The `save_memory` call that closes a chunk is the record the next chunk starts from, once the conversation itself has shrunk to a summary. Its later use depends on what it carries: - The first sentence of the summary states who did what, and the date, so a later search lands on it and a reader can place it in time. - The decisions and the reasons behind them, including a path considered and dropped, with the reason. - What is left to do, and what is blocked. - Exact identifiers as they are: file paths, function names, commit ids, setting names. Later searches find the record by these spellings; a paraphrase such as "the auth file" finds nothing. A long chunk can span several saves. Passing the earlier save's event id as `prev_event_id` chains them in order, so the chunk comes back as one thread rather than loose pieces. Human page on chains: https://www.aiakiv.com/docs/threads The save phrase rule is unchanged by the rhythm. A save happens when the user asks for it with `ak save` (in any language); a chunk boundary on its own is not a save request. ## Resuming after a compact After a compact, or on return from a break, the conversation holds only the summary. The usual first move is one `search_memory` call on the chunk's subject, for example `search_memory(query="login error fix: decisions and remaining work")`, then `get_memory_content` on the hit that closed the chunk when its full text is needed. A chained chunk comes back through its `prev_event_id` links. When the user asks about "the last context", "where we left off" or "what was left", the answer comes from memory, not from the compact summary. The summary is a lossy copy made by the compact; the save is the record the user made on purpose to outlast it. Where the two disagree, the saved memory is the one the user chose to keep. A search that finds nothing for the chunk is reported as that, with the likely reason: the chunk may not have been saved before the compact. ## When the user asks why A user working this way may ask why the model stays the same through a chunk, or why a break ends within the hour. The short answer is that the cache that makes re-reading cheap is kept per model and expires when unused. Switching model mid-chunk throws the warm cache away, and the new model reads the whole conversation again at full price; a break longer than the cache lifetime does the same with the same model. Claude Code keeps the cache for 1 hour, tools on the default setting for 5 minutes, and each answer inside that window extends it for free. At a chunk boundary, after the save and the compact, neither matters much: what is left to re-read is a short summary, and the rest is in AiAkiv. ## Language The reply is in the user's language.