Skip to content
← Back to News

Engineering11 min read

Session memory without dragging the whole transcript around

The obvious way to give an agent memory is to paste the conversation back in every turn. It works until it doesn't, and it fails on the day you least want it to. Here is what to do instead.

Almost every agent starts the same way: keep the messages in a list, send the list on every turn. It is the right thing to build first — it is three lines and it works.

What it is not is a design. It has a cost curve nobody chose, a failure mode that arrives without warning, and a privacy property that will be a problem the first time someone asks about it.

1. What concatenating actually costs

Three things grow at once, and they compound rather than add:

  • Tokens per turn. Turn twenty does not cost what turn one cost. It carries the previous nineteen. Cost per conversation grows with the square of its length, not with its length.
  • Latency. Time to first token rises with the prompt. The conversation gets slower exactly as it gets more valuable — that is, as it approaches a decision.
  • Memory in the process. Every turn held in the execution is memory the process cannot release while the conversation is open.

The version of this that hits hardest is the one that runs out of process memory rather than context. If you have seen a worker die mid-conversation, that is the same design showing a different symptom — we wrote that one up separately.

2. Three layers, because they answer different questions

«Memory» is one word for three jobs. Keeping them apart is most of the work, because each one has a different lifetime, a different size and a different reason to exist:

LayerWhat it holdsLifetimeWhat it answers
WorkingThe current turn and the few before itThe turnWhat are we talking about right now?
OperationalThe facts the conversation has established: product, quantity, destination, payment methodThe sessionWhat have we already agreed, so I don't ask twice?
Long-termThe documents, the catalogue, the history that predates this conversationPermanentWhat do we know that this conversation never said?
Most systems that feel like they have a bad memory have collapsed the middle one into the first.

The operational layer is the one that earns its keep. It is small, it is structured, and it is what lets a system say «so, the two you wanted, to the same address?» on turn fifteen without re-reading turns one through fourteen.

3. A summary that survives the window

Between the transcript and the facts there is a third thing worth keeping: a running summary of the thread. Not the messages, and not a list of fields — the state of the conversation in a few sentences.

It matters because of what happens when the window trims. If the only record is the transcript, trimming loses whatever fell off the front. If there is a consolidated summary kept outside the transcript, trimming costs you the wording and not the meaning. That summary is also what should go into a retrieval query — searching your documents with the whole conversation pasted in retrieves noise.

4. Where the session should live

If session state lives inside the process handling the turn, it dies with a restart and it grows with traffic. Outside the process it needs three things, and they are not optional:

  1. An expiry that slides. A conversation that has been idle for an hour is over. One that is active should not expire mid-sentence because it started an hour ago.
  2. An eviction rule. Memory is finite even when it is on disk. Dropping the least recently used session is boring and correct.
  3. A boundary per customer. One tenant's session data is not reachable from another's, and that is enforced where the data is read, not by convention.

There is a fourth thing that is easy to defer and expensive to retrofit: what you persist is not what you hold in memory. A live session can hold a phone number; a stored one should hold a label that points at it. Keep the operational facts — the product, the quantity, the destination — because those are what let you rebuild a quote. Pseudonymise the rest on the way to disk.

5. Anti-patterns

  • Summarising every turn. You pay a model call per turn to save tokens on the next one. Do the arithmetic before you assume it wins.
  • Retrieving with the whole conversation as the query. Long queries retrieve vaguely. Use the summary, or the turn, not the transcript.
  • One vector index for every tenant, filtered after the search. The filter is not the boundary; the index should be.
  • Storing the transcript forever because it might be useful. It will be useful to whoever breaches you. Decide the retention before you need it.
  • Treating the summary as authoritative. It is a compression of what was said, not a record of what was agreed. The operational facts are the record.

6. Where to start

If you have the concatenating version running, you do not need to rebuild it. The order that gets you the most for the least is:

  1. Pull the established facts out of the transcript into a small structured object. You will find the agent stops asking twice.
  2. Stop sending the full history; send the facts plus the last few turns.
  3. Add the running summary, written off the critical path.
  4. Move the session out of the process, with expiry and eviction.
  5. Decide what gets pseudonymised on the way to storage, and do it before you have a year of transcripts.

Frequently asked questions

Why not just use a bigger context window?

It moves the wall without changing the slope. Cost and latency still grow with every turn, and you still lose the beginning of the conversation when you eventually cross the line — just later, and with a bigger bill.

Isn't a summary lossy?

Yes, and that is why it is not the record. Keep the established facts as structured data — those are authoritative — and let the summary carry the tone and the context around them.

How long should a session live?

Long enough that a customer who steps away comes back to the same conversation, short enough that idle sessions are not accumulating. A sliding expiry measured in tens of minutes, with eviction of the least recently used, handles both.

Where does RAG fit in this?

It is the long-term layer, and it should be queried with the summary or the current turn — not with the transcript. Retrieval quality falls as the query gets longer and vaguer.

Keep reading