News · New this week

How ChatGPT remembers your chat: stateless models explained

The model remembers nothing between messages. The app around it stores the history and rebuilds the prompt on every turn.

If an interviewer asks how ChatGPT remembers your conversation, the short answer is that in a typical setup, the model doesn't. The app does.

That's the twist worth knowing. The model remembers nothing between messages, so everything that looks like memory is built outside it.

What happens when you send a message

Your browser sends the message to a chat server. The server knows very little. It only knows which conversation you're in.

Using that conversation ID, it fetches the past turns from a normal database. The history lives in that history DB, keyed by conversation ID, outside the model. Nothing exotic, just stored rows that get read back.

Next comes a prompt builder. It stacks three things into one prompt: your saved memories, the conversation history and your new message.

Saved memories are a separate thing from history. They're facts the app saved earlier, kept in a per-user store, and they get pasted into the prompt as text. The model sees them the same way it sees anything else you typed.

Turn 2 resends turn 1

Once the prompt is built, the full thing goes to the LLM. Every turn, the history is sent again.

Take the example from the reel. You tell the assistant your name is Sam, and it replies with a greeting to Sam. Then you ask what your name is. The model can answer only because the prompt for that second turn contains the first exchange along with the new question. It isn't recalling anything. It's reading it.

The reply goes back to you and also into the history DB, so the next turn has it available. Then the cycle repeats. The model is stateless, and the app remembers.

Why the app can't just send everything

Every model has a token limit. More tokens also cost more and run slower, so even without a hard limit there's a reason to be selective.

When the assembled prompt is too long for the context window, old turns get dropped or summarized first. The prompt builder is where that decision gets made, before anything reaches the model.

This also explains a familiar behavior. If a long chat seems to forget something from early on, the likely reason is that the early turns were trimmed to fit, so the model never saw them.

Using this in a system design answer

If you get this question in an interview, walk the path in order. Browser, chat server, history DB, prompt builder, LLM, then the reply flowing back into the history DB. Name the two stores separately, history by conversation ID and memories per user, because mixing them up is an easy mistake.

Then raise the context window limit yourself and say what you'd do about it. The reel leaves open whether you'd summarize old turns or just drop them. Both are valid, and having a reason for your pick is what makes the answer land.

Next time a chat app seems to remember you, trace where that text came from. It was either stored history or a saved memory, pasted into the prompt on that turn.

  • #systemdesign
  • #chatgpt
  • #llm
  • #interviewprep

More reels

All news →