YanMate
How much of your context window is left for the conversation?
A character forgets because its context window fills up. The model can read only so many tokens at once, and the character card takes its share of them on every single message — the persona and scenario are re-sent each turn, and the frontend holds more back for the reply. What remains is the room for your conversation, and once the chat outgrows it the oldest messages are cut to make space. Set your context window and your card below, drag the conversation forward, and see how many exchanges the character still remembers and exactly when it starts to forget. It runs in your browser, and a card you drop is never uploaded.
Simulate a conversation
The card
Optional. A card fills in the first three numbers; without one, they are a typical mid-sized card.
50 exchanges in, the character remembers only the last 34 exchanges. The first 16 exchanges are gone.
Forgetting started at exchange 35. From there on, every new exchange pushes the oldest one out. The example dialogue was dropped at exchange 34 to make room for the conversation.
The card's permanent fields take the room of about 5 exchanges — memory it gives up on every single message.
- Held for the reply300
- Card (every message)900
- Conversation6,800
- Free192
The same chat in other windows
| Window | Remembers up to | Card's share |
|---|---|---|
| 4k | 14 exchanges | 22.0% |
| 8k (yours) | 34 exchanges | 11.0% |
| 16k | 75 exchanges | 5.5% |
| 32k | 157 exchanges | 2.7% |
| 128k | 649 exchanges | 0.7% |
To see where a card's tokens come from field by field, use the token counter.
What fills the window
Five things share the window on every message, and only the last of them is your conversation. The first two never shrink, however long you talk.
Held back for the reply
Before anything else, the frontend reserves room for the answer the model is about to write. It is usually a few hundred tokens and it is set in your frontend's settings as the response length. Room held for a reply cannot hold history.
The card, every message
The persona, the scenario and any system prompt are sent on every turn for the whole life of the chat. This is the cost that decides how much a character can remember, and it is the one the token counter's permanent subtotal measures.
Lorebook entries, when they match
A lorebook entry is injected only on turns where one of its keys appears. When you drop a card here, the simulator assumes one average entry is active per turn — raise it if your chats tend to trigger several at once.
Example dialogue, until it is needed
Most frontends send the examples while there is room and drop them first when the conversation needs the space. They buy the character's voice early on, and they are the first thing given up for memory.
The conversation
Everything left is history, kept newest first. The greeting is part of it — it is the first message, so it is the first thing to fall out. When the next exchange does not fit, the oldest one goes.
How to make a character remember more
Trim the permanent fields first. A persona is paid for on every message, so cutting two hundred tokens from it returns two hundred tokens of history on every turn for as long as the character exists. Cutting the same amount from the greeting saves it once. The token counter shows which fields are permanent and how big each one is.
Move facts out of the persona. A detail the character needs only when a subject comes up — a hometown, a sibling's name, how a spell works — belongs in a lorebook entry that is sent when its key appears, not in the persona that is sent always. The lorebook vs memory guide goes through which is which.
Use a bigger window, if you can. Doubling the window more than doubles the history, because the card does not grow with it. Whether you can depends on the model and on what your frontend or provider charges for it.
Shorter replies last longer. An exchange's size is mostly the reply. A character that answers in one paragraph instead of three keeps more than twice as many exchanges in the same window.
What this simulator leaves out
It is a model of the part every frontend shares, and it is honest to say where that stops. Frontends add system prompts of their own, and some summarise old messages instead of cutting them, or keep a separate long-term memory that survives outside the window entirely — none of that is simulated here, so the forgetting point it shows is the one you would hit without those features.
Every token count is an estimate at about four characters per token, which drifts 10 to 20 percent on English and much further on Japanese, Chinese, Korean or emoji. Real conversations are also uneven: one long reply takes the room of several short ones. Treat the result as the shape of the trade-off — where the card's cost lands and roughly when forgetting begins — rather than a prediction of which message a particular app will drop.
For how memory works when it is kept outside the window, the guide to how AI companion memory works covers summaries, memory stores and what each of them can and cannot recall.
Questions
- Why does my AI character forget what I said earlier?
- Because the model can only read a fixed number of tokens at once — its context window — and every message you send has to fit in it alongside the character card. Once the conversation outgrows the room the card leaves, the frontend starts cutting the oldest messages to make space for the newest. Nothing is broken and nothing is deleted from your chat log; the model simply is not shown those messages any more, so from its side they never happened.
- How many messages does a character remember?
- It depends on three numbers: the size of the context window, how much of it the card takes on every message, and how long the messages are. In an 8k window, a 900-token card and replies of a paragraph each leave room for roughly thirty-five exchanges. The same card in a 32k window keeps more than 150. The simulator above does this arithmetic for your own numbers, and the table under it shows the same chat across the common window sizes.
- Does a longer character description make it forget faster?
- Yes, directly. The persona and scenario are re-sent on every message, so every token they use is a token that can never hold anything you said. A 1,500-token persona in an 8k window costs the room of seven or eight exchanges on every single turn. That is why trimming a persona is worth more than trimming a greeting: the greeting is sent once and scrolls out like any other message, while the persona is paid for forever.
- What happens to the example dialogue?
- Most frontends include it while there is room and drop it first once the conversation needs the space — before any of your messages are cut. This simulator models it that way: examples stay in the window until they no longer fit alongside the history, then they go all at once. Some frontends let you choose to keep them permanently instead, in which case count them as part of the card.
- Is this exactly what SillyTavern or Character.AI does?
- No, and it does not pretend to be. Every frontend assembles its prompt a little differently — some add their own system prompt, some summarise old messages instead of dropping them, some keep a separate long-term memory. This is a model of the part they share: a fixed window, a card that is always in it, and history that is cut oldest-first. The token counts are also estimates at about four characters per token. Use it to see the shape of the trade-off, not to predict the exact message a particular app will forget.
- Is my card uploaded anywhere?
- No. If you drop a card, it is read by JavaScript in your browser using the same parser this site uses for imports, and the only thing kept from it is three numbers, which go into the page address so you can share the result. The file itself never leaves your device, and without a card the simulator does not even load the parser.
Other free tools
- Character card viewerDrop a character card — PNG, JSON or .charx — and read every field, warning and hidden value inside it, then convert between V2, V3 and PNG. Runs in the browser; nothing is uploaded.
- Character.AI export viewerOpen the ZIP Character.AI emails you and read it: your characters as cards you can take anywhere, your chat history as plain text. Runs in the browser; the archive is never uploaded.
- Character card token counterDrop a card or paste a persona and see the token cost field by field — separating what is re-sent on every single message from what is sent once or only sometimes. Runs in the browser.
- Character voice analyzerPaste a character's replies or a chat transcript and see how it talks — sentence length, habits, repeated phrases — with the best replies laid out as example dialogue for a character card. Runs in the browser; nothing is uploaded.
- Character card makerFill in a form and download a real character card PNG carrying both the V2 and V3 chunks, so it opens anywhere. Any image works — a JPEG is converted in the browser. Nothing is uploaded.
- Character card templatesEight original V3 starters — mentor, rival, cozy companion, game master, language partner, study coach, noir detective, historical figure. Download PNG or JSON, or open one in the maker. Nothing is uploaded.
- AI companion cost calculatorPick how long you would keep an AI companion app and see what each plan costs over exactly that stretch, monthly against prepaid annual — every price sourced and dated, blanks where none is published. Runs in the browser.
- Character safety checkPaste a character's persona, its greeting or a reply it gave you and check it against a published list of 30 guilt, dependency, exclusivity and minor-coded language patterns. Every finding names the rule it broke. Runs in the browser.
An AI character on YanMate.