A research team at Chroma put 18 of the world's leading AI models through the same simple retrieval task twice. Once with a short input. Once with a long one. The task never changed. The accuracy did, and on several models it fell by tens of percentage points.
Nothing was broken. No model was having a bad day. This is context rot, and it explains why the AI assistant that felt sharp at 9am feels unreliable by the time you are forty messages deep into the same chat.
What is context rot?
Context rot is the fall in an AI assistant's answer quality as the conversation it has to re-read gets longer. The model is not tired and it is not faulty. It is being asked to find the one relevant line inside a pile of text that grows with every message, and its accuracy at that job drops as the pile grows.
The term comes from a July 2025 research report by Chroma, "Context Rot: How Increasing Input Tokens Impacts LLM Performance", written by Kelly Hong, Anton Troynikov and Jeff Huber. The team tested 18 leading models, including GPT, Claude, Gemini and Qwen families, and reached a blunt conclusion: models do not process a long input evenly. Performance varies significantly as input length changes, even on tasks a model handles perfectly when the input is short.
Read that again, because it is the part most business owners miss. The task was the same. Only the amount of surrounding text changed.
How does an AI chatbot actually remember your conversation?
It does not remember. Every time you press send, the assistant re-reads the whole conversation from the first line, then writes the next reply. There is no memory in the human sense, only a transcript that is read again from the top, and that transcript gets longer every single turn.
A useful way to picture it: imagine that every message you send is answered by a brand new temporary staff member. Capable, fast, and completely new. Before replying, that person must read the entire file of everything said so far.
On message three, the file is one page. That person reads it in seconds and answers precisely.
On message sixty, the file is forty pages. Same capable person, same few seconds, forty pages. Something will be skimmed.
This is also why the numbers on your subscription page matter more than they look. OpenAI's own business pricing page lists the GPT Instant total context window on a ChatGPT Business seat at 54K, with a maximum user input of roughly 40 pages, against 128K and roughly 250 pages on the Enterprise plan. The same page adds a detail worth knowing: the portion of that window available for your input is smaller than the total, because system instructions, saved memories and the model's own internal processing all occupy space in it.
You are not the only tenant in that window.
Why does a longer chat make the answers worse?
Three findings from the Chroma work explain most of what you see day to day.
Long inputs are not read evenly. The common assumption is that a model treats the ten-thousandth word with the same care as the tenth. It does not. Reliability degrades as input length grows, which is why a model can ace a task at 10,000 tokens and fumble the identical task at 100,000.
Position decides what survives. Models recall information placed at the beginning or the end of a long input far more reliably than information buried in the middle. Your carefully negotiated instruction from message twelve is now the middle of the document. That is the worst seat in the house.
Difficulty is not the trigger. Length is. Chroma's tests included deliberately easy tasks, including a synthetic exercise of repeating words back. Even there, longer inputs produced less reliable output. The lesson for a business owner is uncomfortable and useful: a simple request can fail purely because of what came before it in the same chat.
What does context rot look like inside a Hong Kong small business?
It rarely announces itself. It shows up as small, expensive drift.
The quotation that forgot the margin. A trading company owner opens a chat, states that every quote must hold a 15% margin, and works through eleven product lines. By the ninth, the AI is producing quotes that look right and price at 11%. Nobody notices until the customer accepts.
The reply that stopped sounding like your shop. A retailer sets the tone at the start: warm, brief, no exclamation marks, always offer a store pickup option. Fifty replies later the drafts are chirpy, long, and never mention pickup. The instruction did not vanish. It sank.
The glossary that quietly changed. A restaurant group agrees a bilingual menu glossary in message six: one specific dish name in Traditional Chinese, one in English, no variation. By message sixty, three variations are in circulation and the printer has used two of them.
None of these are dramatic failures. All three cost real money, and all three are invisible in the moment because the output still looks finished.
What do most people get wrong about long AI chats?
Three misconceptions do most of the damage.
Misconception 1: a bigger context window fixes it. A larger window means the model can accept more text before it is forced to drop anything. It does not mean the model reads that text with even attention. Chroma's central point is precisely that capacity and reliability are different things. More room is not more care.
Misconception 2: the AI is lying to you. When a model contradicts something agreed earlier in the same thread, it is not being deceptive. From where it sits, that instruction is a faint line in the middle of a long document it has just skimmed. It is not defying you. It genuinely did not weight it.
Misconception 3: switching to a newer model solves it. Newer models are better, and the effect still applies. Chroma found degradation across all 18 models tested, including the strongest ones available at the time. This is a property of how these systems read long inputs, not a bug in one vendor's product.
How do you stop context rot? A five-step routine
You do not need technical skill to fix this. You need a working habit, and it takes about a minute to learn.
--- One job per chat. Quotations in one conversation, customer replies in another, the staff roster in a third. Mixing three jobs in one thread triples the pile the model must re-read before every answer.
--- Front-load the brief. Put the rules that must never move into your very first message: the margin, the tone, the glossary, the deadline. The opening of a long input is one of the two positions models recall most reliably.
--- Ask for a handover note before you restart. When a thread gets long, ask: "Summarise the decisions and rules we have agreed so far, as a numbered list I can paste into a new chat." Read it. Correct it. Then open a fresh chat and paste it in. A new chat with no briefing is not a fix, it is amnesia.
--- Re-state the three rules that cannot move. Every ten or so messages, repeat them in one short line. Placing them near the end of the input puts them back in the other high-recall position.
--- Watch for the tell. When the assistant starts re-asking something you already answered, or repeating a paragraph it produced earlier, the rot has begun. Stop. Take the handover note. Start again. Do not push through.
The related discipline of deciding what to feed a model in the first place is covered in our guide to context engineering, and the question of where an answer's facts actually come from is covered in AI grounding.
Frequently asked questions about context rot
Does starting a new chat lose all my work?
Only if you start it empty. The handover note is the transfer mechanism. Ask for a numbered summary of decisions, check it yourself, paste it into the new chat, and you keep the conclusions while dropping the forty pages of discussion that produced them.
How long is too long?
There is no universal number, because it depends on the model, the plan and how much of the window is already spent on system instructions and memories. A practical rule for a busy owner: if you cannot scroll to the top of the thread in a couple of flicks, it is time for a handover note.
Is this the same as the AI hallucinating?
No, and the distinction matters. A hallucination is an invented fact. Context rot is a real instruction being under-weighted because of where it sits in a long input. The fixes are different. Hallucination is addressed by giving the model verified sources. Context rot is addressed by managing conversation length.
Do paid plans suffer from it?
Yes. A paid plan usually buys a larger window, which delays the point at which text is dropped. It does not change how evenly the model reads what is inside that window.
The takeaway for a business owner
The most common AI complaint in Hong Kong offices right now is some version of "it was good at first, then it got careless". In most cases nothing degraded except the length of the conversation.
That is unusually good news, because it means the fix costs nothing. No new subscription, no new tool, no consultant. One job per chat, rules at the top, a handover note before you restart. Three habits, learnable in a morning, and they recover most of the reliability people assume they have lost.
It also fits what Hong Kong SMEs keep saying about why AI stalls. In a Dah Sing Bank survey of more than 340 local SMEs conducted in May 2026, around 23% had already adopted AI or generative AI while 45% had not started at all, and among those who had not, 57% named a lack of relevant knowledge or skills as the main obstacle. Not price. Not the technology. Knowing how to use it properly.
That gap is a teaching problem, and teaching is what a long partnership is for. We understand AI. UD stands with you.
Reviewed by the UD AI team, Hong Kong, September 2026.
Want to put this into practice with your own team?
Knowing why your AI drifts is the first step. Turning that into a working routine your staff actually follow is the next one, and it is where most teams need a hand. UD has spent 28 years helping Hong Kong businesses make technology behave, and we will walk you through it step by step, from a plain-language assessment of where AI fits in your operation to a setup your team can run without us.