Context growing unboundedly: 147 → 425 tokens across 3 calls
Input token counts are climbing with each LLM call in this trace (147 → 496 → 425). This is the unbounded conversation history pattern — cost and latency grow O(N²) as the session continues. At this rate, long sessions will become expensive and slow.
Fix: trim context with a rolling window or summarisation
# Option 1: rolling window — keep only the last N messages
MAX_MESSAGES = 10
messages = messages[-MAX_MESSAGES:]
# Option 2: summarise when context exceeds a threshold
def trim_context(messages, max_tokens=3000):
total = count_tokens(messages)
if total <= max_tokens:
return messages
# Summarise the oldest half, keep the recent half verbatim
old = messages[:-5]
recent = messages[-5:]
summary = llm.summarise(old)
return [{"role": "system", "content": f"Earlier context: {summary}"}] + recentWaterfall
agent.run
anthropic.messages.create
Summarize the Q3 earnings report attached. Focus on revenue, margin, and guidance.
(answer to: Summarize the Q3 earnings report attache…)
anthropic.messages.create
Given the customer ticket below, draft a refund response that follows policy P-204.
(answer to: Given the customer ticket below, draft a…)
anthropic.messages.create
Translate the user manual section to Japanese, preserving the table structure.
(answer to: Translate the user manual section to Jap…)
tool.search
tool.calculator
tool.code_exec