P
Peekr Cloud
DemoAcme Agents

← Traces

agent.run

trace_id: trace_000053 · tenant: acme

6 spans·5427msok

Context growing unboundedly: 75 → 260 tokens across 3 calls

Input token counts are climbing with each LLM call in this trace (75 → 257 → 260). This is the unbounded conversation history pattern — cost and latency grow O(N²) as the session continues. At this rate, long sessions will become expensive and slow.

Fix: trim context with a rolling window or summarisation

fix.py
# Option 1: rolling window — keep only the last N messages
MAX_MESSAGES = 10
messages = messages[-MAX_MESSAGES:]

# Option 2: summarise when context exceeds a threshold
def trim_context(messages, max_tokens=3000):
    total = count_tokens(messages)
    if total <= max_tokens:
        return messages
    # Summarise the oldest half, keep the recent half verbatim
    old = messages[:-5]
    recent = messages[-5:]
    summary = llm.summarise(old)
    return [{"role": "system", "content": f"Earlier context: {summary}"}] + recent

Waterfall

agent › run

agent.run

user: u_heavy_48
anthropic › messages › create

anthropic.messages.create

claude-opus-4-7351 tok$0.014faithful 1.00user: u_heavy_48
Input

Summarize the Q3 earnings report attached. Focus on revenue, margin, and guidance.

Output

(answer to: Summarize the Q3 earnings report attache…)

anthropic › messages › create

anthropic.messages.create

claude-opus-4-7456 tok$0.018faithful 0.89user: u_heavy_48
Input

Given the customer ticket below, draft a refund response that follows policy P-204.

Output

(answer to: Given the customer ticket below, draft a…)

anthropic › messages › create

anthropic.messages.create

claude-opus-4-7319 tok$0.012faithful 0.90user: u_heavy_48
Input

Translate the user manual section to Japanese, preserving the table structure.

Output

(answer to: Translate the user manual section to Jap…)

Tool: web_fetch

tool.web_fetch

user: u_heavy_48
Tool: web_fetch

tool.web_fetch

user: u_heavy_48