P
Peekr Cloud
DemoAcme Agents

← Traces

agent.run

trace_id: trace_00003f · tenant: acme

7 spans·3802msok

Context growing unboundedly: 147 → 425 tokens across 3 calls

Input token counts are climbing with each LLM call in this trace (147 → 496 → 425). This is the unbounded conversation history pattern — cost and latency grow O(N²) as the session continues. At this rate, long sessions will become expensive and slow.

Fix: trim context with a rolling window or summarisation

fix.py
# Option 1: rolling window — keep only the last N messages
MAX_MESSAGES = 10
messages = messages[-MAX_MESSAGES:]

# Option 2: summarise when context exceeds a threshold
def trim_context(messages, max_tokens=3000):
    total = count_tokens(messages)
    if total <= max_tokens:
        return messages
    # Summarise the oldest half, keep the recent half verbatim
    old = messages[:-5]
    recent = messages[-5:]
    summary = llm.summarise(old)
    return [{"role": "system", "content": f"Earlier context: {summary}"}] + recent

Waterfall

agent › run

agent.run

user: u_607
anthropic › messages › create

anthropic.messages.create

claude-opus-4-7674 tok$0.026faithful 0.99user: u_607
Input

Summarize the Q3 earnings report attached. Focus on revenue, margin, and guidance.

Output

(answer to: Summarize the Q3 earnings report attache…)

anthropic › messages › create

anthropic.messages.create

claude-opus-4-7813 tok$0.032faithful 1.00user: u_607
Input

Given the customer ticket below, draft a refund response that follows policy P-204.

Output

(answer to: Given the customer ticket below, draft a…)

anthropic › messages › create

anthropic.messages.create

claude-opus-4-7835 tok$0.033faithful 0.97user: u_607
Input

Translate the user manual section to Japanese, preserving the table structure.

Output

(answer to: Translate the user manual section to Jap…)

Tool: search

tool.search

user: u_607
Tool: calculator

tool.calculator

user: u_607
Tool: code_exec

tool.code_exec

user: u_607