P
Peekr Cloud
DemoAcme Agents

← Traces

agent.plan

trace_id: trace_00004d · tenant: globex

4 spans·4481msok

Context growing unboundedly: 1004 → 2682 tokens across 3 calls

Input token counts are climbing with each LLM call in this trace (1004 → 876 → 2682). This is the unbounded conversation history pattern — cost and latency grow O(N²) as the session continues. At this rate, long sessions will become expensive and slow.

Fix: trim context with a rolling window or summarisation

fix.py
# Option 1: rolling window — keep only the last N messages
MAX_MESSAGES = 10
messages = messages[-MAX_MESSAGES:]

# Option 2: summarise when context exceeds a threshold
def trim_context(messages, max_tokens=3000):
    total = count_tokens(messages)
    if total <= max_tokens:
        return messages
    # Summarise the oldest half, keep the recent half verbatim
    old = messages[:-5]
    recent = messages[-5:]
    summary = llm.summarise(old)
    return [{"role": "system", "content": f"Earlier context: {summary}"}] + recent

Waterfall

agent › plan

agent.plan

user: u_832
anthropic › messages › create

anthropic.messages.create

claude-opus-4-72.6k tok$0.100faithful 0.39user: u_832
Input

Summarize the Q3 earnings report attached. Focus on revenue, margin, and guidance.

Output

Q3 revenue hit $4.2B, a 38% YoY jump, with operating margin expanding to 27%…

anthropic › messages › create

anthropic.messages.create

claude-opus-4-71.5k tok$0.059faithful 0.97user: u_832
Input

Given the customer ticket below, draft a refund response that follows policy P-204.

Output

(answer to: Given the customer ticket below, draft a…)

anthropic › messages › create

anthropic.messages.create

claude-opus-4-73.9k tok$0.151faithful 0.90user: u_832
Input

Translate the user manual section to Japanese, preserving the table structure.

Output

(answer to: Translate the user manual section to Jap…)