Delegation Depth Explosion
Issue: Agents Delegate to Sub-Agents Creating Unbounded Depth
Frequency: Occasional
Symptoms
- Task passes through many agent layers
- Context diluted at each delegation level
- Latency compounds with each delegation
- Token costs multiply per delegation
- Original intent lost in delegation chain
Root Cause Agents that can spawn or delegate to other agents may create deep delegation chains. Each level adds latency, token overhead, and potential for context loss. A simple task that could be handled directly gets delegated through 5+ layers, each agent adding its overhead. Without depth limits, delegation can spiral into unbounded recursion.
Example
Scenario: Research agent with delegation capability
User query: "What's the weather in Tokyo?"
Agent delegation chain:
L0: Main Agent
"I'll delegate this to my research agent"
L1: Research Agent
"I'll delegate this to my data gathering agent"
L2: Data Gathering Agent
"I'll delegate this to my API specialist agent"
L3: API Specialist Agent
"I'll delegate this to my weather API agent"
L4: Weather API Agent
Finally calls weather API
Returns: "72°F, Sunny"
L4→L3: "The weather API reports 72°F, Sunny in Tokyo"
L3→L2: "My API specialist found it's 72°F and Sunny"
L2→L1: "Data gathering confirms 72°F, Sunny weather"
L1→L0: "Research indicates Tokyo is 72°F and Sunny"
L0→User: "Based on my research, Tokyo is 72°F and Sunny"
Cost analysis:
Direct call: 1 API call + 1 LLM call
With delegation: 1 API call + 10 LLM calls (2 per level)
Latency: 5x (each level adds ~500ms)
Tokens: 8x (context passed up and down)
Cost: $0.002 vs $0.016
Worse case: Infinite delegation
Agent A delegates to Agent B
Agent B delegates to Agent A
System hangs or crashes
Key Statistics From Delegation Research (2026):
- Average delegation depth in production: 2-3 levels
- Each delegation level adds 400-800ms latency
- Token overhead per level: 150-300 tokens
- Delegation loops cause 8% of agent hangs
- Context fidelity drops 10-20% per level
Delegation Problems
| Problem | Cause | Impact |
|---|---|---|
| Deep chains | No depth limit | High latency/cost |
| Circular delegation | A→B→A | Infinite loop |
| Context loss | Summarization at each level | Wrong output |
| Over-delegation | Simple tasks delegated | Waste |
| Accountability loss | “Who did what?” unclear | Debug difficulty |
Contributing Factors
- No delegation depth limits
- Agents optimized to delegate
- No task complexity assessment
- Missing delegation tracking
- No direct execution preference
- Recursive agent architectures
Test Scenario & Reproduction
Scenario Setup
- A hierarchical multi-agent system where each agent defaults to delegating rather than checking whether it can execute directly
- No hard depth-budget limit or complexity-based delegation gate configured
- A trivially simple task (single API call) submitted at the top of the hierarchy
Trigger Mechanism
- Submit a simple, single-tool-call task to the top-level agent (L0) in a system with delegation capability at every layer
- Let each layer choose to delegate further rather than execute directly, and record the resulting chain depth
- Measure the cumulative latency, token cost, and final-answer fidelity against a direct single-call baseline
Example Reproduction Steps:
1. Submit "What's the weather in Tokyo?" to the Main Agent (L0) with delegation enabled
2. Trace the chain as L0 delegates to a Research Agent (L1), which delegates to a Data Gathering Agent (L2), which delegates to an API Specialist Agent (L3), which delegates to a Weather API Agent (L4)
3. Record that only L4 actually calls the weather API, returning "72F, Sunny"
4. Trace the response as it is re-narrated at each level going back up (L4->L3->L2->L1->L0->User)
5. Measure total LLM calls (expect ~10 vs. 1 for direct execution), cumulative latency (~5x), and token cost (~8x, $0.016 vs $0.002)
6. Additionally configure Agent A to delegate to Agent B and Agent B to delegate back to Agent A, and observe whether the system hangs
Expected Failure State
- The simple weather lookup passes through 5 agent layers (L0-L4) instead of being resolved with a single direct tool call
- Total LLM calls, latency, and token cost are several times higher than the direct-execution baseline (approximately 10x calls, 5x latency, 8x cost per the documented figures) for identical output
- The final answer to the user is a re-narrated paraphrase (“Based on my research, Tokyo is 72F and Sunny”) several hops removed from the actual API response, with measurable fidelity drift at each hop
- In the circular-delegation variant, the system hangs or crashes with no depth limit catching the A->B->A loop
Mitigation Strategies
Prevention
- Hard depth budget passed through the call chain: The Tokyo weather example shows a simple lookup passed through 5 layers (L0-L4) purely because each agent’s default behavior is “delegate” rather than “check if I can just do this.” Attach a decrementing depth budget (e.g., start at 3) to the task at L0, and require any agent receiving budget=0 to execute directly or fail rather than delegate further — this caps the chain before it reaches L4-level absurdity. Trade-off: a hard cap can force a genuinely complex task to be handled by an under-equipped agent if the budget is set too low.
- Complexity-gated delegation instead of default-delegate: “What’s the weather in Tokyo?” is a single API call, yet every layer chose to delegate rather than execute — the stats confirm average production depth should be 2-3 levels, not 5. Require each agent to score task complexity (e.g., “does this need one tool call or multiple reasoning steps?”) before delegating, and force direct execution for single-tool-call tasks like weather lookups. Trade-off: complexity scoring itself costs an LLM call and can be wrong, occasionally delegating something that should have been direct or vice versa.
- Circular-delegation guard via visited-agent set: The “worse case” in the example — Agent A delegates to B, B delegates back to A — causes a hang/crash with no depth limit needed to trigger it. Attach a visited-agent-ID list to the task context and reject any delegation that would re-enter an agent already in the chain. Trade-off: requires every agent in the system to honor and propagate the visited-list faithfully; a non-compliant agent reintroduces the loop risk.
Detection & Response
- Per-level latency accumulation tracking: Since each delegation level in the example adds ~500ms (400-800ms per the stats) and the chain compounds to 5x direct latency, instrument each hop to log cumulative latency and flag any task exceeding 2x the direct-execution baseline latency for its task type.
- Token/cost multiplier monitor: The example shows 8x token cost ($0.016 vs $0.002) purely from delegation overhead with zero added value (same answer, just re-narrated at each level). Track the ratio of total tokens consumed across the chain vs. tokens the final tool call itself required, and flag ratios above a threshold (e.g., >3x) as excessive delegation overhead.
- Context fidelity decay check: The stats note 10-20% fidelity drop per level as “72°F, Sunny” gets re-narrated (L4→L3→L2→L1→L0). Sample final output against the ground-truth tool result and flag semantic drift (paraphrase divergence) beyond a threshold as a sign the chain is too deep for the content it’s carrying.
Architecture Patterns
- Supervisor with bounded delegation depth: A single supervisor (L0) holds the depth budget and a registry of available direct-execution tools (like the weather API), so it can call the tool itself instead of delegating to a Research Agent that delegates to a Data Gathering Agent, etc. Deployment consideration: requires the supervisor to maintain an up-to-date tool/capability registry, or it will fall back to delegation anyway.
- Flat tool-calling instead of nested agent-to-agent delegation: For tasks like the weather lookup that resolve to one API call, replace the L1-L4 agent hierarchy with direct tool access from L0 — i.e., treat “weather API” as a callable tool, not a sub-agent to delegate to. Deployment consideration: blurs the architectural line between “agent” and “tool,” requiring clear guidelines on when a capability should be a tool vs. a full agent.
- Delegation ledger with loop detection: Maintain an explicit, task-scoped ledger recording every delegation hop (who delegated to whom, at what depth) that is checked before each new delegation for depth-budget and cycle violations — directly preventing the “Agent A delegates to Agent B, Agent B delegates to Agent A” infinite loop. Deployment consideration: the ledger must be passed reliably through every hop; losing it mid-chain reopens the loop risk.
Metrics
- avg_delegation_depth: Target 2-3 levels (per observed production baseline); Alert if p95 exceeds 4 levels for any task category.
- delegation_overhead_ratio: Target < 2x tokens/latency vs. direct execution baseline; Alert if > 5x (matching the example’s 8x/5x blowup pattern).
- circular_delegation_rate: Target 0% of tasks hitting a repeated agent ID in the delegation chain; Alert on any occurrence (P1, since this causes hangs per the 8% hang-rate stat).
- context_fidelity_at_final_hop: Target > 90% semantic similarity between final output and ground-truth tool result; Alert if < 80%.
Alerts
- Circular Delegation Detected (P1): Condition - a delegation chain re-enters an agent ID already present in the visited-agent list. Action: immediately abort the chain, return an error to the originating caller, and log the full chain for debugging (matches the “system hangs or crashes” failure mode in the example).
- Depth Budget Exhausted (P2): Condition - a task’s delegation depth budget reaches 0 without task completion. Action: force the current agent to execute directly with available tools or return a “cannot complete without further delegation” response rather than silently continuing.
- Excessive Overhead Ratio (P3): Condition - delegation_overhead_ratio for a completed task exceeds 5x the direct-execution baseline. Action: log for delegation-policy review; if recurring for the same task pattern, add a complexity-threshold rule to route that pattern to direct execution.
References
- MAST Taxonomy - Multi-agent coordination failures
- Augment Code: Multi-Agent Failures - Delegation patterns
- DEV.to: $47,000 Agent Loop - Runaway agents
- Microsoft: Failure Modes in Agentic AI - Agent coordination
- LeanOps: Token Cost Analysis - Cost patterns