106 lines
3.5 KiB
Markdown
106 lines
3.5 KiB
Markdown
# Cascade Routing: Cut Your API Costs by 66%
|
|
|
|
**June 11, 2026**
|
|
|
|
Most teams use one model for everything. They default to Claude Opus or GPT-5 for every task, which means they are paying premium prices for simple work.
|
|
|
|
## The Price Range
|
|
|
|
| Model | Input ($/M) | Output ($/M) |
|
|
|-------|-------------|----------|
|
|
| Gemini 2.5 Flash | $0.15 | $0.60 |
|
|
| DeepSeek V3 | $0.27 | $1.10 |
|
|
| Claude Sonnet 4 | $3.00 | $15.00 |
|
|
| Claude Opus 4 | $15.00 | $75.00 |
|
|
|
|
That is a 100x range between the cheapest and most expensive.
|
|
|
|
## The Cascade Pattern
|
|
|
|
Route different steps to different models. Use cheap models for simple retrieval and formatting. Use expensive models only for complex reasoning.
|
|
|
|
```
|
|
Retrieve context -> Gemini Flash ($0.15/$0.60)
|
|
Analyze data -> Claude Sonnet ($3/$15)
|
|
Make decision -> Claude Opus ($15/$75)
|
|
Format output -> Gemini Flash ($0.15/$0.60)
|
|
```
|
|
|
|
## The Savings
|
|
|
|
| Pattern | Cost/Task | Savings |
|
|
|---------|-----------|--------|
|
|
| All Opus | $2.50 | Baseline |
|
|
| Cascade | $0.85 | 66% |
|
|
| All Sonnet | $0.50 | 80% (but quality loss on complex steps) |
|
|
|
|
## Implementation
|
|
|
|
```python
|
|
def route_task(task_complexity: str) -> str:
|
|
if task_complexity == "retrieval":
|
|
return "gemini-2.5-flash"
|
|
elif task_complexity == "analysis":
|
|
return "claude-sonnet-4"
|
|
elif task_complexity == "decision":
|
|
return "claude-opus-4"
|
|
elif task_complexity == "formatting":
|
|
return "gemini-2.5-flash"
|
|
```
|
|
|
|
## Dynamic Cascade — Route Based on Content
|
|
|
|
A smarter approach: measure the complexity of each input and route dynamically:
|
|
|
|
```python
|
|
def dynamic_cascade(prompt: str) -> str:
|
|
"""Measure prompt complexity and route to appropriate tier."""
|
|
token_count = len(prompt.split())
|
|
has_code = "```" in prompt or "def " in prompt
|
|
has_reasoning = any(word in prompt.lower() for word in
|
|
["explain", "analyze", "compare", "why", "how"])
|
|
|
|
if token_count > 2000 or (has_code and has_reasoning):
|
|
return "claude-opus-4" # complex: needs reasoning
|
|
elif token_count > 500 or has_reasoning:
|
|
return "claude-sonnet-4" # moderate: needs some analysis
|
|
else:
|
|
return "gemini-2.5-flash" # simple: cheap model is fine
|
|
```
|
|
|
|
This catches the case where a supposedly "simple" task turns out to need reasoning. The dynamic approach typically saves 50-70% while keeping quality high.
|
|
|
|
## Real-World: Multi-Agent Cascade
|
|
|
|
In a multi-agent system, cascade routing applies at the agent level too:
|
|
|
|
```yaml
|
|
# cascade-config.yaml
|
|
agents:
|
|
research-agent:
|
|
model: gemini-2.5-flash # cheap — bulk web scraping
|
|
max_tokens: 4000
|
|
|
|
analyzer-agent:
|
|
model: claude-sonnet-4 # mid — pattern recognition
|
|
max_tokens: 8000
|
|
|
|
synthesis-agent:
|
|
model: claude-opus-4 # premium — report generation
|
|
max_tokens: 16000
|
|
```
|
|
|
|
Each agent gets the model tier appropriate for its function. The fleet costs 70% less than running all agents on Opus.
|
|
|
|
## When Not to Cascade
|
|
|
|
If your task is a single critical decision, use the best model. Cascade routing shines when you have a pipeline of steps with varying complexity, which is most real-world agent systems. Also avoid cascading for:
|
|
|
|
- **Single-turn tasks** — On/off decisions that need the best reasoning
|
|
- **Creative work** — Writing, design, strategy — quality > cost
|
|
- **Safety-critical actions** — Database operations, deployments — use the most reliable model
|
|
|
|
---
|
|
|
|
*From Module 6 of the [Agentic Engineering Course](/). The full module includes a cost optimization lab with working code.*
|