9.4 KiB
Module 6: Economics & Evaluation
Lesson 6.0: The Compute Advantage Equation
Core insight: Your value as an agentic engineer scales with the amount of compute you can harness effectively.
The Equation
Compute Advantage = (Compute Scaling × Autonomy) ÷ (Time + Effort + Monetary Cost)
| Variable | What It Means | How to Improve It |
|---|---|---|
| Compute Scaling | How much AI compute you can throw at problems | More agents, better models, larger context windows |
| Autonomy | How much the agent does without your intervention | Better prompts, better tools, verification layers |
| Time | How long it takes to get results | Faster models, parallel execution, fewer iterations |
| Effort | How much you have to craft prompts/instructions | Reusable skills, templates, mental models |
| Monetary Cost | What you pay for API calls | Cascade routing, cheaper models for simple steps |
What It Tells You
- High Compute Advantage = You get more output for less input. You're leveraging agents effectively.
- Low Compute Advantage = You're spending too much time/cost for too little gain. Fix the bottleneck variable.
Real-World Application (Your Stack)
| Tool | Compute Scaling | Autonomy | Time | Effort | Cost | Advantage |
|---|---|---|---|---|---|---|
| Claude Code (lead) | 8 | 7 | 6 | 5 | 4 | (8×7)÷(6+5+4)=3.7 |
| Pi Agent (custom) | 7 | 8 | 7 | 6 | 6 | (7×8)÷(7+6+6)=2.9 |
| OpenCode (OSS) | 5 | 5 | 7 | 7 | 9 | (5×5)÷(7+7+9)=1.1 |
| Gemini Flash (fast) | 4 | 3 | 9 | 8 | 9 | (4×3)÷(9+8+9)=0.5 |
Higher score = more output per unit of investment. Use this to decide which tool for which task.
How to Optimize
- Improve numerator: Run more agents in parallel (P-threads), increase autonomy with verification
- Reduce denominator: Use cascade routing (cheap model for simple steps), reuse skills/mental models
- Track over time: Your Compute Advantage should increase as you build better harnesses and mental models
Lesson 6.1: LLM Pricing Landscape 2026
Per-Million-Token Pricing
| Model | Input ($/M) | Output ($/M) | Best For |
|---|---|---|---|
| Gemini 2.5 Flash | $0.15 | $0.60 | High-volume, simple tool calls |
| DeepSeek V3 | $0.27 | $1.10 | Structured tasks, batch processing |
| DeepSeek R1 | $0.55 | $2.19 | Reasoning-heavy single steps |
| Gemini 2.5 Pro | $1.25 | $5.00 | Long-context sessions (1M+ tokens) |
| GPT-4o | $2.50 | $10.00 | Balanced cost/quality |
| Claude Sonnet 4 | $3.00 | $15.00 | General agentic reasoning |
| GPT-5 | $10.00 | $40.00 | Frontier research, complex plans |
| Claude Opus 4 | $15.00 | $75.00 | Complex multi-step agent orchestration |
Price Range: 100x Difference
Cheapest (Gemini Flash) to most expensive (Claude Opus) is a 100x multiplier. Choosing the right model for each step is your highest-leverage cost optimization.
Lesson 6.2: Cascade Routing
The Pattern
Don't use one model for everything. Route different steps to different models:
Step 1: Retrieve context → Gemini Flash ($0.15/M input)
Step 2: Analyze → Claude Sonnet ($3/$15)
Step 3: Make decision → Claude Opus ($15/$75)
Step 4: Format output → Gemini Flash ($0.15/$0.60)
Savings Profile
| Pattern | Cost/Task | Savings |
|---|---|---|
| All Opus | $2.50 | Baseline |
| Cascade (Flash → Sonnet → Opus → Flash) | $0.85 | 66% savings |
| All Sonnet | $0.50 | 80% savings (but quality loss on complex steps) |
Implementation
def route_task(task_complexity: str) -> str:
if task_complexity == "retrieval":
return "gemini-2.5-flash"
elif task_complexity == "analysis":
return "claude-sonnet-4"
elif task_complexity == "decision":
return "claude-opus-4"
elif task_complexity == "formatting":
return "gemini-2.5-flash"
Lesson 6.3: Cost Per Session Math
Where Costs Come From
| Component | Share | Notes |
|---|---|---|
| Output tokens | ~70% | Model generation is most expensive |
| Input tokens | ~20% | Context + tool results |
| Cached tokens | ~10% | Can be zero if not configured |
The Multipliers
| Factor | Multiplier | Why |
|---|---|---|
| Retry rate | 1.2-2.0x | Failed tool calls retry |
| Tool overhead | 3-5x per tool call | Planning + execution + error recovery + result parsing |
| Context growth | 1.5x per 10 turns | Every turn adds tokens to context |
The 3x Rule
Production agent costs 3x your prototype estimate.
1x = ideal path (everything works first time) 2x = retries + edge cases 3x = monitoring + error handling + observability overhead
Quick Estimation
def estimate_cost(turns: int, avg_tokens_per_turn: int, model_price_per_m: float):
"""Very rough estimate."""
base = turns * avg_tokens_per_turn * model_price_per_m / 1_000_000
retry = base * 1.5
overhead = base * 3.0
return {"base": base, "with_retries": retry, "production": overhead}
Lesson 6.4: Agent Evaluation Metrics
The Key Metrics
| Metric | What It Measures | Target |
|---|---|---|
| pass@k | % of k attempts where at least one succeeds | >80% |
| pass^k | % where ALL k attempts succeed (consistency) | >60% |
| Tool Call Accuracy | Correct tool + correct params | >90% |
| Task Completion Rate | End-to-end success | >70% |
| Cost Per Task | Total API cost per unit | Varies |
| Loop Efficiency | Steps taken vs optimal | <2x overhead |
| Grind Rate | % of attempts that re-run identical code | <5% |
pass@k Explained
pass@k = probability that at least one of k attempts succeeds
k=1: 60% pass rate (single attempt)
k=3: 1 - (0.4)^3 = 93.6% (best of 3)
k=5: 1 - (0.4)^5 = 98.9% (best of 5)
Higher k = higher reliability but higher cost. The trade-off is the core optimization problem.
Lesson 6.5: Automated Evaluation
Golden Dataset Approach
eval_cases = [
{
"input": "Find the user with email john@example.com",
"expected_tool": "query_database",
"expected_params": {"query": "SELECT * FROM users WHERE email = 'john@example.com'"},
"expected_output_contains": ["john@example.com"],
"weight": 1.0
}
]
VCR-Style Recording
For non-deterministic tool results (search, API calls), record the response once, then replay it deterministically:
# Record mode: capture real responses
# Replay mode: use recorded responses
# Test: verify agent makes correct decisions with known data
Lesson 6.6: A/B Testing Agents
Canary Deployment
- Route 5% of traffic to new agent config
- Compare against 95% on current config
- Metrics: success rate, cost, latency, loop depth
- If new config wins on all metrics → roll out to 100%
- If new config loses → rollback, investigate
What to A/B Test
- System prompt wording
- Model selection
- Temperature settings
- Tool descriptions
- Iteration limits
Lesson 6.7: Human Evaluation
What Automated Evals Miss
- Quality of reasoning — agent made right decision for wrong reasons?
- Tone and style — output technically correct but poorly written?
- Edge cases — agent handled happy path but not real-world variation?
- Hallucination cascades — plausible-looking but wrong intermediate steps?
Spot-Check Sampling
Review 5-10% of all agent sessions manually. Focus on:
- Sessions with high cost (>2x average)
- Sessions with tool loops
- Sessions where automated eval scored low
- Random sample for baseline
Lesson 6.7b: Real-World Cost Scenarios
Scenario A: Research Agent (High Token Burn)
Pattern: Agent reads 20 web pages, synthesizes report
Cost breakdown:
├── 20 web page reads @ 2K tokens each = 40K input tokens
├── 5 LLM reasoning turns @ 4K tokens = 20K tokens
├── 1 report generation @ 8K output tokens = 8K output tokens
├── Claude Sonnet: ~$0.08/run
├── Claude Haiku: ~$0.02/run (80% cheaper, 90% as good for extraction)
└── Cascade: Haiku for reads, Sonnet for synthesis → $0.03/run
The cascade saves 62% on this exact pattern. Use cheap models for bulk work, expensive models for synthesis.
Scenario B: Code Generation (Iterative)
Pattern: Agent writes code, tests, fixes, repeats
Cost without optimization:
├── Average: 8 iterations × 6K tokens = 48K tokens
├── Cost: ~$0.35 per feature
├── With guardrails (limit to 3 iterations): $0.13 per feature
└── Savings: 63%
Scenario C: Always-On Brand Monitor
Pattern: Scans 50 sources every 5 minutes, 24/7
Cost without optimization:
├── 288 runs/day × $0.08 = $23.04/day = $691/month
├── With cascade + dedup + scheduling: $4.15/day = $125/month
└── Savings: 82%
The 80/20 Rule
90% of cost savings come from three changes:
- Model cascade — use cheap models for routine work (saves 50-80%)
- Iteration limits — cap loops at 3-5 turns (saves 40-60%)
- Deduplication — don't re-read the same context (saves 20-30%)
Do these three first before any other optimization.
Lab 6.8: Build an Eval Harness
Objective: Create golden Q&A pairs + automated pass/fail scoring.
Starter: course/labs/L6-eval-harness/starter.py
Lab 6.9: Cost Optimization
Objective: Profile a session, identify savings, implement cascade routing.
Starter: course/labs/L6-cost-optimization/starter.py