diff --git a/plans/meta-prompts/loop_engineering.md b/plans/meta-prompts/loop_engineering.md index 2e475bf..83bb193 100644 --- a/plans/meta-prompts/loop_engineering.md +++ b/plans/meta-prompts/loop_engineering.md @@ -230,3 +230,29 @@ Reference: Jeremy Howard talk at Answer.AI (YouTube `SUZwYV5JYBM`), built on Dec - Prefer architectures where the primary agent does the reasoning and delegates execution, not where it becomes a blind router. - Size your loops so the human stays in the learning loop — the agent handles toil, the human handles judgment. - If a day of agentic work leaves you energized and knowledgeable, you're augmenting. If it leaves you with output you can't explain or debug, you're in dark flow. + +## Model routing — tiered model strategy + +Reference: model-routing cost-savings tutorial (YouTube `SUZwYV5JYBM`), Coinbase engineering blog (GLM 5.2 for most coding tasks). + +**Core thesis:** use a frontier model for **planning/research/review** and a cheaper model for **execution/coding**. The plan-execute split saves 68%+ on token costs because coding produces ~6x more output tokens than planning, and output tokens are 5x more expensive than input on frontier models. + +**TAC model tier mapping (current provider config):** + +| Tier | Models | Role | Best for | +|------|--------|------|----------| +| **1 — Frontier planning** | ZAI GLM-5.2, Nous/StepFun 3.7 Flash | Architecture, design, review, complex debugging | Planning phase, PR review, spec writing | +| **2 — Execution coding** | Go/DeepSeek V4 Flash, opencode-go/qwen3.6-plus | Writing code from specs, implementing defined features, running tests | Execution phase after spec is written | +| **3 — Fast/cheap** | Haiku 4.5, Nous/StepFun (free tier) | Linting, simple git ops, status checks, deploy steps | Mechanical tasks that don't need reasoning | + +**The math:** Planning with Tier 1 (100K input + 20K output tokens) = ~$2/feature. Coding with Tier 1 (150K input + 120K output tokens) = ~$7.50/feature. Total: ~$9.50/feature. Coding with Tier 2 instead: $0.30 + $0.72 = ~$1.02/feature. Total with split: ~$3.02/feature. **68% savings.** + +**Coinbase pattern:** rising total token usage with flat/falling costs by routing most tasks to cheap open-source models (GLM 5.2) and reserving frontier models for planning. Also uses aggressive context caching and thinking-effort tiering. + +**Thinking effort matters:** low thinking for simple tasks (deploy this code, check status, format this file), high/max thinking only for hard architectural problems. The default "high" on most platforms is overkill for >50% of agentic work. + +**Implementation patterns for TAC:** +- pi's `-m` flag already implements manual model routing — use `-m opencode-go/deepseek-v4-flash` for execution, `-m zai/glm-5.2` for planning +- The brainstorms/spec skills naturally produce the architect's spec → handoff to builder pattern +- Third-party harnesses (Cursor auto mode, Not Diamond) handle this automatically if configured +- The Agent multiplexer note's orchestrator can route planning to one pane (Tier 1) and execution to another (Tier 2)