# Module 4: Multi-Agent Orchestration ## Lesson 4.1: Why One Agent Is Not Enough A single agent, no matter how smart, hits three ceilings: 1. **Context ceiling** — One agent doing everything means one context window holding everything. File contents, database schemas, business logic, deployment configs — it all competes for space. 2. **Capability ceiling** — A generalist agent is mediocre at everything. A specialized agent (backend dev, security reviewer, UI designer) outperforms the generalist at its domain. 3. **Reliability ceiling** — One agent failing = everything fails. Multi-agent systems degrade gracefully — one team's failure doesn't take down the whole system. ### The Stacking Insight ``` Single smart model < Two specialized agents < Orchestrator + 3 teams ``` Benchmarks measure single models. Production wins with multi-agent orchestration. --- ## Lesson 4.2: Orchestration Patterns ### Pattern 1: Dispatcher (Hub-and-Spoke) ``` User → Orchestrator → dispatches to: ├── Specialist A (parallel) ├── Specialist B (parallel) └── Specialist C (parallel) ← synthesizes results ``` **Best for**: Tasks that need multiple perspectives or parallel work. The orchestrator doesn't work — it delegates, then synthesizes. **Implementation**: `agent-team` extension. Teams defined in `teams.yaml`. Orchestrator uses `dispatch_agent` tool. ### Pattern 2: Pipeline (Sequential) ``` Step 1: Planner → Step 2: Builder (gets $INPUT=plan) → Step 3: Reviewer (gets $INPUT=code) ``` **Best for**: Workflows with clear stages where each depends on the previous. **Implementation**: `agent-chain` extension. Pipelines in `agent-chain.yaml`. `$INPUT` carries forward, `$ORIGINAL` preserves user prompt. ### Pattern 3: Peer-to-Peer (Flat) ``` Agent A ←→ Agent B ←→ Agent C (no orchestrator, all peers) ``` **Best for**: Cross-device work, heterogeneous model teams, flat information flow. **Implementation**: `coms` / `coms-net` extensions. 4 tools: list, send, get, await. --- ## Lesson 4.2b: P-Threads — Parallel Agent Execution **Concept**: Run N agents simultaneously on the same task. Each agent works independently. Pick the best result. ### Pattern ``` User prompt │ ├── Agent 1 (Claude Opus) ──► output_1 ├── Agent 2 (Gemini Pro) ──► output_2 ├── Agent 3 (DeepSeek V3) ──► output_3 └── Agent 4 (Qwen) ──► output_4 │ ▼ Judge (or human) picks best ``` ### When to Use P-Threads - Creative work (UI generation, content writing, architecture design) - Benchmarking (compare models side-by-side on same task) - High-stakes decisions (multiple perspectives reduce blind spots) - Exploration (try N approaches, pick the winner) ### Implementation with mprocs ```yaml # mprocs-teams.yaml procs: claude-lead: cmd: ["claude", "-p", "{{PROMPT}}"] gemini: cmd: ["gemini", "-p", "{{PROMPT}}"] opencode: cmd: ["opencode", "--model", "opencode-go/deepseek-v4-flash", "-p", "{{PROMPT}}"] ``` Run all three simultaneously, collect outputs, pick best. ### Implementation with agent-mux Your Tauri UI (`agent-mux`) is a P-thread controller: it spawns agents into panes, monitors progress, and presents results for comparison. ### Key Design Rules - Each agent gets the **same prompt** (fair comparison) - Agents are **isolated** (no cross-contamination) - Results are **collected and compared** by a judge or human - Cost = N × single-agent cost (budget accordingly) --- ## Lesson 4.2c: F-Threads — Fusion (N Agents, One Winner) **Concept**: N agents produce outputs. A synthetic judge (or rubric) selects the winner. More sophisticated than P-threads — includes evaluation. ### Pattern ``` Agent 1 ──► Output 1 ──┐ Agent 2 ──► Output 2 ──┤ Agent 3 ──► Output 3 ──┤──► Judge → Winner Agent 4 ──► Output 4 ──┤ Agent 5 ──► Output 5 ──┘ ``` ### When to Use F-Threads - Code generation (generate N variants, pick best compiled/tested) - UI generation (your UI Agents system uses this — 3 Vue generators, 1 winner) - Architecture decisions (N proposals, judge evaluates against criteria) - Bug fixes (N approaches, pick the one that passes tests) ### Your UI Agents System as F-Thread Example Your `ui-agents` project uses F-threads natively: - 3 UI Generation agents (Sonnet + open source variants) - Validation team screenshots + checks each - Lead picks the best result - Loser outputs are discarded (or logged for learning) ### Judge Criteria Template ```json { "criteria": [ {"name": "correctness", "weight": 0.4}, {"name": "efficiency", "weight": 0.2}, {"name": "maintainability", "weight": 0.2}, {"name": "completeness", "weight": 0.2} ] } ``` --- ## Lesson 4.3: Depth-2 Delegation The production-proven hierarchy from lead-agents and ui-agents: ``` User │ v Orchestrator (Opus-level, thinking only) │ ├── Planning Team Lead (synthesizes, delegates) │ ├── Product Manager (domain worker) │ └── UX Researcher (domain worker) │ ├── Engineering Team Lead (synthesizes, delegates) │ ├── Frontend Dev (domain worker) │ └── Backend Dev (domain worker) │ └── Validation Team Lead (synthesizes, delegates) ├── QA Engineer (domain worker) └── Security Reviewer (domain worker) ``` ### The Rule **Leads and orchestrators are thinkers, planners, and managers. They are not doers.** They delegate to workers who write files and make changes. Workers don't make strategic decisions. --- ## Lesson 4.4: Agent Experts That Remember ### Mental Models Every agent maintains its own expertise file: ```yaml # engineering-lead-mental-model.yaml observations: - type: architectural_pattern observation: "We use vertical slice architecture for new features" evidence: "PR #142, PR #156" - type: common_failure observation: "WebSocket reconnection logic needs retry with backoff" evidence: "Incident log 2026-03-15" ``` ### Self-Improve Commands Agents periodically run self-improve commands that: 1. Read the mental model 2. Compare against actual codebase (grep, read files) 3. Update stale entries 4. Add new findings ### Compounding Knowledge Session 1: Agent learns project structure Session 2: Agent learns common patterns Session 3: Agent learns failure modes Session N: Agent operates at senior engineer level for this codebase --- ## Lesson 4.5: Domain Locking Every agent has explicit domain permissions: ```yaml # Engineering Lead domain domain: - path: .pi/multi-team/ read: true upsert: true delete: false - path: src/ read: true upsert: true delete: false - path: . read: true upsert: false delete: false # Can't delete anything at root ``` Three levels: - `read: false` — can't even see the files - `upsert: false` — can read, can't write - `delete: false` — can read+write, can't delete --- ## Lesson 4.6: TillDone Task Discipline ### The Pattern 1. Orchestrator creates task list before delegating 2. Tasks are assigned to specific agents 3. Agents mark tasks in progress → done 4. Footer shows live progress 5. If turn ends with incomplete tasks, agent is nudged to continue ### Why It Matters Prevents the #1 multi-agent failure mode: an orchestrator that delegates, gets results, and then wanders off without completing the plan. The task list is the forcing function. --- ## Lesson 4.7: Agent Chains ### YAML-Defined Pipelines ```yaml # agent-chain.yaml steps: - agent: planner prompt: "Create a detailed plan for: $INPUT" - agent: builder prompt: "Implement the following plan: $INPUT" - agent: reviewer prompt: "Review this code: $INPUT" ``` ### Variables - `$INPUT` — Output of the previous step - `$ORIGINAL` — The user's original prompt (always accessible) --- ## Lesson 4.8: Pi-to-Pi Communication ### The Shift: Hierarchy → Flat Traditional agent communication is top-down: ``` Orchestrator → Team Lead → Worker → (result goes back up) ``` Peer-to-peer flips this. Agents are equals: ``` Agent A ←→ Agent B ↕ ↕ Agent C ←→ Agent D ``` ### Reference Implementation: MCP Agent Mail (Jeff Emanuel, 1,955★) Jeff's MCP Agent Mail provides a more formal coordination layer: - **Agent identities**: Each agent has a registered identity - **Inboxes**: Messages are delivered to agent inboxes - **Searchable threads**: Full thread history with search - **Advisory file leases**: Agents can reserve files to prevent conflicts This complements our simpler P2P approach. Use direct coms for speed, MCP Agent Mail for audit trails. ### Four Tools | Tool | What It Does | |------|-------------| | `*_list` | List peer agents (name, model, context usage) | | `*_send` | Send a prompt; returns msg_id | | `*_get` | Non-blocking poll on msg_id | | `*_await` | Block until reply or timeout | ### Safety - Max 5 hops (prevents A→B→A→B loops) - Audit log (msg_id, sender, hops — never prompt bodies) - Self-healing (stale sockets pruned, heartbeats every 10s) --- ## Lesson 4.9: Conversation Awareness Every agent reads a shared JSONL conversation log before responding: ``` .pi/multi-team/sessions//conversation.jsonl ``` This means: - Agents know what's already been discussed - Agents don't repeat work - Full audit trail of every interaction ## Lesson 4.9b: Service Connectors Agents need to interact with external services: Twitter, GitHub, Linear, Slack, email. ### The Pattern Each service gets an MCP connector: ``` Agent → MCP Client → Service Connector → External API ├── twitter-connector: post, search, DM, timeline ├── github-connector: PRs, issues, actions, code search ├── linear-connector: issues, comments, projects, cycles ├── slack-connector: messages, channels, threads, search ├── email-connector: send, read, search, folders └── custom: build your own via connector SDK ``` ### Design Rules 1. **One connector per service** — composable, not monolithic 2. **MCP protocol** — standard interface for all connectors 3. **Scoped credentials** — each connector has its own auth (no shared tokens) 4. **Rate-limited** — connectors enforce service rate limits 5. **Audit-logged** — every external action is logged ### Reference Implementation Jeff Emanuel's `flywheel_connectors` (79★) provides a mesh-native protocol for this. His approach: - Each connector is a standalone Rust binary - Connectors communicate via a shared mesh protocol - Agents discover connectors dynamically - Built-in connectors for Twitter, Linear, GitHub ### Integration with Agent Teams ``` Orchestrator → delegates to agent Agent → calls service connector via MCP Connector → authenticates → calls API → returns result Agent → processes result → continues work Orchestrator → synthesizes final output ``` --- ## Lesson 4.10: The CEO Board System 8 specialist agents for strategic decision-making: | Board Member | Focus | Time Horizon | |-------------|-------|--------------| | Revenue | Cash flow, short-term wins | 30-90 days | | Compounder | Trust, long-term value | 6-24 months | | Contrarian | Assumptions, blind spots | 3x weight on dissent | | Technical Architect | System durability | Ongoing | | Product Strategist | Problem selection | Quarterly | | Customer Oracle | User behavior | Ongoing | | Market Strategist | Positioning | Quarterly | | Moonshot | Asymmetric upside | 1-5 years | ### Flow 1. CEO frames the decision 2. Board debates (sources required, no "I think") 3. Verifier checks facts (2+ sources per claim) 4. Executor creates execution plan 5. Tracker logs for quarterly review --- ## Lesson 4.11: UI Agents System 12 agents across 4 teams for brand-consistent UI generation: ``` Brand → Product → Tree → Branch → Leaf ``` Each level has its own `brand.yaml` with CSS custom properties. Zero hardcoded values. Agents generate Vue components that use only CSS custom properties from the brand config. ## Lab 4.12: Build an Agent Chain **Objective**: Create a plan→build→review pipeline in YAML. **Starter**: `course/labs/L4-agent-chain/starter.yaml` **Solution**: `course/labs/L4-agent-chain/solution.yaml` **Checkpoints**: 1. Planner produces structured plan with tasks and file paths 2. Builder creates code matching the plan 3. Reviewer identifies issues with severity (critical/major/minor) 4. Verifier confirms each claim with file:line evidence ## Lab 4.13: Deploy a Multi-Team System **Objective**: Set up orchestrator + 2 teams with domain locking. **Starter**: `course/labs/L4-multi-team/starter-config.yaml` **Solution**: `course/labs/L4-multi-team/solution-config.yaml` **Checkpoints**: 1. Orchestrator delegates, never executes 2. Each team has lead + members with distinct roles 3. Domain permissions restrict each agent to its scope 4. Mental model files exist for each agent