Non-Technical Track: Decision Frameworks for AI Agents
For Managers, PMs, and Decision-Makers
Framework 1: "Should We Use an Agent for This?"
The 4-Question Filter (for non-technical stakeholders)
Q1: Can a human do this task with clear step-by-step instructions?
- Yes → Could be agentizable
- No → Humans struggle too; don't agentize
Q2: Does the task have clear success criteria?
- Yes → We'll know if the agent succeeded
- No → You can't evaluate an agent if you can't define "done"
Q3: What's the cost of failure?
- Low (wrong suggestion, formatting error) → Full autonomy possible
- Medium (wrong code, wrong data) → Human review required
- High (financial loss, safety, legal) → Human-in-the-loop mandatory
Q4: How many times does this task need to be done?
- Once → Just do it manually
- 10x → Build a script
- 100x → Build an agent
- 1000x → Build a multi-agent system
Decision Matrix
Low Failure Cost High Failure Cost
──────────────┼─────────────────────────────────────
High Volume │ Build agent │ Agent + HITL
Low Volume │ Do it manually │ Do it manuallyFramework 2: The 4 Agent Design Patterns (Andrew Ng)
1. Reflection
The agent checks its own work before presenting it.
Agent writes code → reviews own code → fixes issues → presents finalBusiness value: Reduces human review time by 40-60%. Catches obvious errors before they reach you.
2. Tool Use
The agent uses external tools to gather information or take actions.
Agent reason → call search API → read results → call database → compose answerBusiness value: Agents can access real-time data, not just training data. Turns a static LLM into a live system.
3. Planning
The agent decomposes a complex request into sub-steps.
"Build a landing page" → plan steps (1. Design, 2. Code, 3. Test, 4. Deploy)
→ execute each step → verify resultsBusiness value: Complex tasks get done without step-by-step human instruction. The agent figures out the steps.
4. Multi-Agent Collaboration
Multiple specialized agents work together on different parts of a task.
PM agent plans → Designer creates mockup → Frontend builds → QA testsBusiness value: Each agent is an expert in its domain. Together they outperform one generalist agent at a fraction of the cost.
Framework 3: Cost Estimation for Agent Projects
The 3x Rule
A production agent costs 3x your prototype estimate. Plan accordingly.
| Phase | Cost Multiplier | What's Included |
|---|---|---|
| Prototype (works on happy path) | 1x | Single agent, one model, no safety |
| Production (handles edge cases) | 2x | Retry logic, error handling, testing |
| Production+ (operational) | 3x | Monitoring, security, CI/CD, budgets |
Quick Cost Calculator
Per-task cost = (input_tokens × input_price + output_tokens × output_price) × retry_rate × overhead
Where:
retry_rate = 1.5 (typical for agent systems)
overhead = 3.0 (production multiplier)
Example:
Agent session: 4000 input tokens, 1000 output tokens
Model: Claude Sonnet ($3/$15 per M tokens)
Base cost: (4000 × $3/M + 1000 × $15/M) = $0.012 + $0.015 = $0.027
With retries: $0.027 × 1.5 = $0.041
Production cost: $0.041 × 3.0 = $0.123 per task
At 1000 tasks/month: $123/month
At 10000 tasks/month: $1230/monthCost by Autonomy Level
| Level | Cost/Task | Human Time/Task | Best For |
|---|---|---|---|
| No agent | $0 | 30 min | One-off tasks |
| AI-assisted | $0.01-0.05 | 10 min | Human does the work, AI helps |
| Agent with review | $0.05-0.30 | 5 min | Agent works, human reviews |
| Full autonomy | $0.10-1.00 | 0 min | Agent works, human audits |
| Multi-agent | $0.50-5.00 | 2 min | Complex workflows, human supervises |
Framework 4: Autonomy vs Risk Matrix
Low Risk Medium Risk High Risk
(suggestions, (code gen, data (financial, medical,
content gen) processing) infrastructure)
─────────────┼─────────────────────────────────────────────────────────────
Tier 1: │ Full autonomy │ Agent + review │ Human only
Simple │ │ │
│ │ │
Tier 2: │ Agent + review │ Agent + verifier + │ Human + agent
Complex │ │ human spot-check │ as tool
│ │ │
Tier 3: │ Multi-agent + audit │ Multi-agent + │ Not recommended
Strategic │ │ verifier + human │ for agentsFramework 5: The "Build vs Buy" Decision
Build an Agent When:
- Your task is unique to your business
- You need to keep data in-house
- No existing tool solves the problem
- You have the engineering capacity
Buy/Rent When:
- It's a common task (customer support, code review, monitoring)
- A SaaS product already exists
- You don't have AI engineering expertise
- Speed to market matters more than customization
Hybrid When:
- Core task is unique, but supporting tasks are common
- You can build a small agent that calls existing SaaS tools
Framework 6: Measuring Success
Leading Indicators (week 1-4)
- Task completion rate (is the agent finishing its work?)
- Human review time reduction
- Cost per task vs cost of human doing it
- Error rate (how often does the agent need correction?)
Lagging Indicators (month 2-3+)
- Time saved per week
- Quality compared to human baseline
- Number of tasks automated vs manual
- Maintenance burden (how often does the agent break?)
Red Flags
- Agent costs more than the human it replaces
- Agent requires more oversight than doing it yourself
- Agent breaks every time the model updates
- Team doesn't trust the agent's output
Framework 7: Governance & Compliance Quick Reference
Regulatory Requirements by Region
| Requirement | EU (AI Act) | California (SB 1047) | UK/Japan |
|---|---|---|---|
| Human oversight | Mandatory for high-risk | Mandatory kill-switch | Voluntary |
| Audit trails | Required | Required | Recommended |
| Incident reporting | Required | Required | Voluntary |
| Disclosure (AI-generated) | Required | Required | Recommended |
| Data privacy (GDPR) | Full compliance | Similar | Different standard |
Minimum Viable Governance
For teams not in regulated industries, start with:
- Every agent action is logged with timestamp, agent, input, output
- Human review gate for destructive operations (deletes, writes, spending money)
- Monthly cost report per agent per task type
- Quarterly security review — has the agent's behavior changed?
Summary: Decision Flow
1. Can this task be automated?
├── No → Do it manually
└── Yes →
2. What's the failure cost?
├── High → Human-in-the-loop mandatory
└── Low →
3. What's the volume?
├── <10/month → Do it manually
├── 10-100/month → Build a script
└── >100/month → Build an agent
4. How many specialists needed?
├── One → Single agent
└── Multiple → Multi-agent system
5. What's the budget?
├── <$50/month → Use cheapest model (Gemini Flash)
├── $50-500/month → Cascade routing (mix models)
└── >$500/month → Multi-agent with verification