agentic-ai-engineering/course/ASSESSMENTS.md

12 KiB
Raw Blame History

Course Assessments

Module 1: Foundations Quiz

  1. What three components make an AI agent? a) LLM, Tools, Loop ✓ b) LLM, Database, UI c) Prompt, Context, Memory d) Model, API, Cache

  2. True or False: The quality of the model matters more than the design of the harness.

    • False ✓ — The harness determines cost, safety, and reliability. A good harness with a mediocre model outperforms a bad harness with the best model.
  3. Scenario: You need to process 10,000 customer support emails per month, categorizing them into "billing", "technical", or "other". Each misclassification costs $0.50 in manual re-routing. Should you use an agent? a) No, write deterministic code b) Yes, full autonomy agent c) Yes, agent with human review ✓ — Volume is high (10K/mo), but failure cost is non-zero. Agent with review is the sweet spot. d) No, do it manually

  4. What does the reasoning parameter do in tool calls? a) Enables chain-of-thought in the model b) Forces the agent to explain why it's calling each tool ✓ c) Provides error messages on failure d) Logs tool execution time

  5. Calculate: An agent has a 2% failure rate per turn and runs 50 turns. What's the probability of at least one failure?

    • P = 1 - (0.98)^50 ≈ 1 - 0.364 = 63.6%
  6. Which of the following is NOT a component of the agent loop? a) Think b) Act c) Compile ✓ d) Observe

  7. True or False: The harness is the product, not the model.

    • True ✓
  8. Scenario: Your CEO wants to use an agent for quarterly financial reporting. What's your recommendation? a) Full autonomy — agents can do anything b) Human-in-the-loop mandatory ✓ — Financial reporting is high-risk, needs human oversight c) Don't use an agent at all d) Use an agent but don't tell anyone


Module 2: Architecture Quiz

  1. What are the four pillars of every agent system? a) Tools, Loop, Context, Memory ✓ b) Model, API, Database, UI c) Speed, Cost, Quality, Safety d) Planning, Building, Testing, Deploying

  2. Which tool distribution channel has the HIGHEST context cost? a) CLI b) MCP Server ✓ c) File System Scripts d) Skills

  3. True or False: A generalist agent with many tools outperforms specialized agents on every task.

    • False ✓ — Specialized agents outperform generalists at their domain.
  4. Match the agent loop variant to its description:

    • Simple Prompt → Execute: Single API call, no loop
    • Tool-Use Loop: while not terminal_tool_called
    • Task-Completion: while not complete_task
    • Sub-Agent: Main agent spawns sub-calls
  5. What's the best practice for context window management? a) Keep everything forever b) Hybrid: system prompt + recent turns + summary + current results ✓ c) Summarize everything immediately d) Use the maximum context window size always

  6. Which codebase architecture is best for multiple independent agent capabilities? a) Atomic/Composable b) Layered c) Vertical Slice ✓ d) Pipeline

  7. What's the purpose of a mental model file? a) Store the agent's system prompt b) Let the agent accumulate domain knowledge across sessions ✓ c) Log tool call history d) Configure model parameters

  8. True or False: The reasoning parameter on tool calls reduces hallucinated tool calls.

    • True ✓ — Forcing the LLM to explain WHY reduces hallucinated calls by ~40%.

Module 3: Safety Quiz

  1. What is the L3 marquee break? a) Agent uses rm -rf directly b) Agent writes a Python script that does the destructive work, then runs it ✓ c) Agent ignores the system prompt d) Agent escapes the sandbox

  2. If an agent has a 1% per-turn failure rate, what's the probability of failure over 100 turns? a) 1% b) 10% c) ~63% ✓ d) 100%

  3. Which security level should be the MINIMUM for production systems? a) L1 (skill) b) L2 (system prompt) c) L3 (blacklist hook) ✓ d) L4 (whitelist hook) — Strongly preferred

  4. What are the three access levels in damage-control? a) Allow, Deny, Ask b) Zero-access, Read-only, No-delete ✓ c) Read, Write, Execute d) Public, Private, Secret

  5. True or False: The Verifier agent can write files.

    • False ✓ — The verifier has read-only tools by design.
  6. Which hook event fires BEFORE a tool executes and CAN block the call? a) UserPromptSubmit b) PreToolUse ✓ c) PostToolUse d) Stop

  7. What does the CONFIDENCE ladder measure? a) Model confidence in its answer b) Verifier's completeness + correctness of verification ✓ c) User satisfaction score d) Token efficiency

  8. Scenario: Your agent needs to run npm test and git status. Which L4 approach is correct? a) Whitelist exactly ^npm test$ and ^git status$ ✓ b) Allow all bash but ask for confirmation c) Use L2 system prompt saying "be careful" d) Whitelist ^uv run .*\.py$


Module 4: Orchestration Quiz

  1. What are the three orchestration patterns? a) Dispatcher, Pipeline, Peer-to-Peer ✓ b) Sequential, Parallel, Random c) Single, Dual, Multi d) Local, Remote, Hybrid

  2. In depth-2 delegation, what do Leads do? a) Write all the code themselves b) Synthesize and delegate, never execute ✓ c) Monitor the orchestrator d) Handle security

  3. True or False: In the agent-team dispatcher pattern, the orchestrator does the actual work.

    • False ✓ — The orchestrator delegates, it doesn't execute.
  4. What does $INPUT represent in agent-chain pipelines? a) The user's original prompt b) The output of the previous step ✓ c) Configuration variables d) Tool call results

  5. How are Pi-to-Pi agents different from hierarchical orchestration? a) They're faster b) They're peers, not parent-child ✓ c) They use a different API d) They only work on the same machine

  6. What prevents runaway A→B→A→B loops in Pi-to-Pi communication? a) Rate limiting b) Max 5 hops ✓ c) Session timeout d) Manual review

  7. How many board members does the CEO Board System use? a) 4 b) 6 c) 8 ✓ d) 10

  8. True or False: In the UI Agents system, brand identity uses hardcoded CSS values.

    • False ✓ — All styling uses CSS custom properties from brand.yaml. Zero hardcoded values.

Module 5: Production Quiz

  1. What's the 3x rule? a) Agents are 3x faster than humans b) Production agents cost 3x your prototype estimate ✓ c) You need 3 agents minimum d) 3 retries per task maximum

  2. What is a golden dataset? a) The most valuable customer data b) Curated input/behavior/output pairs for regression testing ✓ c) Training data for fine-tuning d) API benchmark results

  3. True or False: Shadow deployments run agent outputs in production.

    • False ✓ — Shadow deployments run identical inputs but outputs are NOT served to users.
  4. What four things must be rolled back for an agent? a) Prompt, Model, Parameters, Configuration ✓ b) Code, Database, UI, Tests c) API keys, Endpoints, Models, Prompts d) Git commit, Docker image, Environment, Config

  5. What is decision tracing? a) A/B test results b) Full chain of every tool call + LLM completion + decision point ✓ c) User behavior analytics d) Model training logs

  6. Which metric indicates a tool loop? a) High latency b) Same tool called 5+ times with same params ✓ c) High error rate d) Low token usage

  7. What's the difference between a warning and a hard stop budget policy? a) Warning logs, hard stop blocks b) Warning notifies, hard stop pauses the agent ✓ c) Warning is optional, hard stop is mandatory d) They're the same

  8. True or False: CI/CD for agents is the same as CI/CD for traditional software.

    • False ✓ — Agent CI/CD requires golden datasets, pass@k evaluation, and prompt versioning, not just unit tests.

Module 6: Economics Quiz

  1. What's the price range (cheapest to most expensive) of LLM models in $/M tokens? a) 2x b) 10x c) 100x ✓ (Gemini Flash at $0.15 to Claude Opus at $15/$75) d) 1000x

  2. What percentage of token cost is typically output tokens? a) 30% b) 50% c) 70% ✓ d) 90%

  3. What is cascade routing? a) Using multiple models in parallel b) Using cheaper models for simple steps and expensive models for complex steps ✓ c) Routing tasks to different servers d) Cascading agent failures

  4. What does pass@k measure? a) Percentage of tokens passed b) Probability that at least one of k attempts succeeds ✓ c) Pass rate of model benchmarks d) Password strength

  5. True or False: Automated evaluation catches all agent failures.

    • False ✓ — Human eval catches 2-3x more nuanced failures, especially reasoning quality and hallucination cascades.
  6. What's the recommended canary deployment traffic split? a) 1% b) 5% new config, 95% current ✓ c) 50/50 d) 100% new config immediately

  7. What is grinding in an agent context? a) The agent working slowly b) Re-running identical code hoping for different results ✓ c) CPU-intensive operations d) Token optimization

  8. Calculate: An agent session costs $0.05 prototype. What's the estimated production cost?

    • $0.05 × 3 = $0.15 (the 3x rule)

Module 7: Advanced Quiz

  1. What three integrity threats did the Mythos paper find in autoresearch? a) Data leaks, slow queries, bad models b) Reward hacking, grinding, test set leakage ✓ c) Prompt injection, jailbreaking, data poisoning d) Cost overruns, latency spikes, accuracy drops

  2. What's the best defense against grinding in autoresearch? a) Rate limiting b) Code hashing + comparing against median (not best) ✓ c) Manual review of every run d) Using a single model

  3. What does a meta-agent do? a) Manage other agents' memory b) Build new agents from documentation ✓ c) Monitor agent costs d) Handle authentication

  4. True or False: MCP is always better than CLI for tool distribution.

    • False ✓ — For 1-2 agents, CLI + prime prompt is often better due to lower context cost.
  5. Which approach has the LOWEST context cost for tool distribution? a) MCP Server b) CLI c) File System Scripts (progressive disclosure) ✓ d) All approaches have similar context costs

  6. What is the sentinel pattern in the Mac Mini Agent? a) A security guard for file access b) A marker pattern that makes async terminal work deterministic ✓ c) A notification system d) A model selection strategy

  7. True or False: Always-on agents run continuously without sleeping.

    • False ✓ — They use heartbeat execution: wake, work, sleep on a schedule.
  8. What's the purpose of code hashing in autoresearch? a) Version control for experiments b) Detect when identical code is being re-run (grinding) ✓ c) Cache experiment results d) Encrypt agent outputs


Capstone Evaluation Rubric

See M8-CAPSTONE.md for full rubric. Summary:

Area Weight Pass Target
Architecture 20% Diagram present Clear, justified
Implementation 25% Runs on happy path Handles errors
Security 20% L3 minimum L4+ with verifier
Testing 15% Computed pass@k pass@k + cost + grind detect
Documentation 10% Basic Architecture + retrospective
Cost 10% Cascade routing Cascade + verified savings

Answer Key

All quiz answers are marked with ✓ in the questions above.