8.4 KiB
Loop Engineering Meta-Prompt
Use this meta-prompt for every TAC Product Factory / Plan F3 plan that touches agents, deploy, RSI, or self-improvement.
Required loop sections
Every plan must explicitly map the work to four loops:
- Agent loop — what agent/tool loop runs until done.
- Verification loop — what deterministic checks, reviewers, receipts, or rubrics gate success.
- Event-driven loop — what webhook, cron, queue, file event, or dashboard trigger runs the work without manual prompting.
- Hill-climbing loop — what trace/evidence can safely modify prompts, skills, tools, or config later.
Required safety fields
Every plan/receipt must include:
- Canonical deploy route:
git-proxy:8099/deploy. - Legacy note:
deploy-webhook:8098is internal/stale for guarded product deploy. auto_patch_proven:falseunless a degraded-skill recovery receipt exists.rsi_canary_recovery_evidence:trueonly when linking the canary receipt.- Evidence receipt paths for any RSI/self-improvement claim.
Required validation
At minimum, plans must name the checks that prove:
- plan schema/sections validate;
- deploy route is signed-action-only;
- tests pass;
- claims do not exceed receipts.
Taste rule
Prefer the boring loop that compounds over the clever prompt that works once. If the plan needs a new framework, first prove a markdown file, JSON receipt, and pytest assertion cannot hold the invariant.
Agent harness note
A loop is not just repeated model calls. Every agent plan must name the harness around the model:
- context and working memory;
- durable semantic memory / RAG sources;
- episodic memory / traces from prior runs;
- tool allowlist and signed deploy boundaries;
- end-loop guardrails that define when to stop;
- tracing for retrievals, tool calls, latency, token use, and errors;
- evals/receipts that decide whether prompt or config changes can feed back into the next run.
LLMOps/hill-climbing only promotes a changed prompt, retrieval config, tool config, or model parameter when traces plus evals show the new loop is safer or better.
Reference: Sean's AI Stories, "You Can Learn AI Agent Harness & Loop Engineering In 19 Min" (YouTube GrNbuWWJYiI).
Open-source model portability note
A strong harness should make the task portable across model tiers, including open-source or lower-cost models. Plans must spell out enough context, phases, and quality gates that a non-frontier model can follow the work without relying on hidden taste.
For outbound/productized work, require:
- explicit prospect/input target;
- crawler or source gathering step;
- report generation step;
- cover letter or executive summary when the output is meant to sell or persuade;
- design/report review phase;
- quality gate checklist that loops back to the builder when review fails;
- model/cost notes when choosing GLM/open-source/local models over frontier models.
Reference: Jordan Urbs, "GLM 5.2 Proves Open Source AI Can Match Fable 5 (AI Harness Engineering)" (YouTube dJI2GRG1GEE).
Autonomous work agent note
For business adoption, distinguish two operating modes before designing the harness:
- Team manages agents — an agent factory or admin team creates, updates, and governs shared agents.
- Agents assist people — each human gets a personal work agent that learns their context, skills, preferences, and second-brain material.
Every autonomous work-agent plan must include:
- per-agent isolation boundary, preferably one container or equivalent sandbox per agent;
- credential separation: no raw secrets inside the agent environment;
- vault/proxy credential injection for only the requests the agent is allowed to make;
- explicit access policies for tools, files, APIs, and outbound network;
- skill/instruction iteration loop for tuning the agent's actual output;
- management surface for upgrading, maintaining, and revoking agents after deployment;
- second-brain/context source plan when personal-agent behavior depends on private knowledge.
Reference: Latent Space, "The Blueprint for Autonomous Work Agents" with Gavriel Cohen / NanoClaw (YouTube hLUGXO5DSpo).
Product work with coding agents note
When implementation gets cheap, do not delete product discipline. Plans should use prototypes to explore, but still preserve the product loop:
- PRD/spec captures why, users, constraints, and success criteria;
- prototype makes options concrete and testable;
- every artifact declares its stage: exploration, prototype, beta, or ship candidate;
- production-looking prototypes must not imply production readiness without receipts;
- human taste and systems thinking review whether the result fits the whole product;
- scheduled/background agents may gather context, but promotion still needs receipts and review;
- role boundaries can flex, but ownership, decision rights, and phase state must stay explicit;
- model-timing assumptions are explicit: some artifacts are kept to retest when model capability changes;
- complexity budget is tracked because autonomous loops tend to add code; deletion/simplification is a first-class review criterion.
Reference: Lenny's Podcast, "OpenAI Codex lead on the new shape of product work" with Andrew Ambrosino (YouTube P3KDebPTUrw).
Activity modes note
Treat AI-era role archetypes as activity modes in the delivery loop, not permanent job titles. A person or agent can move across modes, but each mode needs its own invariant:
- Prototyper — creates options and working artifacts; invariant: stage is marked and prototype output cannot skip PRD/spec or receipts.
- Grower — compounds a rough artifact into a durable product; invariant: every iteration preserves traceability, tests, and user/problem fit.
- Sweeper — deletes, simplifies, migrates, and pays down prototype debt; invariant: complexity budget and deletion/simplification review are explicit.
- Reviewer/Taster — applies human taste, systems thinking, security, and product judgment; invariant: production readiness requires review evidence, not vibes.
- Architect/Primitive keeper — maintains shared primitives, interfaces, and harness boundaries; invariant: tools, credentials, deploy actions, and context sources stay governed.
- Operator/SRE — keeps the loop running in production; invariant: observability, rollback, incident ownership, and revocation paths are named.
Do not claim role collapse means expertise is obsolete. The safer claim is that boundaries flex while specialties, ownership, and best practices remain necessary.
External research anchors
Ground agent plans against external production guidance, not just video notes:
- NVIDIA Secure Agent Workspace: always-on agents need a managed workspace with credential proxy, enterprise tool access, governance, and runtime controls.
- Infisical Agent Vault: secrets stay in the vault and requests receive credentials through a proxy; agents must not hold raw secrets.
- Anthropic agent evals: judge trajectories and environment outcomes, not just final answers.
- OWASP AI Agent Security Cheat Sheet: treat prompt injection, memory poisoning, tool abuse, sandboxing, and access control as first-class risks.
These anchors support TAC's local rule: deploy remains signed-action-only through git-proxy:8099/deploy; do not reintroduce raw command deploy paths.
References: NVIDIA Secure Agent Workspace (docs.nvidia.com/enterprise-reference-architectures/secure-agent-workspace-reference-design), Infisical Agent Vault (infisical.com/blog/agent-vault-the-open-source-credential-proxy-and-vault-for-agents), Anthropic "Demystifying evals for AI agents" (anthropic.com/engineering/demystifying-evals-for-ai-agents), OWASP AI Agent Security Cheat Sheet (cheatsheetseries.owasp.org/cheatsheets/AI_Agent_Security_Cheat_Sheet.html).
Model Workspace Protocol note
For sequential workflows with human review between stages, prefer folder-structured orchestration before multi-agent framework code. The Model Workspace Protocol pattern treats numbered folders as stages, markdown files as role/context carriers, and local scripts as the boring mechanical layer.
Use this when it fits:
00-intake/— raw request, sources, constraints.10-plan/— Plan F3 artifact and assumptions.20-build/— implementation notes and changed files.30-verify/— test output, verifier notes, receipts.40-ship/— signed deploy request, outcome, rollback notes.
Reference: arXiv 2603.16021, "Interpretable Context Methodology: Folder Structure as Agentic Architecture".