2.3 KiB
2.3 KiB
Autoresearch Kit
Agents that improve themselves. Run experiments, measure results, keep what works.
Included Skills
- Experiment Loop — Run → Measure → Log → Decide
- Integrity Guard — Code hashing to detect grinding
- Median Over Best — Compare against median, not best
- Reward Hack Defense — Detect timing manipulation
- Log to JSONL — Structured experiment logging
The Experiment Loop
1. RUN: Execute agent with current configuration
2. MEASURE: Collect metrics (latency, cost, success rate)
3. LOG: Record to experiments.jsonl
4. DECIDE: Keep if improvement, discard if regression
5. REPEAT: Try next experiment
Integrity Guards
Code Hashing (against grinding)
code_hash = sha256(open("agent.py").read()).hexdigest()[:12]
if last_run_code_hash == code_hash:
flag_grind_run() # Same code re-run — discard result
Median Over Best (against noise-chasing)
scores = [52, 48, 47, 46, 46, 51, 53]
median = sorted(scores)[len(scores)//2] # 48
best = max(scores) # 46 — don't use this
improvement = (baseline - median) / baseline # Use median
Test Set Integrity (against data leakage)
test_data_hash_before = hash(open("benchmark.json").read())
# ... run experiment ...
test_data_hash_after = hash(open("benchmark.json").read())
assert test_data_hash_before == test_data_hash_after
Experiment Log Format
{"run": 1, "status": "baseline", "metric": {"name": "latency", "value": 52, "unit": "ms"}, "description": "Initial measurement"}
{"run": 2, "status": "keep", "metric": {"name": "latency", "value": 46, "unit": "ms"}, "deltaPct": -11.5, "description": "Remove N+1 count queries", "code_hash": "a1b2c3d4"}
{"run": 3, "status": "discard", "metric": {"name": "latency", "value": 53, "unit": "ms"}, "deltaPct": +1.9, "description": "Add connection pool", "code_hash": "e5f6g7h8"}
Usage
python autoresearch.py --max-iterations 10
# Results in experiments.jsonl
cat experiments.jsonl | python -m json.tool
Related Course Material
- M7 Advanced Topics — Autoresearch lesson + lab
mythos-learnings.md— original research on reward hacking, grindingbrand-monitor/autoresearch.jsonl— real experiment data (52→46ms)