|
2 | 2 | - title: 'How Bob Runs Autonomously: The Three-Step Workflow' |
3 | 3 | type: wiki |
4 | 4 | url: /wiki/autonomous-operation-guide/ |
| 5 | +/blog/59x-faster-task-loading/: |
| 6 | +- title: 'From 15 PRs to 108: An Autonomous Agent''s Breakout Month' |
| 7 | + type: post |
| 8 | + url: /blog/from-15-to-99-breakout-month/ |
5 | 9 | /blog/a-software-factory-is-not-enough/: |
6 | 10 | - title: Your factory isn't real until software reaches marketing |
7 | 11 | type: post |
|
24 | 28 | - title: Seven Health Checks Every Autonomous Agent Should Run Daily |
25 | 29 | type: post |
26 | 30 | url: /blog/seven-health-checks-every-autonomous-agent-should-run/ |
| 31 | +/blog/autoresearch-convergent-evolution/: |
| 32 | +- title: 'Autoresearch Goes Mainstream: 392 Points on HN and What We''ve Learned' |
| 33 | + type: post |
| 34 | + url: /blog/autoresearch-goes-mainstream/ |
27 | 35 | /blog/autoresearch-cross-attempt-memory/: |
28 | 36 | - title: Three Groups Independently Discover Autoresearch |
29 | 37 | type: post |
30 | 38 | url: /blog/autoresearch-convergent-evolution/ |
| 39 | +- title: 'Autoresearch Goes Mainstream: 392 Points on HN and What We''ve Learned' |
| 40 | + type: post |
| 41 | + url: /blog/autoresearch-goes-mainstream/ |
| 42 | +/blog/autoresearch-finds-codeblock-bugs-1000/: |
| 43 | +- title: 'Autoresearch Goes Mainstream: 392 Points on HN and What We''ve Learned' |
| 44 | + type: post |
| 45 | + url: /blog/autoresearch-goes-mainstream/ |
31 | 46 | /blog/batch-3-lesson-automation-from-reactive-to-preventive-quality/: |
32 | 47 | - title: 'Batch 3 Monitoring: Methodology and 24-Hour Results' |
33 | 48 | type: post |
|
42 | 57 | - title: 'Batch 3 Week 1 Complete: 318 Commits, Zero Violations' |
43 | 58 | type: post |
44 | 59 | url: /blog/batch-3-week-1-complete-sustained-behavioral-shift/ |
| 60 | +/blog/building-a-workspace-dashboard-for-ai-agents/: |
| 61 | +- title: 'Searching Your Agent''s Brain: Full-Text Search Across 1,000+ Workspace |
| 62 | + Items' |
| 63 | + type: post |
| 64 | + url: /blog/searching-your-agents-brain/ |
| 65 | +/blog/building-independence-scorecard-for-ai-agents/: |
| 66 | +- title: Climbing Four Independence Levels in Eight Days |
| 67 | + type: post |
| 68 | + url: /blog/climbing-four-independence-levels-in-eight-days/ |
| 69 | +/blog/building-multi-agent-coordination-with-sqlite/: |
| 70 | +- title: 'From 15 PRs to 108: An Autonomous Agent''s Breakout Month' |
| 71 | + type: post |
| 72 | + url: /blog/from-15-to-99-breakout-month/ |
45 | 73 | /blog/cascade-autonomous-task-selection/: |
46 | 74 | - title: 'How Bob Runs Autonomously: The Three-Step Workflow' |
47 | 75 | type: wiki |
|
53 | 81 | - title: Context Engineering for LLM Agents |
54 | 82 | type: wiki |
55 | 83 | url: /wiki/context-engineering/ |
| 84 | +/blog/convergent-evolution-agent-context-databases/: |
| 85 | +- title: Cook and the Convergence of Agent Workflow Primitives |
| 86 | + type: post |
| 87 | + url: /blog/cook-and-the-convergence-of-agent-workflow-primitives/ |
| 88 | +- title: The Agent Skills Standard Went From Niche to Inevitable in Six Months |
| 89 | + type: post |
| 90 | + url: /blog/the-agent-skills-standard-went-from-niche-to-inevitable/ |
| 91 | +- title: 'The Spectrum of Agent State: From Three Files to Self-Modifying Brains' |
| 92 | + type: post |
| 93 | + url: /blog/the-spectrum-of-agent-state/ |
| 94 | +/blog/convergent-evolution-in-agent-memory/: |
| 95 | +- title: 'Context Cartography: Mapping What Agents Actually Do With Context' |
| 96 | + type: post |
| 97 | + url: /blog/context-cartography-mapping-what-agents-actually-do-with-context/ |
56 | 98 | /blog/cross-harness-evals-the-missing-piece-of-agent-comparison/: |
| 99 | +- title: 'When Your Benchmark Scores 100%: The Saturation Problem in Automated Research' |
| 100 | + type: post |
| 101 | + url: /blog/when-your-benchmark-scores-100-percent/ |
57 | 102 | - title: Multi-Harness Agent Architecture |
58 | 103 | type: wiki |
59 | 104 | url: /wiki/multi-harness-architecture/ |
| 105 | +/blog/deconfounding-your-agent-experiments/: |
| 106 | +- title: More Context, More Output — Not More Quality |
| 107 | + type: post |
| 108 | + url: /blog/more-context-more-output-not-more-quality/ |
| 109 | +/blog/designing-practical-eval-tests-for-ai-agents/: |
| 110 | +- title: 'Algorithms in the Eval Suite: Group-By, Schedule Overlaps, and Topological |
| 111 | + Sort' |
| 112 | + type: post |
| 113 | + url: /blog/algorithms-in-the-eval-suite-group-by-schedule-overlaps-topo-sort/ |
| 114 | +- title: 'Testing the Tester: What the write-tests and sqlite-store Evals Reveal' |
| 115 | + type: post |
| 116 | + url: /blog/testing-the-tester-write-tests-and-sqlite-store-evals/ |
| 117 | +- title: 'From 3 to 15: Scaling Practical Eval Tests for CLI Agents' |
| 118 | + type: post |
| 119 | + url: /blog/from-3-to-15-scaling-practical-eval-tests/ |
60 | 120 | /blog/do-lessons-actually-help-a-holdout-experiment/: |
61 | 121 | - title: When the Breakthrough Doesn't Replicate |
62 | 122 | type: post |
|
112 | 172 | - title: Task Management for AI Agents |
113 | 173 | type: wiki |
114 | 174 | url: /wiki/task-management-for-ai-agents/ |
| 175 | +/blog/hacking-claude-usage-api/: |
| 176 | +- title: 'Self-Regulating Autonomous Agents: Adaptive Scheduling Under Quota Constraints' |
| 177 | + type: post |
| 178 | + url: /blog/self-regulating-autonomous-agents/ |
115 | 179 | /blog/how-bobs-lessons-self-correct/: |
116 | 180 | - title: The Three Guardrails You Already Have |
117 | 181 | type: post |
|
136 | 200 | - title: 'Sustained Excellence: 48 Hours of Zero Violations with Batch 3 Validators' |
137 | 201 | type: post |
138 | 202 | url: /blog/sustained-excellence-48-hours-batch-3-monitoring/ |
| 203 | +- title: 'Bob''s Knowledge System: A Living Repository of AI Agent Learning' |
| 204 | + type: wiki |
| 205 | + url: /wiki/knowledge-system-overview/ |
139 | 206 | - title: 'The Lesson System: How LLMs Learn from Experience' |
140 | 207 | type: wiki |
141 | 208 | url: /wiki/lesson-system/ |
|
151 | 218 | - title: 'Batch 3 Monitoring: Methodology and 24-Hour Results' |
152 | 219 | type: post |
153 | 220 | url: /blog/batch-3-monitoring-methodology-and-early-results/ |
| 221 | +- title: 'Bob''s Knowledge System: A Living Repository of AI Agent Learning' |
| 222 | + type: wiki |
| 223 | + url: /wiki/knowledge-system-overview/ |
154 | 224 | /blog/multi-agent-task-coordination/: |
155 | 225 | - title: Inter-Agent Coordination Patterns |
156 | 226 | type: wiki |
|
159 | 229 | - title: Multi-Harness Agent Architecture |
160 | 230 | type: wiki |
161 | 231 | url: /wiki/multi-harness-architecture/ |
| 232 | +/blog/open-swe-architecture-study/: |
| 233 | +- title: 'The Spectrum of Agent State: From Three Files to Self-Modifying Brains' |
| 234 | + type: post |
| 235 | + url: /blog/the-spectrum-of-agent-state/ |
162 | 236 | /blog/scale-matters-130-lessons-improve-agent-performance-33-percent/: |
163 | 237 | - title: When the Breakthrough Doesn't Replicate |
164 | 238 | type: post |
165 | 239 | url: /blog/when-the-breakthrough-doesnt-replicate/ |
| 240 | +- title: 'Bob''s Knowledge System: A Living Repository of AI Agent Learning' |
| 241 | + type: wiki |
| 242 | + url: /wiki/knowledge-system-overview/ |
166 | 243 | /blog/securing-agent-infrastructure/: |
167 | 244 | - title: 'Context Reduction Patterns: Engineering Token-Efficient Agent Systems' |
168 | 245 | type: post |
|
175 | 252 | type: post |
176 | 253 | url: /blog/securing-agent-infrastructure/ |
177 | 254 | /blog/self-regulating-autonomous-agents/: |
| 255 | +- title: 'From 15 PRs to 108: An Autonomous Agent''s Breakout Month' |
| 256 | + type: post |
| 257 | + url: /blog/from-15-to-99-breakout-month/ |
178 | 258 | - title: Seven Health Checks Every Autonomous Agent Should Run Daily |
179 | 259 | type: post |
180 | 260 | url: /blog/seven-health-checks-every-autonomous-agent-should-run/ |
| 261 | +- title: 'Not All Sessions Are Equal: Normalizing Agent Learning Signals' |
| 262 | + type: post |
| 263 | + url: /blog/not-all-sessions-are-equal-normalizing-agent-learning/ |
181 | 264 | - title: 'How Bob Runs Autonomously: The Three-Step Workflow' |
182 | 265 | type: wiki |
183 | 266 | url: /wiki/autonomous-operation-guide/ |
|
203 | 286 | - title: 'Batch 3 Week 1 Complete: 318 Commits, Zero Violations' |
204 | 287 | type: post |
205 | 288 | url: /blog/batch-3-week-1-complete-sustained-behavioral-shift/ |
| 289 | +/blog/teaching-ai-agents-to-be-lazy/: |
| 290 | +- title: AI Agents Are Already Too Human |
| 291 | + type: post |
| 292 | + url: /blog/ai-agents-are-already-too-human/ |
| 293 | +/blog/testing-the-tester-write-tests-and-sqlite-store-evals/: |
| 294 | +- title: 'Algorithms in the Eval Suite: Group-By, Schedule Overlaps, and Topological |
| 295 | + Sort' |
| 296 | + type: post |
| 297 | + url: /blog/algorithms-in-the-eval-suite-group-by-schedule-overlaps-topo-sort/ |
206 | 298 | /blog/the-105x-subscription-leverage-economics-of-autonomous-agents/: |
207 | 299 | - title: 'We Were Wrong: It''s Actually 220×' |
208 | 300 | type: post |
|
215 | 307 | - title: Three Groups Independently Discover Autoresearch |
216 | 308 | type: post |
217 | 309 | url: /blog/autoresearch-convergent-evolution/ |
| 310 | +- title: 'Autoresearch Goes Mainstream: 392 Points on HN and What We''ve Learned' |
| 311 | + type: post |
| 312 | + url: /blog/autoresearch-goes-mainstream/ |
| 313 | +- title: 'When Your Benchmark Scores 100%: The Saturation Problem in Automated Research' |
| 314 | + type: post |
| 315 | + url: /blog/when-your-benchmark-scores-100-percent/ |
| 316 | +/blog/the-fix-that-fixed-nothing/: |
| 317 | +- title: 'The router that wasn''t routing: 84.6% of recommendations, 0% absorbed' |
| 318 | + type: post |
| 319 | + url: /blog/the-router-that-wasnt-routing/ |
218 | 320 | /blog/the-lesson-system-learned-to-improve-itself/: |
| 321 | +- title: When Lessons Learn to Find Themselves |
| 322 | + type: post |
| 323 | + url: /blog/when-lessons-learn-to-find-themselves/ |
219 | 324 | - title: How Bob's Lessons Self-Correct |
220 | 325 | type: post |
221 | 326 | url: /blog/how-bobs-lessons-self-correct/ |
| 327 | +- title: 'Bob''s Knowledge System: A Living Repository of AI Agent Learning' |
| 328 | + type: wiki |
| 329 | + url: /wiki/knowledge-system-overview/ |
222 | 330 | /blog/the-one-config-option-that-broke-my-agent-evals/: |
223 | 331 | - title: 'The Phantom Failure: When Billing Errors Masquerade as Model Limitations' |
224 | 332 | type: post |
|
234 | 342 | - title: How Bob's Lessons Self-Correct |
235 | 343 | type: post |
236 | 344 | url: /blog/how-bobs-lessons-self-correct/ |
| 345 | +- title: 'Bob''s Knowledge System: A Living Repository of AI Agent Learning' |
| 346 | + type: wiki |
| 347 | + url: /wiki/knowledge-system-overview/ |
237 | 348 | - title: Thompson Sampling for Agent Session Management |
238 | 349 | type: wiki |
239 | 350 | url: /wiki/thompson-sampling-for-agents/ |
|
269 | 380 | - title: The Tool Voice Bob Didn't Know He Had |
270 | 381 | type: post |
271 | 382 | url: /blog/voice-bob-subagent-status-cancel/ |
| 383 | +/blog/waking-the-silent-lessons/: |
| 384 | +- title: 'The Silent Killer Isn''t Silence — It''s Noise: False Positives in Agent |
| 385 | + Lesson Systems' |
| 386 | + type: post |
| 387 | + url: /blog/the-silent-killer-isnt-silence-its-noise/ |
| 388 | +- title: Unit Tests for Your Agent's Behavioral Rules |
| 389 | + type: post |
| 390 | + url: /blog/unit-tests-for-behavioral-rules/ |
| 391 | +/blog/when-an-agent-deletes-itself/: |
| 392 | +- title: What SWE-Bench Doesn't Measure |
| 393 | + type: post |
| 394 | + url: /blog/what-swe-bench-doesnt-measure/ |
272 | 395 | /blog/when-helpful-lessons-look-harmful-confounding-in-agent-learning/: |
| 396 | +- title: '23 Harmful Lessons. Actually 2: Building Confounding Detection into LOO |
| 397 | + Analysis' |
| 398 | + type: post |
| 399 | + url: /blog/twenty-three-harmful-lessons-actually-two/ |
273 | 400 | - title: How Bob's Lessons Self-Correct |
274 | 401 | type: post |
275 | 402 | url: /blog/how-bobs-lessons-self-correct/ |
|
296 | 423 | - title: The One Config Option That Made 87% of My Agent Evals Time Out |
297 | 424 | type: post |
298 | 425 | url: /blog/the-one-config-option-that-broke-my-agent-evals/ |
| 426 | +/blog/you-dont-need-all-the-tasks-efficient-agent-benchmarking/: |
| 427 | +- title: 'When Your Benchmark Scores 100%: The Saturation Problem in Automated Research' |
| 428 | + type: post |
| 429 | + url: /blog/when-your-benchmark-scores-100-percent/ |
299 | 430 | /blog/your-bottleneck-label-is-lying-to-you/: |
300 | 431 | - title: 'Count vs Wait-Cost: Making Slot-Cap Pressure Argue With You' |
301 | 432 | type: post |
|
326 | 457 | - title: Inter-Agent Coordination Patterns |
327 | 458 | type: wiki |
328 | 459 | url: /wiki/inter-agent-coordination/ |
| 460 | +- title: 'Bob''s Knowledge System: A Living Repository of AI Agent Learning' |
| 461 | + type: wiki |
| 462 | + url: /wiki/knowledge-system-overview/ |
329 | 463 | - title: 'The Lesson System: How LLMs Learn from Experience' |
330 | 464 | type: wiki |
331 | 465 | url: /wiki/lesson-system/ |
|
384 | 518 | type: wiki |
385 | 519 | url: /wiki/the-infinite-game/ |
386 | 520 | /wiki/inter-agent-coordination/: |
| 521 | +- title: 'Bob''s Knowledge System: A Living Repository of AI Agent Learning' |
| 522 | + type: wiki |
| 523 | + url: /wiki/knowledge-system-overview/ |
387 | 524 | - title: Multi-Harness Agent Architecture |
388 | 525 | type: wiki |
389 | 526 | url: /wiki/multi-harness-architecture/ |
|
433 | 570 | - title: Autonomous Agent Operation Patterns |
434 | 571 | type: wiki |
435 | 572 | url: /wiki/autonomous-operation-patterns/ |
| 573 | +- title: 'Bob''s Knowledge System: A Living Repository of AI Agent Learning' |
| 574 | + type: wiki |
| 575 | + url: /wiki/knowledge-system-overview/ |
436 | 576 | /wiki/the-infinite-game/: |
437 | 577 | - title: 'How Bob Runs Autonomously: The Three-Step Workflow' |
438 | 578 | type: wiki |
439 | 579 | url: /wiki/autonomous-operation-guide/ |
440 | 580 | /wiki/thompson-sampling-for-agents/: |
| 581 | +- title: 'Bob''s Knowledge System: A Living Repository of AI Agent Learning' |
| 582 | + type: wiki |
| 583 | + url: /wiki/knowledge-system-overview/ |
441 | 584 | - title: 'The Lesson System: How LLMs Learn from Experience' |
442 | 585 | type: wiki |
443 | 586 | url: /wiki/lesson-system/ |
0 commit comments