Context Management
The model is rarely the bottleneck. What reaches it is.
I build software with Claude as a collaborator, and I keep a structured build log of every unit of work. For this paper I treated 133 of those entries, from STEAP and Workshop, as primary-source data on where an LLM development workflow actually fails, then built a retrieval prototype and tested it against my own documented failures.
"The right information existed in the corpus. It was simply not in front of the model at the right time."
Three Failure Classes
- Retrieval. The context existed but never surfaced. One string-vs-integer ID bug recurred four times on a single branch. A FIFO-ordering bug in a sync queue cost about four hours of speculation while the answer sat in the log.
- Staleness. The context surfaced but was no longer true. Withdrawn diagnoses and superseded architectures read almost exactly like valid lessons. Only the Status field and timestamps tell them apart.
- Routing. Context tangled across subsystems. I had hand-developed a three-prompt split (data table, backend route, frontend UI) to stop it, which is itself evidence that scoping carries real weight.
The Prototype
Logs are pulled once from Notion and cached with every schema field intact. A thin retrieve → act → verify loop hands Claude a scoped, ranked bundle of prior entries alongside the task. Writeback stays manual on purpose, to protect the corpus. I implemented the retrieval step three ways:
Arm A
Manual baseline: me as the retrieval layer
Arm B
Keyword index + structured re-ranking on Status and recency
Arm C
Vector retrieval by semantic similarity
What I Found
- Recurrence failures: clearly helped. Both automated arms surfaced the earlier ID-bug entries, including the one whose lesson proposes the real fix. Four debugging sessions could have been one.
- Staleness: only if retrieval respects supersession. The vector arm surfaced withdrawn entries as readily as valid ones, because similarity has no idea that one entry replaced another. The structured arm suppressed them.
- The case I didn't expect. One log entry held both a still-valid lesson and a superseded architecture. No single Status flag serves both. Finer-grained supersession matters more than the choice of retriever, and I would not have seen that without building it.
- Routing: mostly out of reach. A bounded bundle is not multi-prompt orchestration, and the paper says so.
- Limits. One developer, two logs, one model, and queries I wrote knowing the targets. It is a reasoned assessment, not a benchmark. A real evaluation needs held-out queries from someone who has never seen the corpus.
Where It Went Next
The paper's constraints became the way I actually run STEAP. Planning happens in chat; execution happens in Claude Code; a one-way Notion mailbox carries runbooks between them.
- Entries are never edited after creation. They move New → Ingested (committed into the repo, hash recorded) → Executed (stamped only once the merge is verified on main) or Superseded.
- A rules file sets the gates: a diff review before any merge, phone-approvable versus desk-review tiers, and "stop beats workaround."
- A headless GitHub Actions runner can only perform steps I have tagged for it, one step per run, and it never merges. Phone notifications tell me when something needs a decision.
- As of September 2026: 168 build-log entries and 99 mailbox entries on STEAP alone.