Part 06 of 13 · Memory & Learning

How a correction becomes reusable knowledge

A reviewer’s correction becomes reusable only after quality, scope and permission checks not merely because it was observed.

Read the article on Substack ↗ (opens in a new tab)

Summary

A reviewer notices that the extraction agent recorded the service-period start as the invoice date, corrects the field and approves the case. Keeping the corrected case is obviously useful; the interesting question is whether that correction should change what happens to the fifth invoice from the same supplier.

Evidence arrives from both sides of the platform — agent executions and outputs, human reviews and decisions, workflow edits and design-session interactions, scoring outcomes, failures, cost, latency, integration and marketplace activity — which is what allows the system to ask whether a recurring correction became less frequent after a design change. Capture, though, is not learning: the trace contains the agent’s wrong output alongside the reviewer’s fix, so feeding it straight into future behaviour would teach the mistake and the correction together.

Interpretation happens through agent scoring (grading executions on precision, accuracy, completeness, speed, cost and review outcome), human-review analysis (distinguishing consensus about a repeatable error from contradiction worth surfacing), and workflow insights (evaluating versions against an evaluation framework, policy and historical performance). What emerges is a candidate pattern, which must then pass the promotion gate — the mechanism every earlier article referred to — answering three questions: is it good enough, where may it be used, and who may see it. Only reviewed traces, corrected outputs and evaluated patterns pass; a reviewed failure can still serve as negative evidence under the same conditions.

Promoted learnings are indexed for semantic retrieval alongside workflows, agents and code, with the Work Knowledge Graph supplying structural context, so the runtime cognitive layer and design-time agents can both use them. Success is measured not by events captured or patterns promoted but by repeat-error rate and reviewer effort — and the article says explicitly that it describes the mechanism rather than reporting results, while naming the real risk the gate reduces without removing: mistaken generalization from one supplier to all.

All thirteen parts →