Published on Medium ·

How Task Knowledge Makes Small Models Practical

Define the task, evaluate the model and preserve the approved assignment, so compatible workflows can reuse it at execution time.

Read the post on Medium ↗ (opens in a new tab)

Summary

A workflow often knows a task’s responsibility before it sees the next case. Extracting specified claim fields is a narrower job than planning the whole claims process, and that distinction makes it possible to evaluate a small language model for the particular work it will carry. The article describes how Opus Remodel connects that evaluation to the Work Knowledge Graph, preserving an approved model assignment alongside the reusable task and the evidence that supports it.

The assignment has a scope. A shared task name is insufficient: the schema, document population, language, data restrictions and operating requirements must be compatible. A route evaluated on short English forms cannot simply be applied to scanned Arabic documents. The approved binding records model, prompt and schema versions, input limits and fallback rules, and each workflow can pin the version it uses. A routing table handles dispatch; the graph supplies the shared task identity and evidence needed to judge where reuse is justified.

Routine execution then becomes a lookup and bounded applicability checks. Where the required decision inputs are already available, no model call is needed merely to choose the default model. Determinism makes that selection reproducible for the same policy and inputs; it does not make the generated answer reproducible or correct. Unsupported inputs, unavailable models and failed checks need explicit routes to an eligible alternative, human review or a controlled stop.

Quality is a condition of the choice. The smaller model must meet the same task requirements on representative held-back inputs and difficult cases, with a stated tolerance for degradation. Evaluation also has to cover the workflow: an extraction error can pass a local check and mislead a later task. The economic measure is cost per successfully completed case, including retries, fallback, hosting and reviewer time. A larger model can remain the cheaper choice once those costs are counted.

Failures, reviewer corrections and changes in the input population feed reassessment outside routine dispatch. The Snowflake-backed Remodel design keeps datasets and evaluation traces for that work, while approval produces a new route version with a history and a way back. The useful accumulation is evidence about which model performs a defined task well, under which conditions and at what cost.

All Medium posts →