The problem
Outsourced business work — medical coding is the paper’s example — follows detailed procedures that are mostly proprietary, closed or never written down. A general-purpose language model has not seen them, so it cannot reliably produce a workflow that matches how the work is actually done; in practice, the paper notes, even a partial match takes several rounds of conversation with a domain expert.
What the paper does
Opus starts from the “Workflow Intention”: what the client supplies, what they expect back, and any context about the process. That intention is matched against a “Work Knowledge Graph”, whose nodes are tasks and whose edges join tasks that have followed one another in real historical workflows, weighted by how often they did.
The matched parts of the graph are handed, with the intention, to a “Large Work Model”: a language model fine-tuned on work rather than on text in general. It proposes candidate workflows, each a directed acyclic graph of tasks, where a task is a sequence of instructions and an instruction can be a model call, a piece of code or a review by a human expert. An AI agent is one kind of instruction among several, not the unit everything else is built around. The candidates are merged into one “Workflow Graph”, and the cheapest route through it, by compute, time and model cost, is the workflow chosen.
What it shows
The test is evaluation and management coding for hospital outpatient visits, mapped with Healthpoint Abu Dhabi, against a reference workflow written by senior professional medical coders. Two Opus models, built on work models of 70 billion and 8 billion parameters, were compared with o1-preview, GPT-4o, Gemini 1.5 Pro and Flash and Claude 3.5 Sonnet over ten trials, on five measures of how closely the generated workflow matched the reference: which tasks it covered, and in what order.
Both Opus models beat the others by at least a factor of two on most measures. The large model covered 72 per cent of the reference tasks against 25 per cent for Claude 3.5 Sonnet, and averaged across the measures the large and small models outperformed Claude 3.5 Sonnet by 38 and 29 per cent. The small model covered slightly more tasks but put them in a much worse order, which the paper attributes to its smaller, less noisy graph.
What it does not claim
It covers generation and optimisation only. Running the workflow is out of scope, the cost model is not developed, and the architecture and training of the work model and the design of the graph are left out. The results are for one use case, with GPT-4o judging whether each task matches, and the small model’s advantage holds only when the knowledge the use case needs is inside its smaller graph.
Where it sits
This is where the Opus papers start, and the terms it introduces — intention, the Work Knowledge Graph, a workflow as a graph of tasks — are the ones the later papers make precise. Kingston is one of four authors, all at AppliedAI.