Generating workflows ·

Opus: A Quantitative Framework for Workflow Evaluation

A way to score a workflow on what it is for: how likely it is to succeed, what it costs, what a correct result is worth, and how well it is built.

Read the paper PDF ↗ (opens in a new tab)

The problem

There was no principled way to say that one generated workflow is better than another. Text metrics such as BLEU, ROUGE and BERTScore say how similar two descriptions are, not whether a workflow runs efficiently or reliably. The established process and software-quality metrics assume execution is deterministic, so they cannot express a step that sometimes fails or a cost that accumulates along the way.

What the paper does

Each task in the workflow graph is given a probability of success when everything before it succeeded, another for when something before it failed, a duration and a resource cost. From those the framework computes the workflow’s overall chance of success, its total cost, its running time along the longest path and its peak demand for resources. The “Opus Workflow Reward” is the expected value of the correct outputs minus the weighted cost.

Reward alone does not see how a workflow is built, so four “Normative Penalties” score each task between 0 and 1 for cohesion, coupling, observability and information hygiene, combined into a single penalty. Workflows are ranked by reward first, with the penalty as the tie-breaker.

What it shows

The worked case turns customer complaint emails into support tickets, with the cost of a mistake priced from customer lifetime value and the chance of losing the customer. Of three designs, the most thorough — separate extract, classify and review steps — succeeds slightly more often, 0.9262 against 0.9250, but costs more and takes 6.1 seconds against 3.8, so a compact design with one review earns the higher reward. Two compact designs tie on reward. The one that folds its review into the action scores worse on penalty, because nobody can then see what the review contributed, and it is ranked last.

The paper notes what the numbers imply. For a small customer the model favours speed over reliability; for an enterprise customer, where a lost account is worth far more, reliability would win.

What it does not claim

The framework assumes each task’s duration and cost are fixed, which the paper notes often fails, for example in OCR; it treats failures as independent; it charges for tasks that may never run; and it offers no guarantee that a best workflow exists or can be found. Rankings mean something only between workflows with the same inputs, outputs and objective, scored with consistent parameters, and the evidence offered is the single case study.

Where it sits

It builds on the three earlier Opus papers and gives the series a measure of quality that is about the workflow rather than its wording. It is written to plug into reinforcement learning, where the reward says how well a workflow performs and the penalty how much room it has to improve.

All 10 papers →