Generating workflows ·

Opus: A Prompt Intention Framework for Complex Workflow Generation

Ask a language model for several workflows in one request and it tries to build one that does everything. Separating the request into intentions first fixes that, and the fix matters more as the request grows.

Read the paper PDF ↗ (opens in a new tab)

The problem

People describe the work they want automated informally, and one request often bundles several goals. Handed straight to a language model, such a request produces a single workflow that tries to serve them all. The paper finds that this leads to overlapping contexts and lower accuracy, and that it gets markedly worse the more goals are packed in.

What the paper does

The framework puts an “Intention Capture” layer between the request and the generation. One model call extracts the input, process and output signals from the request; a second groups them into separate intentions; and one workflow is then generated per intention. Missing detail is flagged as an incomplete intention, to be asked about rather than guessed.

To test it the team built a synthetic benchmark of 1,000 requests, drawn from a pool of 1,000 services across 100 industries, with 100 at each level of mixing from one objective to ten. Generation with and without the layer was compared on nine models, using the same model for the reference and for both runs so that the only difference is the intention step.

What it shows

Across all the models and all the scores, the average difference between generating with intention capture and without it was positive, and the gap widens with complexity. For Claude 3.7 Sonnet, similarity to the reference was about the same either way for a single objective; at ten objectives it held at 0.89 with intention capture and fell to 0.08 without. A GPT-4o judge told the same story, with gains of up to 40 per cent on its total score.

What it does not claim

The paper says its text-similarity metrics measure surface overlap and are not enough on their own, that a single model acting as judge brings bias and limited reproducibility, and that output-length limits kept some models from being tested at the higher levels. The gains in grouping intentions level off, or reverse, for the most densely mixed requests.

Where it sits

It puts the Workflow Intention framework into practice with off-the-shelf models, and separates the value of the intention step from what a model already knows. Its admission that text overlap is not a sufficient measure of a workflow is the gap the next paper, on quantitative evaluation, takes up.

All 10 papers →