Papers · 10 explained

The papers, explained

Each paper has a page of its own here that says in plain English what problem it takes on, what it does, what it shows and what it does not claim, with the figures the paper itself prints. The PDF is one click away on every page. The explainer is a way in, not a substitute.

—

Generating workflows

Five papers on Opus, written at AppliedAI. They run from a model that proposes a workflow, through the two ways of saying what a workflow is for, to scoring the result and finally to proving, as it is built, that it can still be finished.

Why a general language model cannot write the workflow for a specialised business process, and what Opus adds so that it can: a graph of how the work is actually done, and a model trained on it.

On hospital medical coding the two Opus models scored more than twice the general-purpose models on most measures, and beat the best of them, Claude 3.5 Sonnet, by 38 and 29 per cent on average.

The explainer →PDF ↗ (opens in a new tab)

Ask a language model for several workflows in one request and it tries to build one that does everything. Separating the request into intentions first fixes that, and the fix matters more as the request grows.

On 1,000 requests carrying from one to ten objectives each, generation without intention capture collapsed towards zero similarity as the objectives multiplied; with it, quality held.

The explainer →PDF ↗ (opens in a new tab)

When a workflow is built one addition at a time, every addition can pass every local check and still remove the last way to finish it. The paper makes the admission rule exact and carries a proof of it alongside the work.

A certificate travels with the unfinished workflow, and each addition must come with a checked update showing that an acceptable completion still exists. The core results are machine-checked in Lean 4.

The explainer →PDF ↗ (opens in a new tab)

—

Control under delegated authority

Two papers in supervisory control theory on what a controller can and cannot be made to do when it acts through a fixed set of actuators, or when the authority to refuse an action depends on what has already happened.

A controller limited to a fixed catalogue of modes can usually keep a system safe by cutting off risky futures. This paper asks when it can keep exactly the safe ones, and how few modes that takes.

Safety and exactness come apart: one mode can be enough for safety while exactness needs one per state. The application is permission profiles for AI agents.

The explainer →PDF ↗ (opens in a new tab)

In delegated control, who may veto an action can change as the system runs, and a supervisor may not know whether it currently holds the right. The paper gives the exact condition under which correct control is still possible, and computes the best behaviour achievable when it is not.

Knowing an action must be blocked is not enough, and neither is knowing you are allowed to block it: one supervisor has to know both at once.

The explainer →PDF ↗ (opens in a new tab)

—

Life support

Two papers on keeping a breathable atmosphere where nobody can step outside for air, written at Galactic Bioware.

An AI-controlled, semi-closed breathing loop for a firefighting suit, modelled from first principles, with every command to the hardware passing through a safety filter.

Recycling breath could roughly triple a firefighter’s working time, but topping up with pure oxygen slowly drives the suit towards a fire hazard. In simulation the controller gave 18.7 to 33.7 per cent more endurance than conventional control.

The explainer →PDF ↗ (opens in a new tab)

—

Information security

A study of manipulative narratives in the war in Ukraine, written at the Kyiv Aviation Institute.