Explore ArkGraph and select the steps to run.
This paper isolates how task state is coupled to an LLM agent while holding models, rules, and paired episodes fixed. It compares a raw transcript, an exact displayed checklist, state-derived per-turn directives, and an enforcement gate that rejects state-violating actions. Synthetic assigned-project episodes, a real-tool harness, and two external benchmarks test models across reasoning regimes and task sizes. Accurate displayed state remains unreliable; directives improve results according to model obedience; enforcement is robust when errors are state-decidable but inherits compiler and matcher mistakes. Reasoning narrows rung differences, and PM-Bench shows that enforcement can hurt when action depends on uncertain free-text cue recognition.
The study shows that merely showing an agent an accurate record is not equivalent to making that record control behavior. Stronger coupling can prevent repeated, premature, or cancelled actions, but enforcement is not universally safer: it can block correct work when its compiled state or request-to-step judgment is wrong. The evidence comes mainly from scripted single-user tasks in two synthetic domains, plus limited airline and prospective-memory case studies, so it does not establish behavior for broader workflows, unreliable tools, or multi-agent settings.
The paper’s claims are available in Research claims.