How Strongly Should Task State Influence an LLM Agent?
Chenyu Zhang, Wonbin Kweon, Jiawei Han
为什么值得读
The study shows that merely showing an agent an accurate record is not equivalent to making that record control behavior. Stronger coupling can prevent repeated, premature, or cancelled actions, but enforcement is not universally safer: it can block correct work when its compiled state or request-to-step judgment is wrong. The evidence comes mainly from scripted single-user tasks in two synthetic domains, plus limited airline and prospective-memory case studies, so it does not establish behavior for broader workflows, unreliable tools, or multi-agent settings.
核心研究结论
- Returning rejection notices within the same turn eliminated visible false-completion replies without changing strict episode success relative to next-turn notices.
- Retrieval used as the task record failed almost every episode through omissions; verbatim top-k recall added to the full transcript helped but stayed below directives, Mem0 augmentation did not help, and an unverified agent-written ledger outperformed the always-accurate displayed checklist while still leaking re-executions.
- Adding an explicit one-shot rule primarily improved the raw text rung, especially with Qwen3-235B thinking; schedule controls showed smaller gains from the revised probe schedule, while stateful rungs moved little and the measured compile tax on enforcement was at most two episodes per cell.