LLM agents need to sustain goal-consistent reasoning across long multi-turn interactions under strict resource constraints. However, as the multi-turn context accumulates, it can destabilize the underlying LLM's internal representation of task-relevant information from earlier turns, blurring the boundary between constructive reasoning and representation drift. We formulate multi-turn reasoning as a hidden-state trajectory of the underlying LLM that is characterized via two complementary signals: temporal curvature that captures the directional consistency of turn-to-turn updates, and variance slope which measures the expansion or contraction of the exploration space. Across four tasks and three underlying LLMs, we observed that these geometric signals distinguish between correct and incorrect episodes prior to completion. We further decompose each episode into three-action chains formed from four actions (Read, Write, Respond, Transfer) and show that separability is action-dependent, with different signals distinguishing various chain patterns. Our experiments demonstrate that trajectory geometry can identify critical turns in the reasoning process, increasing task success rates on tau-Bench from 24.1% to 39.6% while reducing token cost by 11.2%.
Autonomous LLM agents frequently fail during multi-turn interactions because accumulating interaction history destabilizes their internal representations, causing context saturation, representation drift, and unneeded reasoning overhead. This paper proposes tracking the trajectory of intermediate hidden states across conversational turns using two geometric metrics: temporal curvature (measuring directional consistency of hidden-state displacements) and variance slope (measuring expansion or contraction of the exploration space). Instead of relying on static context engineering or always-on chain-of-thought deliberation, the authors demonstrate that monitoring these latent geometric signals can trigger deliberative reasoning dynamically only when failure or drift is imminent. For practitioners and researchers developing conversational agents in tool-grounded domains like retail and airline customer support, this method offers a potential route to improve task success rates while lowering overall inference token costs. A primary operational limitation is that extracting hidden states requires direct white-box access to the model's internal activations at specific intermediate layers, making the approach inaccessible for closed-source proprietary API-only models.
论文中的结论可在「研究结论」中查看。