REVIEW 5 cited by
Controlling Large Language Model Agents with Entropic Activation Steering
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
The rise of large language models (LLMs) has prompted increasing interest in their use as in-context learning agents. At the core of agentic behavior is the capacity for exploration, or the ability to actively gather information about the environment. But how do LLM agents explore, and how can we control their exploratory behaviors? To answer these questions, we take a representation-level perspective, and introduce Entropic Activation Steering (EAST), an activation steering method for in-context LLM agents. Firstly, we demonstrate that EAST can effectively manipulate an LLM agent's exploration by directly affecting the high-level actions parsed from the outputs of the LLM, in contrast to token-level temperature sampling. Secondly, we reveal how applying this control modulates the uncertainty exhibited in the LLM's thoughts, guiding the agent towards more exploratory actions. Finally, we demonstrate that the steering vectors obtained by EAST generalize across task variants. In total, these results show that LLM agents explicitly encode uncertainty over their actions in their representation space. Our work paves the way for a new understanding of the functioning of LLM agents and to effective control of their decision-making behaviors.
Forward citations
Cited by 5 Pith papers
-
Training with (Swap) Regret Loss in a Single-Layer Self-Attention Model: A Case Study on the Probability Simplex
Training single-layer attention with squared regret loss has stationary points that implement smoothed fictitious play (external regret) and, via a new swap-regret loss, the Blum–Mansour no-swap-regret algorithm.
-
ASA: Backbone-Training-Free Representation Engineering for Tool-Calling Agents
A probe-gated, router-conditioned activation hook raises strict tool-call F1 from 0.18 to 0.50 on Qwen2.5-1.5B while cutting false-positive triggers from 0.15 to 0.05.
-
Let's Get You Hired: A Job Seeker's Perspective on Multi-Agent Recruitment Systems for Explaining Hiring Decisions
A multi-agent LLM chatbot for job seekers was perceived by 20 interviewed participants as more actionable, trustworthy, and fair than their recalled experiences with traditional hiring methods.
-
Beyond Prompt Engineering: Robust Behavior Control in LLMs via Steering Target Atoms
STA selects sparse autoencoder features by activation amplitude and frequency to build steering vectors that improve LLM safety control over prompt engineering and standard steering.
-
SafeSteer: Interpretable Safety Steering with Refusal-Evasion in LLMs
SafeSteer uses category-specific activation vectors to steer LLMs toward safe, on-topic, non-refusing responses at inference time.
Discussion (0). Continue with ORCID to comment.