Pith. sign in

REVIEW 2 cited by

On Value Functions and the Agent-Environment Boundary

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 1905.13341 v3 pith:3XUZ4T7M submitted 2019-05-30 cs.LG cs.AIstat.ML

classification cs.LGcs.AIstat.ML
keywords boundaryagent-environmentassumptionsfunctionsguaranteeslearningparttheoretical
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

When function approximation is deployed in reinforcement learning (RL), the same problem may be formulated in different ways, often by treating a pre-processing step as a part of the environment or as part of the agent. As a consequence, fundamental concepts in RL, such as (optimal) value functions, are not uniquely defined as they depend on where we draw this agent-environment boundary, causing problems in theoretical analyses that provide optimality guarantees. We address this issue via a simple and novel boundary-invariant analysis of Fitted Q-Iteration, a representative RL algorithm, where the assumptions and the guarantees are invariant to the choice of boundary. We also discuss closely related issues on state resetting and Monte-Carlo Tree Search, deterministic vs stochastic systems, imitation learning, and the verifiability of theoretical assumptions from data.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Offline Learning for Combinatorial Multi-armed Bandits

    cs.LG 2025-01 conditional novelty 7.0 of 10

    A pessimistic lower-confidence-bound algorithm achieves suboptimality bounds for offline combinatorial multi-armed bandits with probabilistically triggered arms, under coverage conditions requiring observation of each...

  2. Agency Is Frame-Dependent

    cs.AI 2025-02 conditional novelty 4.0 of 10

    Agency is frame-dependent: whether a system has a boundary, goals, self-causation, and adaptivity depends on arbitrary commitments by the observer.

Pith tools