Pith. sign in

Paper Citation Record · LEDGER

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning

As of 10 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2607.29617.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.29617 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T03:39:00.790790Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a32e3473-0170-4fed-a666-65629cd8e3cc · outbound

This paper cites an unresolved cited work.

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T03:38:59.806437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:38:59.806437Z digest=sha256:c2d9e2c03387169104b58a1b54f1dd930d206839261bfa4563f576a49a4336ba

Observation e01af168-0c2f-4314-a5a7-fd521672d600 · outbound

This paper cites an unresolved cited work.

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T03:38:59.858491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:38:59.858491Z digest=sha256:b1df2276975e1875bb1acefcbb6104c950b3f1724327bc7b6a99bfbe49b09b9d

Observation ce16b7a1-9f83-4eed-8170-318139427623 · outbound

This paper cites 57Appendix table of contents Part II Additional Results We collect here additional results omitted from the main text.

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning 57Appendix table of contents Part II Additional Results We collect here additional results omitted from the main text

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T03:39:00.355093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:39:00.355093Z digest=sha256:b23cf07408f84167497ae58c54fb9128f4adc12c27cf4ec2e120be90f95ef6a5

Observation a047ee89-3ac2-4571-a270-4c4421cad7f2 · outbound

This paper cites It is strictly suboptimal, with J πb Mb = 1/2 and supπ J π Mb = 1.

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning It is strictly suboptimal, with J πb Mb = 1/2 and supπ J π Mb = 1

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T03:38:59.960597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:38:59.960597Z digest=sha256:50527ac5197897d3cd5903593410abebbb9e70671364c567ae53f536ab3e872a

Observation 4cb40963-218d-4bd4-afc8-8ac14b8260ca · outbound

This paper cites If {πb :b∈ Bn} ⊆Π, then, for everyε∈(0,1/2), log2 Nε(Π, dΠ)≥2 n −n−1.

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning If {πb :b∈ Bn} ⊆Π, then, for everyε∈(0,1/2), log2 Nε(Π, dΠ)≥2 n −n−1

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T03:39:00.090686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:39:00.090686Z digest=sha256:3421c7e5a813963570704e982cb9319bc45723d8830fb30f7cff924756e918a5

Observation 55af81c0-6ca3-4570-83d8-cd69e9d956f5 · outbound

This paper cites , QK ∈ Qand coefficients (wk)K k=1 defining the linear combination LC((Qk)K k=1) = PK k=1 wkQk.

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning , QK ∈ Qand coefficients (wk)K k=1 defining the linear combination LC((Qk)K k=1) = PK k=1 wkQk

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T03:39:00.157760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:39:00.157760Z digest=sha256:2f80b71ec641c07c898685e34e811e94900230245d32126bde476c2eb42cec86

Observation d26c9ae6-8355-4920-9dd5-fe5a0e3608a1 · outbound

This paper cites It follows that, with probability 1/2, the next state is x−.

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning It follows that, with probability 1/2, the next state is x−

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T03:39:00.254929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:39:00.254929Z digest=sha256:40f650c3d453145115bde583f27fd9da42de2db114670698d2770de2e39addb1

Observation 4af1816b-a5f5-4da2-a999-8b230c982408 · outbound

This paper cites The resulting algorithm is in Algorithm 3.

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning The resulting algorithm is in Algorithm 3

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T03:39:00.433896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:39:00.433896Z digest=sha256:3c593574aadf8733304695c7a0c2b50251c932f220858142f150c129d601bdbe

Observation 32e96416-ff2d-4fff-ab0e-f565bbaae6f4 · outbound

This paper cites an unresolved cited work.

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T03:39:00.596173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:39:00.596173Z digest=sha256:a16ec4e7b8e89fccafb4a1b7663ec21d146bd49479102b7a0d9c408c13bc6b3a

Observation da324c86-8952-4077-b147-cd90ff6b1c66 · outbound

This paper cites an unresolved cited work.

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T03:39:00.775798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:39:00.775798Z digest=sha256:30aa443aed4e3d98fbaaab9fcf5519bbc43e2f482bd916adb8600b5086a30c56

Observation c088e1c1-8d8e-4e26-84d1-3effefb409d1 · outbound

This paper cites HX h=1 X x∈X dπE h (x) D QπI h (x,·), πE,h(· |x)−πIh h (· |x) E# . Usingd πE h =d h + (1−α)(d πE h −d πout h ), the right-hand side is equal toT 1 +T 2, where we define T1 :=E I.

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning HX h=1 X x∈X dπE h (x) D QπI h (x,·), πE,h(· |x)−πIh h (· |x) E# . Usingd πE h =d h + (1−α)(d πE h −d πout h ), the right-hand side is equal toT 1 +T 2, where we define T1 :=E I

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T03:39:00.790790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:39:00.790790Z digest=sha256:49aa44017d54beabebc1c43887404f3d5633917e6755d486b7445eecdea0cb7e

Observation 4eee1d9d-00ca-4ed2-a928-982f3d09fb48 · outbound

This paper cites value-based IL.

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning value-based IL

Reference 1991

Resolution
unresolved
no resolver link, observed 2026-08-03T03:38:59.609784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:38:59.609784Z digest=sha256:5290df9b401a7616a50721a6c3636f0786952d6224d8d9b8e67751dd7907f584

Observation 7fbca4f4-5c12-46cc-a076-64de5c06e280 · outbound

This paper cites A Theory of Learning with Autoregressive Chain of Thought.

When Does On-Policy Interaction Help? Representational Tradeoffs in Value-Based Imitation Learning A Theory of Learning with Autoregressive Chain of Thought

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-03T03:38:59.481636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:38:59.481636Z digest=sha256:346a8643d6791b37f4721927eed93a2f4f254590970f1b2888c2da483ed90eef

Pith citing papers

No inbound Pith citation observations are available.