Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning with Segment Feedback

As of 19 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 2 inbound Pith citation observations for arXiv:2502.01876.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.01876 v2

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T14:19:51.917733Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T18:21:33.975739Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T15:02:45.745385Z

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a9e2aed3-f7cb-4920-b02b-d9226a645ef8 · outbound

This paper cites write newline.

Reinforcement Learning with Segment Feedback write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T14:19:51.837280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:19:51.837280Z digest=sha256:2100b7120413d2770d14d48951b854cb3dc454a5729ef573190c79b9fe15b222

Observation eeb90841-1345-4e41-b206-70885f44a7a9 · outbound

This paper cites Improved algorithms for linear stochastic bandits.

Reinforcement Learning with Segment Feedback Improved algorithms for linear stochastic bandits

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.176914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.842071Z digest=sha256:824d05043c905d81ccbeb9b4c1ff7d42ae294e200f0f48fd32ffab0d10d66650

Observation 48bb65b3-5523-46c3-8f38-e793d7315b3d · outbound

This paper cites Near-optimal discrete optimization for experimental design: A regret minimization approach.

Reinforcement Learning with Segment Feedback Near-optimal discrete optimization for experimental design: A regret minimization approach

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.166256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.845372Z digest=sha256:3ee5f141fa2d33d686a1e7b737ba8ea57d60c71213b33a8f71e3bed154516562

Observation 797df648-711a-4078-b2ba-0d730ef58bd0 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Segment Feedback Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:19:52.155293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.848793Z digest=sha256:44cc8abc471d65d051c1048d084f96fe23d1829edb9ea841f1700e37c19cbb79

Observation 047d5f60-4c25-45bc-98ea-aad9dcd95919 · outbound

This paper cites G., Osband, I., and Munos, R.

Reinforcement Learning with Segment Feedback G., Osband, I., and Munos, R

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.142619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.852301Z digest=sha256:5f55d15cb2725738598c9c32f55b04cad5f5fcda58b246df9e21a3a37ae51f3f

Observation f64a0e19-359e-4ce6-949e-26d36f838ad3 · outbound

This paper cites and Sundberg, C.-E.

Reinforcement Learning with Segment Feedback and Sundberg, C.-E

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.131058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.855536Z digest=sha256:b4d3a2fdd592c2966380164b5db05d69520bad5a8a744d7c028132899f587891

Observation 9708ab51-edfd-4bf4-b011-f37a9bdfb6ae · outbound

This paper cites On the theory of reinforcement learning with once-per-episode feedback.

Reinforcement Learning with Segment Feedback On the theory of reinforcement learning with once-per-episode feedback

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.119563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.859125Z digest=sha256:6da397ffc9589dd28ab68a9e4e3b40f4381d92f9b2c297d90949f94560d3f4b1

Observation 4254ac76-f449-49fa-b2f2-080697f08db9 · outbound

This paper cites Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning.

Reinforcement Learning with Segment Feedback Unifying pac and regret: Uniform pac bounds for episodic reinforcement learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.107444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.862300Z digest=sha256:89bbf4dbd446dd035ac977d4a7da8afef40cce34e1263557bad18d0220be6c70

Observation e906830f-dac6-4f56-9ed3-365145e617dd · outbound

This paper cites Reinforcement learning with trajectory feedback.

Reinforcement Learning with Segment Feedback Reinforcement learning with trajectory feedback

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.096493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.864992Z digest=sha256:75e211d9b4e7ef3a56b1750a67f0c09e5e0d6ea9a750b5cb7dd5603a8f71e08e

Observation fb8f7803-616b-4706-98c9-9f1248569426 · outbound

This paper cites Improved optimistic algorithms for logistic bandits.

Reinforcement Learning with Segment Feedback Improved optimistic algorithms for logistic bandits

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T14:19:51.868220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:19:51.868220Z digest=sha256:719d082f47e379c6b9b06ff4fed33c990f5f190b7a059e24f5d15da97c8cd4a9

Observation d7d7eccb-65ee-42ee-bcd1-d734a6f62da8 · outbound

This paper cites Parametric bandits: The generalized linear case.

Reinforcement Learning with Segment Feedback Parametric bandits: The generalized linear case

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.080447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.871243Z digest=sha256:29664e7c825eb6019389ae9287d85f523ddde36f6f1f182e21f5169086897980

Observation a9985149-adc6-43e9-91ed-5dadc67c580c · outbound

This paper cites Harnessing causality in reinforcement learning with bagged decision times.

Reinforcement Learning with Segment Feedback Harnessing causality in reinforcement learning with bagged decision times

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.070439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.874352Z digest=sha256:94fa0525311dbd170675bbdc985ac500ac96581107f71e2062477339c2585202

Observation c2cb0f95-f8d3-4aef-9dcf-5b58cf0598e5 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Segment Feedback Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:19:52.058986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.877211Z digest=sha256:3af4434591e79742e820a3f1bd2a1155e759b36df42bdeea270f5755b32ff576

Observation d422c481-9aa1-411d-8d41-79dc6dd9e9cd · outbound

This paper cites Near-optimal regret bounds for reinforcement learning.

Reinforcement Learning with Segment Feedback Near-optimal regret bounds for reinforcement learning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.048612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.879915Z digest=sha256:948d2e0650056fe7ed9b48df31edb546349df0200a7bfd20f3a798acd74ee3d1

Observation 5b0bfe69-5799-4b8a-8fd0-7f183172296f · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Segment Feedback Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:19:52.038642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.882770Z digest=sha256:3e2e17bbea244b310908ae6f0c825bf17b1857942de24a2852e0d7c4ff3fd431

Observation b5dad29f-7bf6-47e8-92c7-26568c55f415 · outbound

This paper cites and Hutter, M.

Reinforcement Learning with Segment Feedback and Hutter, M

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.028471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.886343Z digest=sha256:276e2187679e8e3c37589561f090776aa3111cbb21469a50f8b3d325d2f0392e

Observation 9be24f8c-4c6a-45fe-b684-70e86e12cc16 · outbound

This paper cites and Massart, P.

Reinforcement Learning with Segment Feedback and Massart, P

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.019473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.889347Z digest=sha256:46f23d708ac9812250b8d38ae9281437aa538522040d3225b63015d1673fccad

Observation 22347b28-8b4a-44a4-84d6-ac0ff433e6ce · outbound

This paper cites D., Jonsson, A., Kaufmann, E., Leurent, E., and Valko, M.

Reinforcement Learning with Segment Feedback D., Jonsson, A., Kaufmann, E., Leurent, E., and Valko, M

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T14:19:51.892457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:19:51.892457Z digest=sha256:33d20da9b415d281fea99a63cd6d45ae6ed6a54beaa351f7a1d1960c47ff4c02

Observation efd83852-8ee7-485f-a728-2325face270d · outbound

This paper cites and Moore, A.

Reinforcement Learning with Segment Feedback and Moore, A

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:52.003308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.895685Z digest=sha256:29022bec20cabc51c60581a81741395129b59af4a01dd2b1aeed645fb55192e3

Observation 05dfcaf2-1f5b-4aa8-8b19-d3ae9b83aabb · outbound

This paper cites Optimal design of experiments.

Reinforcement Learning with Segment Feedback Optimal design of experiments

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:51.995256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.898605Z digest=sha256:2e770adcaaef1e37ac7018acc7ef79bb50d40a64ccd67a151b4a5b1e43615fed

Observation f4451481-f3ee-475b-82b5-6f603bcb7e04 · outbound

This paper cites Self-concordant analysis of generalized linear bandits with forgetting.

Reinforcement Learning with Segment Feedback Self-concordant analysis of generalized linear bandits with forgetting

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:51.986161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.901392Z digest=sha256:3be42f4d72735c927cf1bccfa689d7566c5fb137cb2c067a568945d5dc7fb74d

Observation 19dcb54f-00ab-46af-96f0-463ae0cb0835 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Segment Feedback Unresolved cited work

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T14:19:51.904455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:19:51.904455Z digest=sha256:7fd2625be7075fe034d1208b276ff38f492fdc288497946baabadcda06dba67a

Observation ed638af5-d215-4e19-996b-97ec7b8aa83e · outbound

This paper cites Reinforcement learning from bagged reward.

Reinforcement Learning with Segment Feedback Reinforcement learning from bagged reward

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:51.970970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.908053Z digest=sha256:98d0b4deffa80ed0f256da88ee8eff443ba5626e656d826d1a9a42c873333210

Observation 84586d59-6cec-4806-a0e8-c061cd2a6dd2 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Segment Feedback Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T14:19:51.911511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:19:51.911511Z digest=sha256:7657a11623c236277cf3ab340051b9b085dd2112cb1717a2618ce023502a81d3

Observation 4cad5ac9-f212-46e7-903c-f6f62a6953da · outbound

This paper cites an unresolved cited work.

Reinforcement Learning with Segment Feedback Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-09T14:19:51.954438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.914853Z digest=sha256:5362dec5dadf6c9fd09c3edd13a402299bf676178acd3c1f24ab96df1d41d20f

Observation 3bc721f4-5633-47d3-8fc7-93a12b879514 · outbound

This paper cites and Brunskill, E.

Reinforcement Learning with Segment Feedback and Brunskill, E

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T14:19:51.944380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=arxiv_source observed=2026-08-09T14:19:51.917733Z digest=sha256:34ef1f5775a4611cb70acfe0e7215611cd3cbfb63d6d33cae60910060c63ae6d

Pith citing papers

Observation 2f6a9377-259a-45e1-9173-21492ddfe58f · inbound

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling cites this paper.

SP3O: Reinforcement Learning from Segment Preferences without Reward Modeling Reinforcement Learning with Segment Feedback

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-15T15:02:45.752291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T15:02:45.617996Z digest=sha256:164cc03ea2e365fcb82a3a752c7eeda747f3daac90e6426701145cdc9a870446

Observation e5a35949-7b26-4bf3-b165-64b8ebb71333 · inbound

Temporal GRPO: Beyond Trajectory-Level Credit in Vision-Language-Action Reinforcement Learning cites this paper.

Temporal GRPO: Beyond Trajectory-Level Credit in Vision-Language-Action Reinforcement Learning Reinforcement Learning with Segment Feedback

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T18:21:33.975739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T18:21:33.975739Z digest=sha256:cccf4edf4bed92053b3a21c40fc3d21fafbc254c1af0207a0427bc4932450994