Pith. sign in

Paper Citation Record · LEDGER

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods

As of 20 August 2026, this Paper Citation Record lists 46 of 46 outbound references and 1 inbound Pith citation observation for arXiv:1908.03263.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
1908.03263 v1

Coverage vector

measured 46 of 46 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T14:25:39.636614Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-11T11:33:20.892688Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-11T11:33:21.562600Z

Reference resolution

46 of 46 outbound references displayed

  • verified exact2
  • verified fuzzy18
  • unresolved26
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 191a5e9e-b75f-46ec-943b-b200e16d508c · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.924948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.177695Z digest=sha256:57dde5feeac8df84c3192ad1ab06062de211aff71de1e8c438e59884efbaae7b

Observation 7da8213a-70aa-4ee0-829a-f224c1aae600 · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.897151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.187513Z digest=sha256:2e20006f74c70152580a9190d018d2449ac59e1899a3fe1dba4f70a80eed6fd6

Observation b3ca9d18-1357-4dfd-a6ee-27d6b2e53085 · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.869973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.196950Z digest=sha256:de6cc03edec15b0ca1460da83b33cb69eb17a8e20e87cebe41373f1178a1e536

Observation cc2a877b-8086-408d-9614-bd9979352332 · outbound

This paper cites Peters and S.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Peters and S

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.847301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.206498Z digest=sha256:c423cfd42d3ec5a68b459bc0342ff7a17815ebe321102c9ad15c6672b4d171b7

Observation 34cf5d5e-26cf-4e57-a13e-f01d62e27c77 · outbound

This paper cites Schulman, S.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Schulman, S

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T14:25:39.213978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:25:39.213978Z digest=sha256:6afbc5d5d0051572d47755316d95cd8036bdcbfb700529f21530e72391cecbca

Observation 57f71c96-e1ba-4f92-a4ec-cdf0b63a3ba2 · outbound

This paper cites Cheng, X.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Cheng, X

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.804769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.222566Z digest=sha256:81d3b278fc3ea4f28167198482136e43783c10fa971a66adcc94b394011521cb

Observation d31b7ebf-c3c0-4b7f-a6d5-0fa7aa35bdee · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.784745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.234559Z digest=sha256:22263ee752dd965a50b823dd536ba0447d8e8faba4857ee027f9376bbd65a225

Observation 71803ba2-1efc-48fb-bd61-c9da035f4473 · outbound

This paper cites Cheng, X.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Cheng, X

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.750048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.245548Z digest=sha256:0c66ddffa3d67b158a679d16d07c6894b1b6666bed21681e55227ca66b1b1904

Observation 09f63648-3278-48c7-8939-de72168efcb4 · outbound

This paper cites Policy Optimization with Stochastic Mirror Descent.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Policy Optimization with Stochastic Mirror Descent

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-14T14:25:39.867071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.260813Z digest=sha256:102c7c7cdccae10a1c08f9fce91cbcc45dd4b2b07cb7f3839146f2380bfa3f95

Observation d626b641-55e4-4d66-892f-f924b8a01997 · outbound

This paper cites Ghadimi, G.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Ghadimi, G

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.725517Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.274695Z digest=sha256:dcac0eea7409027b8e5d74e23d1fae201f96227d17b2cae050ef7af019faaf86

Observation a671989a-ef87-46c2-8c7c-a3c28e61d926 · outbound

This paper cites Kimura, S.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Kimura, S

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.702491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.282947Z digest=sha256:dc248b09afc98ac0cd8de14728fdf8fb3a68a94c2db86a5208ccfa81a2f42122

Observation 5c46c7f7-eb94-4c5f-bd5e-6990e74346a4 · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.676747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.291379Z digest=sha256:45ad904af7de8de9fe55918529f0b002e522bf29eb5abc734cfd51d47f6d435e

Observation 9aa4f302-9e25-4fe3-a30e-1ca7475dc66a · outbound

This paper cites Silver, G.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Silver, G

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.650996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.300911Z digest=sha256:06e321763cec8a07d81d2176b7e9724ddc5ba057246ef247500319151475b80e

Observation a7042fe8-6800-4ca3-9fbb-cd57f2745822 · outbound

This paper cites Schulman, P.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Schulman, P

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.616211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.306683Z digest=sha256:e53ef43e58b092898d6eb3d0524c4909793e6b8e94d58d878ab352cd34c24c18

Observation c6273e3d-8564-4baf-a028-7b905328105e · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.593088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.315864Z digest=sha256:673472c8e5eb2a9488a479a490ebd5ca46711b195360298ce99519bfe8d11665

Observation 80be2f38-460e-4332-b502-7de517abca12 · outbound

This paper cites Efroni, G.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Efroni, G

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.570376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.322420Z digest=sha256:47f031e8ec6474e80510016c42bd043f5c8ab627cb9b480a480a4842192175e2

Observation f8d14759-c79d-4cfd-910c-c571bb326be1 · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.551194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.328381Z digest=sha256:2841190583f4a939b6630c98898ae9d99b3fcc8fcca70c23250a3a680e88d577

Observation dd6fce16-93ba-4e1f-8598-cb6ec1a9da50 · outbound

This paper cites Greensmith, P.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Greensmith, P

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.531336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.337728Z digest=sha256:8cc3cba481b3419197091d3193d56fe6ee092481d912cb76aaa99391c3a87778

Observation e1079437-cff7-41ef-9154-0d8c8d1d818d · outbound

This paper cites Jie and P.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Jie and P

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.508191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.346992Z digest=sha256:7d78e4f5f65436c65011bfc266256495a7b1ed7ab068575ef76b4ad90adf945c

Observation 24bf1f2f-8d04-49fe-b23c-be04665ed17b · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.483707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.363164Z digest=sha256:4f711f029c50c7a10675e06d63ecaef213e0051962b4f9752c9594e6204b2978

Observation 7ee92dd1-df7f-4d7e-a123-7650d1098d6d · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.462850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.370748Z digest=sha256:95f4c96acaf3b14a77b8f470ea04acb860aebc71655361af4f0afecdcba4452c

Observation f864052a-b272-4a22-885e-ba3824d912c0 · outbound

This paper cites Grathwohl, D.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Grathwohl, D

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.441686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.386784Z digest=sha256:88477559fe82bf41cf7ae19e93508387784cddcde079065b0c0e47c8cabe8c00

Observation 8a087b20-2c97-45a4-8fed-e943814e64e2 · outbound

This paper cites The Mirage of Action-Dependent Baselines in Reinforcement Learning.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods The Mirage of Action-Dependent Baselines in Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T14:25:39.395346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:25:39.395346Z digest=sha256:6d9701c7610d9b9007137ad4f4758775c3d1526f69063170a67b8fee1d57a6df

Observation 371e0090-7036-4b0c-a91b-ec496ff7e2e3 · outbound

This paper cites Reward-estimation variance elimination in sequential decision processes.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Reward-estimation variance elimination in sequential decision processes

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T14:25:39.404839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:25:39.404839Z digest=sha256:5b7bd056aed0dba9202cf32f6213d9359a7d67d4b2ead2861918190eb26c83cd

Observation a2db03ec-246c-485a-a25c-17b62980cfc2 · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.417372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.418128Z digest=sha256:97ce6329b61d7ddd3c06e4ed2b02914d45763a333c1946468dc4bc76ce5b3b4f

Observation c8c16567-bff4-4ed4-a284-9f49c2b7c3b5 · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.397483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.425024Z digest=sha256:0ea9e8363fe901e6f73b718a963f0adbe17cc0a98a2af4b37af0d462403c8b5b

Observation 83f44b55-f1fe-4b9b-9aac-c46a2fb6e61b · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.365337Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.434584Z digest=sha256:4637cedb4b98ee315c55a40592490dae367f91962c0d414cdb2120bcbe22f5b3

Observation b4090b9c-e85c-46ae-a45d-ddb4d4f6ca6b · outbound

This paper cites Beck and M.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Beck and M

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T14:25:39.440612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:25:39.440612Z digest=sha256:f8eb67467e044c220fb741318cae4f9d0303f164da73035227b33fc9694c8228

Observation 07ba154f-3f7f-4ce6-831b-813832b46c4f · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.317137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.449838Z digest=sha256:53be65b48bdfaf2ecddea482603ad9f21943d2452464e03009289b08e92636f0

Observation 2c2a25d3-496d-4314-be2d-c416db887245 · outbound

This paper cites Vemula, W.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Vemula, W

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.284582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.464991Z digest=sha256:b18b824708f31e6a2115c2e88015aeda9101de897e3fb60463607df23c2a8b31

Observation 38408601-f664-4a24-bb12-efedbf9cadeb · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.258593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.482592Z digest=sha256:93af32581aa3453ab4d39e962b83d4c6c505e20cdbc58f7e847a8ea02777bd6c

Observation 895ec48f-b0a8-4842-8570-d3adb223217c · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-14T14:25:39.491821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:25:39.491821Z digest=sha256:e31ba202fa79a76dc0d18e46b8704acbd28624bba8e9f4c1d8fe6f1dd27fb816

Observation 90fad94f-fd03-4d05-a2a9-3772dacca61f · outbound

This paper cites Schmidt, N.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Schmidt, N

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-14T14:25:39.499332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:25:39.499332Z digest=sha256:49bd1634c8f056e7f1785cd3172f6c429ebb9701f47c3ec577795b16f704e0b1

Observation bceff0ea-6228-4960-823e-e87165ca94b5 · outbound

This paper cites Johnson and T.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Johnson and T

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.199064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.506390Z digest=sha256:ea046e16099a6917bf5bd2400c31437cbddac75b6dcd551b989f70473efd334d

Observation 3b37b103-d5da-4510-b3d3-50c58d4bc5b0 · outbound

This paper cites Defazio, F.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Defazio, F

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.168026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.514204Z digest=sha256:e295a4dd29ce0b94c89e68e2876f5590edc4691c7595fa7f186e9f2a90ebeeb2

Observation 1df63345-774d-44a6-8c82-8065e0835dd4 · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.132969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.520508Z digest=sha256:6514ff191708ceac9c1034f703801ecbbd9047bedd2b1e0170f1ed0d2765cce1

Observation 9627ad71-871a-4bb5-8b6f-3f4df6e5262a · outbound

This paper cites Expected Policy Gradients for Reinforcement Learning.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Expected Policy Gradients for Reinforcement Learning

Reference 37

Resolution
verified exact
local_arxiv, observed 2026-08-14T14:25:39.759551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.540027Z digest=sha256:c9319baa1392f5a939177d108e9618736633c10ade6ab2e8d8d7dbfe1253a5a7

Observation 14a9e677-16d0-4d01-9b91-5cb3b06f84c7 · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.105382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.547750Z digest=sha256:179fd1c6159538ba1fd698ffd7c6b84020e39980792477d304223330d6d12ab7

Observation 6716a93d-b318-4de0-9902-a434c6d3db5a · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:40.069970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.568347Z digest=sha256:b1c107e08e08e27649a6ac39daa59eacd28b406f98c67809bef7c4ae7faab857

Observation af4d472a-85e8-4970-b13a-003c7530376e · outbound

This paper cites Baxter and P.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Baxter and P

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.038358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.579352Z digest=sha256:da223596ed017c21b6ddf76b13c2710a03eda4ab7895ac1b21bebcc94719e5e1

Observation 0985f900-b8ab-4169-88f1-305275137236 · outbound

This paper cites Landau and E.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Landau and E

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:40.000894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.593945Z digest=sha256:608ada69acd65debed9ef5c8d73e19ee5e2dbfe6859f646de4e39ffd831b0c02

Observation 9e8624c0-204f-46eb-b5ab-034b6abb15d3 · outbound

This paper cites OpenAI Gym.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods OpenAI Gym

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-14T14:25:39.601235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T14:25:39.601235Z digest=sha256:5a975202c79f7d126b6c6f100f03703f3b9dc3cd69f7b25320a31ae7d81ff571

Observation 62138545-a281-4e10-87c7-5e8f3b0c6a9a · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:39.977301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.612620Z digest=sha256:2dd4f58fc8bfd7101e632e1a0fcb3dbe5d72e90c1e65c0b63bf4096efafc4371

Observation bba63ec3-e7a8-4190-8c9d-92853e5157e5 · outbound

This paper cites an unresolved cited work.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-14T14:25:39.946075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.619676Z digest=sha256:9d73de5cf30e38316ff1af4fdf534b2dfd588125c38cfd0ae6a718958a669321

Observation b27d937a-603c-4b9d-bc23-642f3b55db4b · outbound

This paper cites That is, a feasible ordering must be causal at least in actions: the action randomness that causes a state must be arranged before that state in the ordering.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods That is, a feasible ordering must be causal at least in actions: the action randomness that causes a state must be arranged before that state in the ordering

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:39.922455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.627640Z digest=sha256:c41fe3822e6ad0a29c6642cd7ed82024271f50da0cacb9fbb10a771082a64262

Observation 4d2ba498-63b1-4513-bc65-d36c41ce8afb · outbound

This paper cites We consider the following operations (a) Suppose, in an ordering, there isSv→Su,v >u, then we can exchange them without affecting residue.

Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods We consider the following operations (a) Suppose, in an ordering, there isSv→Su,v >u, then we can exchange them without affecting residue

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-14T14:25:39.894577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-14T14:25:39.636614Z digest=sha256:fb1f6011438ccfb1ded69e8e3e4db967c1b1020aed6f20657665786f3a253dd3

Pith citing papers

Observation 2ac05bc8-c0f4-4185-b19b-7e362f2cf794 · inbound

Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems cites this paper.

Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems Trajectory-wise Control Variates for Variance Reduction in Policy Gradient Methods

Reference 282

Resolution
verified exact
arxiv_id, observed 2026-05-11T11:33:21.565277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-05-11T11:33:20.892688Z digest=sha256:54df6ea0718655d9faa448dd43e9d05673a4220589725449b57b7b352547f9d0