Pith. sign in

Paper Citation Record · LEDGER

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach

As of 15 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2507.21397.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21397 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:56:58.146019Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact6
  • verified fuzzy24
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aa43ecb4-40c0-46fe-9e3d-6329c4a2057b · outbound

This paper cites A practical guide to multi-objective reinforcement learning and planning,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach A practical guide to multi-objective reinforcement learning and planning,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.585229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:56:58.046061Z digest=sha256:d33ad3ffccf59503332e931a6bc2cab2b103c8fe46726bd6a1110b4354a0daf9

Observation 2335be4c-0557-482b-9424-18431d3b1e57 · outbound

This paper cites Two-stage constrained actor-critic for short video recommendation,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Two-stage constrained actor-critic for short video recommendation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.574748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:56:58.049639Z digest=sha256:1542d1846a25fdf65f5f052540c7235092eca9f7fad7f087e4b320c35cf1c11c

Observation c54ef003-d5c1-40be-94fd-620c0c7df2b1 · outbound

This paper cites an unresolved cited work.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T12:56:58.052710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:56:58.052710Z digest=sha256:4a4ad1254dce4bc4ba3c4082d66f265776ccdd43e2f0722139e4528147c45a1d

Observation a0ba49b2-07a2-4011-a9f2-01ce76611306 · outbound

This paper cites Finite-Time Convergence and Sample Complexity of Actor-Critic Multi-Objective Reinforcement Learning.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Finite-Time Convergence and Sample Complexity of Actor-Critic Multi-Objective Reinforcement Learning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:56:58.333724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:56:58.055956Z digest=sha256:2dad79e3e0b0b8110ed619da75ce7dfae0dc3072ffee95436bfbed3e1675cdcb

Observation 8960fc44-6ba0-4421-8ed1-de03e4fee705 · outbound

This paper cites Recent theoretical advances in non-convex optimization,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Recent theoretical advances in non-convex optimization,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.557239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:56:58.059423Z digest=sha256:057a4cf8165d32a62c5c3c83c03d62d030b2eb77d24b5abe4028d931c86c509c

Observation 5dde7745-e0af-4e66-85ac-9f1b645295c2 · outbound

This paper cites Federated multi- objective learning,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Federated multi- objective learning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.548089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:56:58.062732Z digest=sha256:b8002842e27c0f2753af5ed97b727fc9b071a0c9f50f8f68edbcc9a13c31b572

Observation e20c2744-26a7-4b51-a809-cc4b285a74e6 · outbound

This paper cites Adaptive weighted sum method for multiobjective optimization: a new method for pareto front generation,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Adaptive weighted sum method for multiobjective optimization: a new method for pareto front generation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.538401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:56:58.066212Z digest=sha256:8906208bddd48fab690dac131a73b9ca70096b990a8988c6241668b1be92b3d0

Observation 77e9520b-053d-4597-9606-ee38f4a52d3a · outbound

This paper cites Multi-Objective LQR with Linear Scalarization.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Multi-Objective LQR with Linear Scalarization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T12:56:58.069078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:56:58.069078Z digest=sha256:af62625016a34b1b611fef0bd6cbf4233358a8bb3466a02c12ee5066aa8914cb

Observation d4c1cdcf-84a8-444d-851a-890cb9400575 · outbound

This paper cites Random hypervolume scalarizations for provable multi-objective black box optimization,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Random hypervolume scalarizations for provable multi-objective black box optimization,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.528494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:56:58.072291Z digest=sha256:b93ae24b9dff2ab42f59c321c4470d8d6bec1112dd7551e45e764a3aa14c92f7

Observation d9031b6b-fd45-4be1-902a-25375671bd78 · outbound

This paper cites Multiple-gradient descent algorithm (mgda) for multi- objective optimization,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Multiple-gradient descent algorithm (mgda) for multi- objective optimization,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.519273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:56:58.075044Z digest=sha256:f22914c03c4f21ed483766f936de9a6ca3f0d2999c68bce13b6ec59eb691c406

Observation 85f6ac14-0a1f-471b-9c63-cb4199f74da3 · outbound

This paper cites Miettinen, Nonlinear multiobjective optimization.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Miettinen, Nonlinear multiobjective optimization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.509473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:56:58.078152Z digest=sha256:d25f5749114b9aad1fa0ebb13aaade3ab149b2e91bdbfe85ffe1221283557cf2

Observation 1e8e0052-7dc9-4624-bca4-c37aae547588 · outbound

This paper cites Complexity of gradient descent for multiobjective optimization,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Complexity of gradient descent for multiobjective optimization,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.499944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:56:58.081670Z digest=sha256:c9601898a4a88c78ca76f0b0a0aa869933e69bd7c4cc9103b48329ab4532104d

Observation e4f5d7f3-b0af-46d0-b4d9-7da84f84cfb1 · outbound

This paper cites The stochastic multi-gradient algorithm for multi-objective optimization and its application to supervised machine learning,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach The stochastic multi-gradient algorithm for multi-objective optimization and its application to supervised machine learning,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.490024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:56:58.084871Z digest=sha256:19ea1a9f2db39e12f00a9edde86d0e17c4616b17a5f8f5efecd6750d2fbee2a9

Observation c6ea9553-2fe2-459b-a7f3-998ef65bb746 · outbound

This paper cites Mitigating gradient bias in multi-objective learning: A provably convergent approach,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Mitigating gradient bias in multi-objective learning: A provably convergent approach,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.480273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:56:58.088050Z digest=sha256:8bd41924c25cbfa65d95c0773849a48b5e6cca0fc462018eb2c08c7b64d65464

Observation d2a9d917-b656-46a7-8c0b-7a2790349017 · outbound

This paper cites Direction-oriented Multi-objective Learning: Simple and Provable Stochastic Algorithms.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Direction-oriented Multi-objective Learning: Simple and Provable Stochastic Algorithms

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:56:58.307541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:56:58.091135Z digest=sha256:1761727c40ea7e40f1abcded448bcd015ce1ea03cef2eef7042c5b5528b6aa20

Observation dfaec3d1-35ca-4f9e-95cf-ab1a9993cf88 · outbound

This paper cites Multi-criteria reinforcement learning.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Multi-criteria reinforcement learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.470294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:56:58.094657Z digest=sha256:af9f64360a871fd6268c58bd3bc1114bffc9d3a26a767d6bd82f05cc05f8364b

Observation 3deef8f7-69ec-4d7c-9c80-26ccb39754aa · outbound

This paper cites Reinforcement recommenda- tion with user multi-aspect preference,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Reinforcement recommenda- tion with user multi-aspect preference,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.459747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:56:58.097465Z digest=sha256:fedd96f0dcc3e089a603f92d10beee92d588412a5ae00ee5d589b7b61de04f68

Observation dffab944-2702-4da5-8c0d-0a00c4363fd9 · outbound

This paper cites Few for many: Tchebycheff set scalarization for many-objective optimization,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Few for many: Tchebycheff set scalarization for many-objective optimization,

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-06T12:56:58.292300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:56:58.100320Z digest=sha256:b1e5e1d14885505b5daa41ad8e85c71f5055982eb1b569af4a4e553664223549

Observation c4c42fb9-8abf-4deb-8ea3-69f7e795ce73 · outbound

This paper cites A multi-objective/multi-task learning framework induced by pareto stationarity,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach A multi-objective/multi-task learning framework induced by pareto stationarity,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.449710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:56:58.103231Z digest=sha256:9493be5b23bd013f9ede58e42d8ffa824504bffe3a6d432fb5f53bd1b111235b

Observation 51ab15c2-aa1a-4d51-9507-df65c499ee88 · outbound

This paper cites Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T12:56:58.106357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:56:58.106357Z digest=sha256:cf8372ae1f69d3cfba3a478327c4430436aef14722e255e29a7df2fbb5d04d68

Observation f9472c23-5145-4005-a299-b0372f00b82d · outbound

This paper cites Pareto multi- task learning,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Pareto multi- task learning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.440002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:56:58.109784Z digest=sha256:56288521397d6710ede3f5dbddf67d273431a73aeb958ecab54963920c1595ce

Observation 7fc45cfa-8481-4d88-b110-8b46b2dfa19c · outbound

This paper cites Multi-objective reinforcement learning for the expected utility of the return,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Multi-objective reinforcement learning for the expected utility of the return,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.430265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:56:58.112710Z digest=sha256:c481cf86dbf8ac75bc78c7ecf7b1a95e847c1bbb8a3c572bf7ba657347667a58

Observation d335029c-a979-4165-993c-6a91a6e97367 · outbound

This paper cites On finite-time convergence of actor-critic algorithm,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach On finite-time convergence of actor-critic algorithm,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.420122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:56:58.115733Z digest=sha256:e8fb073e87efd408cbb71155dae9af23b19e66f8cb8d10c898af901c1a95e569

Observation 0f909eb7-c329-4fd6-ada5-b76e92e88cd3 · outbound

This paper cites Improving Sample Complexity Bounds for (Natural) Actor-Critic Algorithms.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Improving Sample Complexity Bounds for (Natural) Actor-Critic Algorithms

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:56:58.212083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:56:58.118651Z digest=sha256:ce1d709a5c0e9f14645a34b44da83bacba3daf9bddcd103d5fd1258fa7fd1c6e

Observation 1ac70b17-5ff8-4113-8eab-f5d8b5e15629 · outbound

This paper cites Finite-time convergence and sample complexity of multi-agent actor-critic reinforcement learning with aver- age reward,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Finite-time convergence and sample complexity of multi-agent actor-critic reinforcement learning with aver- age reward,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.409718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:56:58.121863Z digest=sha256:547788439a14796ac872c35593fc8372c1355c3e03bde114ecc6cdf406b8049d

Observation 2d01debc-671d-4c8b-adc3-7035f0a2cc65 · outbound

This paper cites Fully decen- tralized multi-agent reinforcement learning with networked agents,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Fully decen- tralized multi-agent reinforcement learning with networked agents,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.399169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:56:58.124692Z digest=sha256:8859ceac5f2ca27369cad47c1530d7967baf28e124b359f0b7dbf5500636726d

Observation 3ec279f6-9ff0-4d5c-bf80-c36f6d37e54f · outbound

This paper cites Random Hypervolume Scalarizations for Provable Multi-Objective Black Box Optimization.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Random Hypervolume Scalarizations for Provable Multi-Objective Black Box Optimization

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:56:58.197875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:56:58.127643Z digest=sha256:9e3fe94dc76299c4eceefa7160b3adf0b4cfca99a0ccecf318c079c0665ede83

Observation bdca3921-7f6b-48be-b0b5-867a00e93398 · outbound

This paper cites Average cost temporal-difference learning,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Average cost temporal-difference learning,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.387720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:56:58.131124Z digest=sha256:4f7d3f368267e4c30dbc012aa9e24d786bf49cc3767d57636adbd3a5972fd775

Observation df7b27c1-a7c3-459e-9dd3-593f8d39e4d8 · outbound

This paper cites On the convergence of stochastic multi-objective gradient manipulation and beyond,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach On the convergence of stochastic multi-objective gradient manipulation and beyond,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.376462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:56:58.133887Z digest=sha256:7b665d363a1185d08b196a64c9bc0c70c3d47642aafaa3ff38417536ea7953ab

Observation 02f480a6-5ef9-436a-9739-53e2dc8d19b5 · outbound

This paper cites Multi-task learning as multi-objective optimization,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Multi-task learning as multi-objective optimization,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.366148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:56:58.136606Z digest=sha256:7f801227cc2c04d4386633acda34512b3a2beeb682f9ac689eee5d10db0e32a8

Observation 41561cf4-884c-435d-8fe8-edd5dfbe512f · outbound

This paper cites Theoretical Guarantees of Fictitious Discount Algorithms for Episodic Reinforcement Learning and Global Convergence of Policy Gradient Methods.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Theoretical Guarantees of Fictitious Discount Algorithms for Episodic Reinforcement Learning and Global Convergence of Policy Gradient Methods

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:56:58.182805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:56:58.139665Z digest=sha256:d24bdfd6a45c57c02c2f64d45e588ecfffb7d42c8e4613c8c4cab372467b5a04

Observation 95be232f-11ec-4b4a-957f-f401898d6fe8 · outbound

This paper cites Direction-oriented multi-objective learning: Simple and provable stochastic algorithms,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Direction-oriented multi-objective learning: Simple and provable stochastic algorithms,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.355531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:56:58.142832Z digest=sha256:5ea6524ac71ea11f16725163c3cf7ca2bcd43bdb5c5cb55e12c8a53982cd6607

Observation 60d6eb6f-060b-41fe-9663-5e76d6358fff · outbound

This paper cites Reinforcement learning to optimize long-term user engagement in recommender systems,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Reinforcement learning to optimize long-term user engagement in recommender systems,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.344937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-06T12:56:58.146019Z digest=sha256:6a77976258ef56c04aa1623f02992f71b9c483d0ad563200c780b521391844fe

Pith citing papers

No inbound Pith citation observations are available.