Pith. sign in

Paper Citation Record · LEDGER

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach

As of 16 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 0 inbound Pith citation observations for arXiv:2507.21397.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.21397 v1

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T12:56:58.146019Z

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

33 of 33 outbound references displayed

  • verified exact6
  • verified fuzzy24
  • unresolved3
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aa43ecb4-40c0-46fe-9e3d-6329c4a2057b · outbound

This paper cites A practical guide to multi-objective reinforcement learning and planning,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach A practical guide to multi-objective reinforcement learning and planning,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.585229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:56:58.046061Z digest=sha256:c213a3adde5bae6f7d09d4612a7c5ecbaf8c08497c8b7ceaa4c9f73690a77755

Observation 2335be4c-0557-482b-9424-18431d3b1e57 · outbound

This paper cites Two-stage constrained actor-critic for short video recommendation,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Two-stage constrained actor-critic for short video recommendation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.574748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:56:58.049639Z digest=sha256:91aba1ca6ee764335f84632c877b160d645e258011c82605b2ab61931449a1dd

Observation c54ef003-d5c1-40be-94fd-620c0c7df2b1 · outbound

This paper cites an unresolved cited work.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T12:56:58.052710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:56:58.052710Z digest=sha256:4a4ad1254dce4bc4ba3c4082d66f265776ccdd43e2f0722139e4528147c45a1d

Observation a0ba49b2-07a2-4011-a9f2-01ce76611306 · outbound

This paper cites Finite-Time Convergence and Sample Complexity of Actor-Critic Multi-Objective Reinforcement Learning.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Finite-Time Convergence and Sample Complexity of Actor-Critic Multi-Objective Reinforcement Learning

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:56:58.333724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:56:58.055956Z digest=sha256:f1c6dc130f821c000841bcbf9e654d962c10658f0a24925670d2d72d068deeee

Observation 8960fc44-6ba0-4421-8ed1-de03e4fee705 · outbound

This paper cites Recent theoretical advances in non-convex optimization,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Recent theoretical advances in non-convex optimization,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.557239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:56:58.059423Z digest=sha256:9d3d73d57e4d29da8a98013e6348a47d88e2b41a59cf2a14968ff91fba429f46

Observation 5dde7745-e0af-4e66-85ac-9f1b645295c2 · outbound

This paper cites Federated multi- objective learning,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Federated multi- objective learning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.548089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:56:58.062732Z digest=sha256:0d2b46b67eac4d782e070271123d2e02642f352dca12a10b8116c2518c21ce83

Observation e20c2744-26a7-4b51-a809-cc4b285a74e6 · outbound

This paper cites Adaptive weighted sum method for multiobjective optimization: a new method for pareto front generation,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Adaptive weighted sum method for multiobjective optimization: a new method for pareto front generation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.538401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:56:58.066212Z digest=sha256:1a1b0113d2dd6510cdef749bc50f221df9d48e59b3dfc52f776b2ba89da7ad76

Observation 77e9520b-053d-4597-9606-ee38f4a52d3a · outbound

This paper cites Multi-Objective LQR with Linear Scalarization.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Multi-Objective LQR with Linear Scalarization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T12:56:58.069078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:56:58.069078Z digest=sha256:af62625016a34b1b611fef0bd6cbf4233358a8bb3466a02c12ee5066aa8914cb

Observation d4c1cdcf-84a8-444d-851a-890cb9400575 · outbound

This paper cites Random hypervolume scalarizations for provable multi-objective black box optimization,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Random hypervolume scalarizations for provable multi-objective black box optimization,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.528494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:56:58.072291Z digest=sha256:5c70d1224e2eedfdaca0727977f655e5598cc1b87da1bc00be0346c9bea11576

Observation d9031b6b-fd45-4be1-902a-25375671bd78 · outbound

This paper cites Multiple-gradient descent algorithm (mgda) for multi- objective optimization,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Multiple-gradient descent algorithm (mgda) for multi- objective optimization,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.519273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:56:58.075044Z digest=sha256:faf1c9b829b73c5f37ac44caeb89d49a366d1218eaf81bcece89a5e5f0b06fca

Observation 85f6ac14-0a1f-471b-9c63-cb4199f74da3 · outbound

This paper cites Miettinen, Nonlinear multiobjective optimization.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Miettinen, Nonlinear multiobjective optimization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.509473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:56:58.078152Z digest=sha256:e38f8411be387256b56877d1f85283963b95c4d3412526f7741e500afcdcc93d

Observation 1e8e0052-7dc9-4624-bca4-c37aae547588 · outbound

This paper cites Complexity of gradient descent for multiobjective optimization,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Complexity of gradient descent for multiobjective optimization,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.499944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:56:58.081670Z digest=sha256:267559afddab787a4d0814bfe2ee520b20f808841ae87743a365dc5292767d6c

Observation e4f5d7f3-b0af-46d0-b4d9-7da84f84cfb1 · outbound

This paper cites The stochastic multi-gradient algorithm for multi-objective optimization and its application to supervised machine learning,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach The stochastic multi-gradient algorithm for multi-objective optimization and its application to supervised machine learning,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.490024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:56:58.084871Z digest=sha256:9762b2093225d47b1070cb7dc7860c1ba34de747e1b0a38081893588886641fe

Observation c6ea9553-2fe2-459b-a7f3-998ef65bb746 · outbound

This paper cites Mitigating gradient bias in multi-objective learning: A provably convergent approach,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Mitigating gradient bias in multi-objective learning: A provably convergent approach,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.480273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:56:58.088050Z digest=sha256:b358fcf8dcc4aedf1615888a6e8ecff0be269d82635f32d83f4397c37383adf5

Observation d2a9d917-b656-46a7-8c0b-7a2790349017 · outbound

This paper cites Direction-oriented Multi-objective Learning: Simple and Provable Stochastic Algorithms.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Direction-oriented Multi-objective Learning: Simple and Provable Stochastic Algorithms

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:56:58.307541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:56:58.091135Z digest=sha256:25f28ad31ca28ff426262ddd801a53ef6671387d2f8baed10b99a642f3c0ecfb

Observation dfaec3d1-35ca-4f9e-95cf-ab1a9993cf88 · outbound

This paper cites Multi-criteria reinforcement learning.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Multi-criteria reinforcement learning

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.470294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:56:58.094657Z digest=sha256:21713d038e56476cbec6aea23ecf25e0aaf03c8336e2a9a48e580e0cdf12d336

Observation 3deef8f7-69ec-4d7c-9c80-26ccb39754aa · outbound

This paper cites Reinforcement recommenda- tion with user multi-aspect preference,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Reinforcement recommenda- tion with user multi-aspect preference,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.459747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:56:58.097465Z digest=sha256:2abc209247545ece5fd5b72a87de13364acfb9e48c179e957ee706079a1929db

Observation dffab944-2702-4da5-8c0d-0a00c4363fd9 · outbound

This paper cites Few for many: Tchebycheff set scalarization for many-objective optimization,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Few for many: Tchebycheff set scalarization for many-objective optimization,

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-06T12:56:58.292300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:56:58.100320Z digest=sha256:9c0214fa69f8a4b9cbfb02f842c1fb49eb1c4e4a3966771f502481fb318b3b4c

Observation c4c42fb9-8abf-4deb-8ea3-69f7e795ce73 · outbound

This paper cites A multi-objective/multi-task learning framework induced by pareto stationarity,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach A multi-objective/multi-task learning framework induced by pareto stationarity,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.449710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:56:58.103231Z digest=sha256:cde90356c60ef27dae37ca088eea90714f8d98ce64ac39856da380b03179effe

Observation 51ab15c2-aa1a-4d51-9507-df65c499ee88 · outbound

This paper cites Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Traversing Pareto Optimal Policies: Provably Efficient Multi-Objective Reinforcement Learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T12:56:58.106357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T12:56:58.106357Z digest=sha256:cf8372ae1f69d3cfba3a478327c4430436aef14722e255e29a7df2fbb5d04d68

Observation f9472c23-5145-4005-a299-b0372f00b82d · outbound

This paper cites Pareto multi- task learning,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Pareto multi- task learning,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.440002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:56:58.109784Z digest=sha256:d4e251dfc8ccdc7e3fb12d50f2e36ff1749c7472b6cfe68314debd70cb294365

Observation 7fc45cfa-8481-4d88-b110-8b46b2dfa19c · outbound

This paper cites Multi-objective reinforcement learning for the expected utility of the return,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Multi-objective reinforcement learning for the expected utility of the return,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.430265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:56:58.112710Z digest=sha256:9fadb993bf8ceb6a0dae4036fab79b647dd3eb25e60824e0fa5c904c029346c4

Observation d335029c-a979-4165-993c-6a91a6e97367 · outbound

This paper cites On finite-time convergence of actor-critic algorithm,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach On finite-time convergence of actor-critic algorithm,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.420122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:56:58.115733Z digest=sha256:8e9301b32d49206c3d9697b07b896982700a1fc73ec36d9428b57ec97a97d1b1

Observation 0f909eb7-c329-4fd6-ada5-b76e92e88cd3 · outbound

This paper cites Improving Sample Complexity Bounds for (Natural) Actor-Critic Algorithms.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Improving Sample Complexity Bounds for (Natural) Actor-Critic Algorithms

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:56:58.212083Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:56:58.118651Z digest=sha256:def12cf2db7544d473a600ce587ec9be7fa85aff8a27eba4d81fa5dc34222d45

Observation 1ac70b17-5ff8-4113-8eab-f5d8b5e15629 · outbound

This paper cites Finite-time convergence and sample complexity of multi-agent actor-critic reinforcement learning with aver- age reward,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Finite-time convergence and sample complexity of multi-agent actor-critic reinforcement learning with aver- age reward,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.409718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:56:58.121863Z digest=sha256:a30ed18c441177fb245987ebe1b7e6e7a18b887dd6a5894e290c64baab505094

Observation 2d01debc-671d-4c8b-adc3-7035f0a2cc65 · outbound

This paper cites Fully decen- tralized multi-agent reinforcement learning with networked agents,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Fully decen- tralized multi-agent reinforcement learning with networked agents,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.399169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:56:58.124692Z digest=sha256:f1b07df21c22122358b86b1e4a4d75536c6f5940f8c7d304b7138f0012153593

Observation 3ec279f6-9ff0-4d5c-bf80-c36f6d37e54f · outbound

This paper cites Random Hypervolume Scalarizations for Provable Multi-Objective Black Box Optimization.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Random Hypervolume Scalarizations for Provable Multi-Objective Black Box Optimization

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:56:58.197875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:56:58.127643Z digest=sha256:f375391df7b5caffeb9c8d92903aee8f5e1b43008ca5acf97cd0c64aaef6133f

Observation bdca3921-7f6b-48be-b0b5-867a00e93398 · outbound

This paper cites Average cost temporal-difference learning,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Average cost temporal-difference learning,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.387720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:56:58.131124Z digest=sha256:8ac670944a2f28e209722649da89fc4295023e141bbeb75b17ba98c8092c65f4

Observation df7b27c1-a7c3-459e-9dd3-593f8d39e4d8 · outbound

This paper cites On the convergence of stochastic multi-objective gradient manipulation and beyond,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach On the convergence of stochastic multi-objective gradient manipulation and beyond,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.376462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:56:58.133887Z digest=sha256:db0604b1f9c2b431929cedd73c0c30e8258ebaf8ed1e56fa36a32fd493cb9c16

Observation 02f480a6-5ef9-436a-9739-53e2dc8d19b5 · outbound

This paper cites Multi-task learning as multi-objective optimization,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Multi-task learning as multi-objective optimization,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.366148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:56:58.136606Z digest=sha256:bf8529fb896b07a809fa9aaa5a26f7e1f1a06da6e81e8431902448042b5cfa97

Observation 41561cf4-884c-435d-8fe8-edd5dfbe512f · outbound

This paper cites Theoretical Guarantees of Fictitious Discount Algorithms for Episodic Reinforcement Learning and Global Convergence of Policy Gradient Methods.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Theoretical Guarantees of Fictitious Discount Algorithms for Episodic Reinforcement Learning and Global Convergence of Policy Gradient Methods

Reference 31

Resolution
verified exact
local_arxiv, observed 2026-08-06T12:56:58.182805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:56:58.139665Z digest=sha256:2cef17f93de52c366c895f7f129c9909fc052224879f61ac017dda6502fa06dd

Observation 95be232f-11ec-4b4a-957f-f401898d6fe8 · outbound

This paper cites Direction-oriented multi-objective learning: Simple and provable stochastic algorithms,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Direction-oriented multi-objective learning: Simple and provable stochastic algorithms,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.355531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:56:58.142832Z digest=sha256:373d2216c2fc57e52818bfa38134b079705265265ab5bffbe5ecce7a125844d3

Observation 60d6eb6f-060b-41fe-9663-5e76d6358fff · outbound

This paper cites Reinforcement learning to optimize long-term user engagement in recommender systems,.

Enabling Pareto-Stationarity Exploration in Multi-Objective Reinforcement Learning: A Multi-Objective Weighted-Chebyshev Actor-Critic Approach Reinforcement learning to optimize long-term user engagement in recommender systems,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T12:56:58.344937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T12:56:58.146019Z digest=sha256:0af1b1cecb430133d81d88b56b78c20d97a781402a2fc1e1c3b717233df30c3a

Pith citing papers

No inbound Pith citation observations are available.