Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T07:16:24.447043Z
Paper Citation Record · LEDGER
As of 9 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 1 inbound Pith citation observation for arXiv:2601.20753.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-03T07:16:24.447043Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-12T05:06:14.783929Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-12T05:41:23.798033Z
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation a876add4-94d9-46f1-b677-9c80b268a4b3 · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Deep re- inforcement learning algorithm based on graph weight multi-pointer network for solving multiobjective travel- ing salesman problem.IEEE Access, 12:179091–179103,
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 522a511b-3764-4a57-a93d-39de75f3e344 · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning [Guet al., 2022 ] Qinghua Gu, Qingsong Xu, and Xuexian Li
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4a2afd35-bffe-4935-9107-6d9481b131de · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Huband, P
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73c6d813-0606-429d-8c1e-e7688e240984 · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Performance comparison of nsga-ii and nsga-iii on various many-objective test prob- lems.2016 IEEE Congress on Evolutionary Computation (CEC), pages 3045–3052,
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf143e8f-187d-4d0c-afa1-3db3df6b4ebd · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Pareto set learning for neural multi-objective combinato- rial optimization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 55a1d266-76b3-411b-93a1-c6a19ca03870 · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Smooth tchebycheff scalarization for multi-objective optimization
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e7b7fb6-3b1e-4f03-9983-12d02184dd44 · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Profiling pareto front with multi-objective stein varia- tional gradient descent
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0e3d3234-a05d-400e-8b21-f8efa6085d6b · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Pareto set learning for multi-objective reinforcement learning
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3cd51d8e-f59b-40cd-8d7b-93947de9583a · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Role play: Learning adaptive role-specific strategies in multi-agent interactions.Know.-Based Syst., 324(C), January
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d10365f2-990e-44c0-a213-75df90e6c6ee · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Multi-agent reinforcement learning for creating intelligent agents in social networks- oriented role playing games.Entertainment Computing, 54:100941,
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2e707774-5b6f-48f4-89e0-82306fbde7e8 · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Training language models to follow instructions with human feed- back
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 427d4606-ac11-4f9b-8f9a-640d348a24de · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Stable-baselines3: Reliable reinforcement learning implementations.Journal of Machine Learning Research, 22(268):1–8,
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 370b677d-2f12-41ba-8676-2592901e7bd1 · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Constructing complex npc behavior via multi-objective neuroevolution
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a7ea8dd-7966-4677-9603-7bc40fcba3e2 · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Dynamic defender-attacker blotto game
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 106aa7f9-d9f1-480c-a3ad-0d7fe62f85fd · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Steuer and Eng-Ung Choo
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3521586c-ef20-4aba-b0c5-62bf284e6018 · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Empir- ical evaluation methods for multiobjective reinforcement learning algorithms.Mach
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed4d876f-02f2-4486-a223-b541713a0e99 · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Prediction-guided multi-objective reinforcement learning for continuous robot control
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17cfcb28-3879-42b3-8af3-55e6bb55f5f6 · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning A generalized algorithm for multi- objective reinforcement learning and policy adaptation
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98ce69e5-b8c1-4158-b86f-e22784fd22e5 · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Federated reinforcement learning for robot mo- tion planning with zero-shot generalization.Automatica, 166:111709,
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0bb0d2c-1229-4c20-939f-5c0b480e02bd · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Maximum entropy population-based training for zero-shot human-ai coordination.Proceedings of the AAAI Conference on Artificial Intelligence, 37(5):6145–6153, Jun
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation faded5de-da6b-4e1f-9af4-0d1626b8f1f8 · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Scaling pareto-efficient decision making via of- fline multi-objective rl
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a6654283-b7bc-4612-bdc2-65b07539f193 · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Gymnasium: A Standard Interface for Reinforcement Learning Environments
Reference 1983
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c7db0a1b-3f6e-40ed-ad97-67b635f50714 · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Unresolved cited work
Reference 1998
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54966e1c-fb4a-4840-8e2a-e3c3379d4d99 · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Multiple-gradient descent algorithm (mgda) for multiobjective optimization
Reference 2005
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f555d3b8-de8c-4ad7-aeeb-e94564fb5a7d · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning A novel pareto-optimal ranking method for comparing multi- objective optimization algorithms,
Reference 2006
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffcea085-5d3d-4de8-a9ea-fcecd52c7989 · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Evolving multi-modal behavior in npcs
Reference 2008
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83202b15-cb91-41e3-a64a-cd16d84ae533 · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Proximal Policy Optimization Algorithms
Reference 2009
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e91e370e-5cdf-4613-b214-8be72439ce59 · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Drugan, and Ann Now ´e
Reference 2011
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db1c65e2-b45d-4e74-be0b-1e808f4bee35 · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Alegre, Ann Now´e, Ana L
Reference 2012
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cd838726-386f-4764-a5d9-a9c39039bbad · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Graph attention networks
Reference 2013
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 355bea24-fb58-4113-bb88-8da7872e5f28 · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Unresolved cited work
Reference 2014
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7efd88dd-e4cf-4a39-b0df-6fd6254aced9 · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Unresolved cited work
Reference 2016
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fff135a3-cf34-430b-afae-58726948b4ca · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning A Survey of Deep Reinforcement Learning in Video Games
Reference 2017
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 86c61b52-c919-448a-a0f6-59157451e6d2 · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Rlhgnn: Reinforce- ment learning-driven heterogeneous graph neural network for next activity prediction in business processes,
Reference 2018
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 05be720a-d8bf-4321-9e13-327c70b77976 · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Grapheon rl: A graph neural network and reinforcement learning framework for constraint and data- aware workflow mapping and scheduling in heterogeneous hpc systems
Reference 2019
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 720b7142-af1f-4a50-adcd-bb78ae5c0d07 · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning A Review of the Deep Sea Treasure problem as a Multi-Objective Reinforcement Learning Benchmark
Reference 2020
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6f16302-d6a0-403e-93cb-cadd603a9be8 · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Unresolved cited work
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d24440a7-dc84-4103-b41d-6116f2f800de · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Heterogeneous Graph Transformer
Reference 2022
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a066fcf-9907-49f4-bf57-15dc4b2667af · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Blank and K
Reference 2023
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2541a930-ad9b-4bea-b14b-0111582469ea · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Pick your battles: Interaction graphs as population-level objec- tives for strategic diversity
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4fea5578-9688-4f6b-9a03-f225d8aa9c27 · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Unresolved cited work
Reference 2025
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b83132a-13bb-4797-9f2e-69291e465f50 · outbound
GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning Precise and dexterous robotic manipula- tion via human-in-the-loop reinforcement learning.Sci- ence Robotics, 10(105):eads5033,
Reference 2026
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 94b06b96-8c3d-4907-a7d9-3e646fccc147 · inbound
Controllability in preference-conditioned multi-objective reinforcement learning GraphAllocBench: A Flexible Benchmark for Preference-Conditioned Multi-Objective Policy Learning
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.