Pith. sign in

Paper Citation Record · LEDGER

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control

As of 9 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 0 inbound Pith citation observations for arXiv:2606.21925.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.21925 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-26T11:59:18.223000Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

65 of 65 outbound references displayed

  • verified exact5
  • verified fuzzy0
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch7

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 18eff5ca-7a2d-4808-834e-35318cf2d6cb · outbound

This paper cites PhiBE: A PDE-based Bellman equation for continuous time policy evaluation.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control PhiBE: A PDE-based Bellman equation for continuous time policy evaluation

Reference 1

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:19:43.735582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:dcd29f8ce52732c035f3ef5719de5761d3df2b85a086ab73bd29e87a9a37b243

Observation 04c12d0c-ce18-4373-bfe1-5e6fa0c88265 · outbound

This paper cites A deep learning-driven iterative scheme for high-dimensional HJB equations in portfolio selection with exogenous and endogenous costs.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control A deep learning-driven iterative scheme for high-dimensional HJB equations in portfolio selection with exogenous and endogenous costs

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:19:43.725204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:6ac89e3a084592fe3095f92382ea1fbb458b1a33834d4250165084a21b19b5c0

Observation 934246e1-cb9d-45eb-84a2-3eee0650766b · outbound

This paper cites IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics) , volume=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics) , volume=

Reference 3

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:0e8471f4e3085470784500e27e0d88b4eb4592a9020e2793462c799f5b9cd4f0

Observation 19e74262-7d46-4d61-9f3b-3f924dd8b227 · outbound

This paper cites Machine learning , volume=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Machine learning , volume=

Reference 4

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:e0a43cf8144edad43bf0b10b7c324f9d843ad74b06317db30fd58c2a9b9a6997

Observation 20cdcf70-8ac3-493a-9c1c-b0e2849869f6 · outbound

This paper cites Journal of Machine Learning Research , volume=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Journal of Machine Learning Research , volume=

Reference 5

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:d2d568f854caa224658b0639bf44a1edb9e93b0d798087d327818b449eb17aa3

Observation 4ab243e7-2d46-42bf-8a0d-d1c89d485692 · outbound

This paper cites Journal of Machine Learning Research , volume=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Journal of Machine Learning Research , volume=

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:bffd428d56ea60b5c652abc281f3e5b045b77048f62181b264f5ab64739a7909

Observation c1224526-dfcb-4d76-9ee3-2a8d49eec90c · outbound

This paper cites Journal of Machine Learning Research , volume=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Journal of Machine Learning Research , volume=

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:7949ea97d1a4460ea366f9bfe84a5da6ee8a111da784c01b26deba39efd37488

Observation 8fbe1b95-bc16-432d-8202-4af28da38db8 · outbound

This paper cites Journal of Machine Learning Research , volume=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Journal of Machine Learning Research , volume=

Reference 8

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:19eaa10cb7a27e363fcf9fd5e143e9f2de3c465b46698931989243d0d64633f2

Observation 89c20810-68b5-41fa-9374-78d75d8fdd3b · outbound

This paper cites nature , volume=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control nature , volume=

Reference 9

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:c1d4094d00f34301706034571b846dea9cc21e7672d51bd1391344c7b5a760c4

Observation b8778096-3981-4ca8-acda-f4cfc621209d · outbound

This paper cites nature , volume=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control nature , volume=

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:1ff91e456e32b23ac482c1cc8670df84577c5cda4ec7e30fb1efe80039df8b20

Observation 0bd96a63-0a59-4b12-a9b6-d48096270142 · outbound

This paper cites Nature , volume =.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Nature , volume =

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:650cc59b38c8d964554c15c0fb0786d59b1c7b04519e7d9d502b886b41bc0f67

Observation 4550f723-c310-4007-93d9-8312435f2de7 · outbound

This paper cites Proceedings of Thirty Third Conference on Learning Theory , pages =.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Proceedings of Thirty Third Conference on Learning Theory , pages =

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:fea696e5e5ab2cfd097f76c9d8e181e5a2ab93566084eadc0332dfdb3a7c5488

Observation 750b08af-8a36-4a6a-bc19-a502a0cbc4cb · outbound

This paper cites Learning for Dynamics and Control , pages=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Learning for Dynamics and Control , pages=

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:5aad7ff7c94e52fe788604d59107303fed7914bf959c9109570ee59528ef647c

Observation 2ab6bb7f-f4bf-43df-9290-3949d92dd0ce · outbound

This paper cites 2012 , publisher=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control 2012 , publisher=

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:46078386120212786b61be564f36a8289790d0e2108f5ddd11dda5c52a3537f8

Observation 0c6a27e9-382c-4505-b43a-f8eebf07a442 · outbound

This paper cites On Bellman equations for continuous-time policy evaluation I: discretization and approximation.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control On Bellman equations for continuous-time policy evaluation I: discretization and approximation

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:19:43.722513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:fc62b9f7d0c00b02c553ab43da28266ad11ceae9b061c83ff9fa0437ee719353

Observation d9afcbe3-f083-4e61-88ed-de3303f542c4 · outbound

This paper cites Advances in neural information processing systems , volume=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Advances in neural information processing systems , volume=

Reference 16

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:2d6816c2cead273a1694398f8776a3f5f1f3c269a91df9c654f29f496598b7d5

Observation c6008699-7e41-4acf-bf61-0341000a0432 · outbound

This paper cites Proceedings of the AAAI Conference on Artificial Intelligence , volume=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Proceedings of the AAAI Conference on Artificial Intelligence , volume=

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:2df11bb4ebbfa34b19991bd8a64650dd1bb0e8f08fdb034b7fdd5c59209f9036

Observation c0d74310-c4de-4f6e-ad4d-c97510bf9d10 · outbound

This paper cites Proceedings of the 33rd International Conference on Machine Learning , year=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Proceedings of the 33rd International Conference on Machine Learning , year=

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:54179ddafbf0a724f469544ec9a87dd6e4800f31387b60830140ec8947f1b47a

Observation 8ac63e5b-ca1e-4f7b-b4fa-71d0ef671faa · outbound

This paper cites International Conference on Machine Learning , year=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control International Conference on Machine Learning , year=

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:a5b7be93f91e0ea99e1131f1b1268a12c23431b891e82ba8c3fd2edfb3df4ec0

Observation 15ddb8c4-cefe-4f94-ae70-e530de688b26 · outbound

This paper cites International Conference on Machine Learning , year=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control International Conference on Machine Learning , year=

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:182fafc48d08a85aed7e69c607705c289fb082621ad89534097d78e861ef2cb1

Observation 19621f60-3ccd-415a-9074-9ebf79af2bcb · outbound

This paper cites Advances in Neural Information Processing Systems , year=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Advances in Neural Information Processing Systems , year=

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:6b0d9cb2b9b2257a0aa90688031ab5596986d5126b2c2dc251b28c35c274edf9

Observation 053fdde9-d793-47c2-a1db-ebc9126bc6d9 · outbound

This paper cites International Conference on Machine Learning , year=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control International Conference on Machine Learning , year=

Reference 22

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:03a3215bc49d673c4b44402f7f64939b90c496812f11aa49ea98f97b88c19509

Observation bf1cc9bd-5d35-44ba-897b-676e9579f6ca · outbound

This paper cites nature , volume=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control nature , volume=

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:df2049894825b665ed8849c11be65aaa62f44a8b1904448a22ee0dbb84eefaba

Observation 367d7271-099a-4518-a20c-05bb9a19bd78 · outbound

This paper cites Fine-Tuning Language Models from Human Preferences.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Fine-Tuning Language Models from Human Preferences

Reference 24

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T08:19:43.727713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:92decc068e225b0a7a0e05388a18aac25a4e5ca17a4082b24eb0aec5440fdef2

Observation bb19c0c9-32fd-4e67-a652-f445688e1a6f · outbound

This paper cites Journal of Biomedical Informatics , volume=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Journal of Biomedical Informatics , volume=

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:8f8677a4d035a6bcb2a43236e5dea7de28899f1c069e701598562e0ac8acc64c

Observation 855c7827-76bd-4b66-893d-37df84b4e8cd · outbound

This paper cites IEEE reviews in biomedical engineering , volume=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control IEEE reviews in biomedical engineering , volume=

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:f6633620aba6e3205a45b6317fa8eaa4628bc32736426fa80855a6f06b336aef

Observation 3e36efa1-0d45-482c-bc05-fcd60a7accd4 · outbound

This paper cites Advances in neural information processing systems , volume=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Advances in neural information processing systems , volume=

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:8381a42591f0d835a67415e7e151339c8f8c60a76ec00dc81106f7db7f51edd2

Observation f44cb654-dfaf-4dc3-bdfe-6ab863782256 · outbound

This paper cites Diabetes care , volume=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Diabetes care , volume=

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:6cdf706102fc4b7a07248d1abaa6674ac9121799181fc5817432ed81c3fa78e5

Observation 49ebfd26-0602-4599-9382-2cdb3e85c80b · outbound

This paper cites 2009 , PAGES =.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control 2009 , PAGES =

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:b6236c7b863bbb5a2c0150fdc112106f0c46714b3c4e6360f59c7175c11fe516

Observation 0be9f6bc-7a30-4fbc-a4e1-1f40b10b3435 · outbound

This paper cites 1999 , PAGES =.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control 1999 , PAGES =

Reference 30

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:cb5edd00d620a93089a08d249d03b5d1a72208c23876409623a9e0ac3eb8935c

Observation cf9aff3e-d1e9-4aeb-9f94-ef391324fb40 · outbound

This paper cites 2013 , publisher=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control 2013 , publisher=

Reference 31

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:55ebb6ea75430f90130b0f421a4f6046ef67535311e1b02eff01e75f3cf5a7bc

Observation 08f09fbe-dbf9-4e3e-b8af-ef43f1c8c0b1 · outbound

This paper cites 2018 , publisher=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control 2018 , publisher=

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:edb05d01f71c21f736a387b9f70560dbec9b1e3d471329e9b1df33c629e5aa53

Observation 12b26388-1008-47df-86e6-f8ecf738da50 · outbound

This paper cites Communications of the ACM , volume=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Communications of the ACM , volume=

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:f44ca7d7fb89cca80c4981937e65a7425f095381c5d3ecc66b5105db6dee60f2

Observation 0b4a959f-98c3-4204-97ad-8d9828ae82f2 · outbound

This paper cites Machine learning , volume=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Machine learning , volume=

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:9416e07c072f11082bc82d0c9751eed0dda7568fcf72ca4fe7647d820bf668d6

Observation a3f58064-b45e-46ec-bc6b-7ec294a0024e · outbound

This paper cites Neural computation , volume=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Neural computation , volume=

Reference 35

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:cc897c2ba5afb70eba60fd959bb6e96d8ca3a59b955678ce29f406234bad395b

Observation 994c117e-a4a6-447b-8c56-40322671a9e5 · outbound

This paper cites Journal of Machine Learning Research , volume=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Journal of Machine Learning Research , volume=

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:8e9ddb23e246812f329652acdf1c0698f8714cf939d4304315514942f9660975

Observation 8d865529-02cb-459e-bcad-e98ac21a55cf · outbound

This paper cites Mathematical Finance , volume=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Mathematical Finance , volume=

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:93bee4902deddcb0609db2de12566f8c441ee47afeb98ab65bf227a8627cb969

Observation 9f8b2071-5aaf-4f28-b923-af70ebaf0bdb · outbound

This paper cites SIAM Journal on Control and Optimization , volume=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control SIAM Journal on Control and Optimization , volume=

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:f543f6675b0a2889cc105db4167ea041b24dba10b70c3b0ff524e1858477c0a0

Observation 05eef019-2554-42e5-8703-382cf5b13121 · outbound

This paper cites Journal of Machine Learning Research , volume=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Journal of Machine Learning Research , volume=

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:a761816493a7704e4c7e31f10644373f1a99cae7eb61d0247a702af83c8f542a

Observation 67ff858d-e86e-46ef-bdcb-9c6803eaefde · outbound

This paper cites Exploration-exploitation trade-off for continuous-time episodic reinforcement learning with linear-convex models.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Exploration-exploitation trade-off for continuous-time episodic reinforcement learning with linear-convex models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:19:43.714245Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:cc22dd1894030c5cb1066922bf74193473ad52308eb5ddaca53473f9d1339157

Observation 71230e14-7b31-4992-ae3b-8da3e7268cb6 · outbound

This paper cites Proceedings of 1994 IEEE International Conference on Neural Networks (ICNN'94) , volume=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Proceedings of 1994 IEEE International Conference on Neural Networks (ICNN'94) , volume=

Reference 41

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:ec03fb8ccb5971ba29a99a08b1e8671f54389abd6b6498a457f7c8eb5dd59e41

Observation 06b65086-d6d3-4a33-a329-49e9f4da9ce0 · outbound

This paper cites International Conference on Machine Learning , pages=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control International Conference on Machine Learning , pages=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:efe7cdcafcaafb22d0a206063151374fc288cd1444f5c5d6f1325346086acb55

Observation 38dd8a99-14f8-4979-af37-682c35b998f5 · outbound

This paper cites arXiv preprint arXiv:2312.11797 , year=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control arXiv preprint arXiv:2312.11797 , year=

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:19:43.741211Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:43672f9a6a407ed7165d34104b5aad2c796e0bbdd4600708f00d380bca8cedcb

Observation e6df4a6b-d4e8-4945-9440-a0f53987d181 · outbound

This paper cites Accuracy of discretely sampled stochastic policies in continuous-time reinforcement learning.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Accuracy of discretely sampled stochastic policies in continuous-time reinforcement learning

Reference 44

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:19:43.711049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:aedacab530388ab35f5184377baa039b961297e5ddae056ef9985a39169aec8e

Observation a6779442-fbec-4060-a4f5-4003830d89e2 · outbound

This paper cites SIAM Journal on Control and Optimization , volume=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control SIAM Journal on Control and Optimization , volume=

Reference 45

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:07a7306bee84e6059c031a5c5bfcc3d094bf903a7eed7e41ea9e0c74fe06ce68

Observation cbb84136-02e0-4a41-85f4-0d7c5d34707a · outbound

This paper cites Sample and Computationally Efficient Continuous-Time Reinforcement Learning with General Function Approximation.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Sample and Computationally Efficient Continuous-Time Reinforcement Learning with General Function Approximation

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:19:43.738291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:44b750f4149c7b86aad4c42c53dbd1c502b8e050029e40db5490dcae1b51bf33

Observation bd6d39ca-0319-4470-b8c2-21cfa1cac476 · outbound

This paper cites arXiv preprint arXiv:2501.15910 , year=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control arXiv preprint arXiv:2501.15910 , year=

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-07-04T08:19:43.717166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:a09dddde8681bbba9d88f9d66e6f1d15d45352370cc85b0c58536a02c939349d

Observation 03482a95-84e2-4a18-ae21-16ed0955cb42 · outbound

This paper cites Journal of Scientific Computing , volume=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Journal of Scientific Computing , volume=

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:25f49e5656ad039f8a823d8ca6287077d6d8ecf7156c0bea0a3dc6ed4ee667d0

Observation cb9bcc7d-18bf-4e12-a4ee-ebdfe57e34d7 · outbound

This paper cites siam REVIEW , volume=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control siam REVIEW , volume=

Reference 49

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:aeb27820cbb186100ed2bc1564e4bd2d05e386ec4074e27e1b57e9b1ae1ba26f

Observation 530f72c9-89f1-40ed-bbda-ffbeb9089ca3 · outbound

This paper cites an unresolved cited work.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:95150aab3677e0e254d603e59f785c3dfa7cd432a5f18ce9ff8a6b116cfb2a61

Observation e27d9c20-043e-4bae-b7d5-fa6858c9b9cc · outbound

This paper cites GPT-4 Technical Report.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control GPT-4 Technical Report

Reference 51

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T08:19:43.732770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:13c45f4fc252920ad3351284401f31e70549f28a02194a74ae89cb462ab59471

Observation cd053fce-c077-4e74-b0cd-f5cb5fec808d · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 52

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T08:19:43.719766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:872734280d34779a1730687ff3b97fb8a816da63cd06b78fbbbde5f1ea6e8822

Observation c385a958-195c-47d4-ab89-c77c0ea35295 · outbound

This paper cites Proceedings of the 38th International Conference on Machine Learning , pages =.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Proceedings of the 38th International Conference on Machine Learning , pages =

Reference 53

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:8261cae9ce59294d3d4da9afa6225d6314581003a9e228ee3a10303c97fad521

Observation 9f7d6dff-f7bf-42f0-84a6-5afefbb7daa2 · outbound

This paper cites International Journal of Control , volume=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control International Journal of Control , volume=

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:f782b91c07437bd103f8fc9d8950ff1ca8054e4741e1dd1886b6b8aac2f0f2af

Observation dfe4782a-75b7-41fa-9ad9-429e017ec662 · outbound

This paper cites Automatica , volume=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Automatica , volume=

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:84d7571a232b9205ad6404c81fd02ee2f04f32e2ad33d6e3bfcf19e558b6f83a

Observation 5a906685-70ed-4412-9a02-7adfaef95846 · outbound

This paper cites and Soner, H.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control and Soner, H

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:265dc010706548465783c3635e1d084da19b407b4edd361405a44d86bce69256

Observation dfd43f8e-7d29-409b-a7cc-eb86dd3fd42e · outbound

This paper cites Advances in Neural Information Processing Systems , volume=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Advances in Neural Information Processing Systems , volume=

Reference 57

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:587a8c950c28fdc690eb754e6fa583b17926922372e2d61e496af5986ef5c027

Observation b18012b3-f462-4aed-8cd3-ef1127ebfc9e · outbound

This paper cites Optimal-PhiBE: A PDE-based Model-free framework for Continuous-time Reinforcement Learning.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Optimal-PhiBE: A PDE-based Model-free framework for Continuous-time Reinforcement Learning

Reference 58

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T08:19:43.730286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:77d80b68ebc9530383ca66fce2d5d86698777b880d2ae3d4efe923176f1380d0

Observation 279677a0-3c9e-4781-aa95-284ab8c065cd · outbound

This paper cites Foundations and Trends.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Foundations and Trends

Reference 59

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:bcc65607822b6685c540df41f781592b59fe22cdb4fe4019c93bd140f69d0448

Observation 990d68b9-c7ee-4c6d-9d52-9ec1dcc31224 · outbound

This paper cites The Thirty Sixth Annual Conference on Learning Theory , pages=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control The Thirty Sixth Annual Conference on Learning Theory , pages=

Reference 60

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:7571e651f524255a951addb09579a31508f952128e1350b218e50bb267cbc09b

Observation d5ad3f98-74fc-4d20-86c6-6614bfe0eb44 · outbound

This paper cites 1998 , publisher=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control 1998 , publisher=

Reference 61

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:6c46085fa5ebae53c94f514e087cd81de15b3858812f2fda659c3ea517332623

Observation 8c8c50dd-4ccb-4b62-90e5-f08ee8e9ca14 · outbound

This paper cites Proceedings of The Web Conference 2020 , pages=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control Proceedings of The Web Conference 2020 , pages=

Reference 62

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:afe8c34333b9309a443828bec98cc0f99360e75674da2be21ca5e4feb4cfb5cb

Observation 92bee1c0-48ed-4617-b883-8a4cfed91be3 · outbound

This paper cites 2013 , publisher=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control 2013 , publisher=

Reference 63

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:9cdc48753f19e441532a67f28493248b434d03b93bf6dc049a2baa1cb7c4cbf1

Observation feb815fa-22da-4f6f-b8db-a1a88838bbc2 · outbound

This paper cites 2006 , publisher=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control 2006 , publisher=

Reference 64

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:82da578dc19213bb18d1be93a8fe119a9a3a5fe5212afd755893078d3daa7235

Observation 884ddb00-8c2a-446c-aabf-7dd0336dbc9e · outbound

This paper cites 2003 , publisher=.

PhiBE-Q-Learning: Bridging Off-Policy Reinforcement Learning and Continuous-Time Control 2003 , publisher=

Reference 65

Resolution
unresolved
no resolver link, observed 2026-06-26T11:59:18.223000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-06-26T11:59:18.223000Z digest=sha256:2b21a6bf53bebc8cc90179becce1ce652d0700ae66861db066f45ef09d721083

Pith citing papers

No inbound Pith citation observations are available.