Pith. sign in

Paper Citation Record · LEDGER

Reinforcement Learning in hyperbolic space for multi-step reasoning

As of 14 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 1 inbound Pith citation observation for arXiv:2507.16864.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.16864 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:24:27.883078Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-09T23:51:47.724033Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy1
  • unresolved29
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 9f291b92-9ff4-42fc-892d-ecaf5592c824 · outbound

This paper cites Advancing Reasoning in Large Language Models: Promising Methods and Approaches.

Reinforcement Learning in hyperbolic space for multi-step reasoning Advancing Reasoning in Large Language Models: Promising Methods and Approaches

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:25.191374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:25.191374Z digest=sha256:f9b940c7b335a9468e61f0e29609229b9da9de01a923498e35defdc59d14ca92

Observation e791a1a2-9549-4505-a103-4353d0e036f4 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning in hyperbolic space for multi-step reasoning Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:24:30.736509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:24:25.289751Z digest=sha256:89d81c1d12efc4cbb8904d9f882fc1bcb071b20d364f7368c5bb252bae2e57c7

Observation 413e08c6-4fc0-45ed-8c58-1983936e23ed · outbound

This paper cites an unresolved cited work.

Reinforcement Learning in hyperbolic space for multi-step reasoning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:24:30.656103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:24:25.385382Z digest=sha256:adf6a3d48205b6d0ca02759ac450f4f986c73549f95d02309da7a5dc3c222e42

Observation 03a384d6-6603-4d12-a5ca-d098dad036e3 · outbound

This paper cites Multi-step reinforcement learning: A unifying algorithm.

Reinforcement Learning in hyperbolic space for multi-step reasoning Multi-step reinforcement learning: A unifying algorithm

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:24:30.474757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:24:25.600915Z digest=sha256:c6fc20c270fa58707bd16cfd11ec45889a6e7c629a8ff02c8fc15efc57b952ce

Observation 6da6aa9e-b692-4407-a50d-263863717101 · outbound

This paper cites DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model.

Reinforcement Learning in hyperbolic space for multi-step reasoning DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:25.721846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:25.721846Z digest=sha256:ca5d18fe9f5f6ea61960bc8d2ee844f25b90a96861bec57e58390a63ecd155e2

Observation aaffbf05-6329-479d-bf10-f10098639611 · outbound

This paper cites Process-Supervised Reinforcement Learning for Code Generation.

Reinforcement Learning in hyperbolic space for multi-step reasoning Process-Supervised Reinforcement Learning for Code Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:25.819837Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:25.819837Z digest=sha256:96d3404fd451c8a5d086522301b3656173cb4afd9b59c1fb67e00b0e687e0194

Observation 6c9a4a8c-d07e-4988-a050-86cb06ce2e57 · outbound

This paper cites System 2 Reasoning for Human-AI Alignment: Generality and Adaptivity via ARC-AGI.

Reinforcement Learning in hyperbolic space for multi-step reasoning System 2 Reasoning for Human-AI Alignment: Generality and Adaptivity via ARC-AGI

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:25.874748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:25.874748Z digest=sha256:8141f2076506691c050a4676552657348087e9cca19885b608fa23055a868246

Observation 47f297f6-ad03-4245-9655-2cd4e3090e17 · outbound

This paper cites Mechanistic Interpretability for AI Safety -- A Review.

Reinforcement Learning in hyperbolic space for multi-step reasoning Mechanistic Interpretability for AI Safety -- A Review

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:25.984891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:25.984891Z digest=sha256:04e67e492f0ad7fc4e2780715c1ebd771404d5302620b5cd7ddda7fec8a75ea3

Observation 0356ed62-c5dd-4f6b-91f6-952dee763cc6 · outbound

This paper cites Transformers in Reinforcement Learning: A Survey.

Reinforcement Learning in hyperbolic space for multi-step reasoning Transformers in Reinforcement Learning: A Survey

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:26.050901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:26.050901Z digest=sha256:a9f45fca30ca7d13ba4652285285b0c189583566eeae362830feb1cf763bb46b

Observation 5700f723-f5e9-4840-b5d7-8340f3c30fef · outbound

This paper cites A Survey on Transformers in Reinforcement Learning.

Reinforcement Learning in hyperbolic space for multi-step reasoning A Survey on Transformers in Reinforcement Learning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:26.139099Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:26.139099Z digest=sha256:d5ffb8a4d738ed5a9607e664e56f3528d9cc340835aeac1045cfb64a147b1247

Observation 3fbb0660-e539-4f30-93da-fcf734f3e99a · outbound

This paper cites Deep Transformer Q-Networks for Partially Observable Reinforcement Learning.

Reinforcement Learning in hyperbolic space for multi-step reasoning Deep Transformer Q-Networks for Partially Observable Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:26.214913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:26.214913Z digest=sha256:a46e700b3f22f8cda3df2563b4f59928a019868e93189f791cbf9f2ada80e062

Observation ae8e3e84-d7d0-42fe-886a-3e73e6e824f5 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning in hyperbolic space for multi-step reasoning Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:24:30.322606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:24:26.312548Z digest=sha256:af7c06c984b689c7df2be7f96de3b2be336831f6efcedfd027992d3bad1a763c

Observation 3febb1cf-f523-4cb9-8de8-7d08acdfd19c · outbound

This paper cites TransDreamer: Reinforcement Learning with Transformer World Models.

Reinforcement Learning in hyperbolic space for multi-step reasoning TransDreamer: Reinforcement Learning with Transformer World Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:26.376531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:26.376531Z digest=sha256:fce856a570c344ceb943fa7f63fb5a4735c58a4014dd72832dae31f8f498a9d4

Observation 9f2d7c5d-0f5e-4bcf-94f6-8b4c9700e7ec · outbound

This paper cites an unresolved cited work.

Reinforcement Learning in hyperbolic space for multi-step reasoning Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:24:30.170128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:24:26.447389Z digest=sha256:ac07694fb4919c447a4d96d72c33f0d70fd63b1f54f38cbb18b85f9d72adcdbd

Observation 3a5f4cac-de24-471d-8db3-3581bef3df7d · outbound

This paper cites an unresolved cited work.

Reinforcement Learning in hyperbolic space for multi-step reasoning Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:24:30.007428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:24:26.491830Z digest=sha256:ec98551e8c4a3c7f1df580b47da2e8394972ff7aeba2997bcb8326844b9ea866

Observation a906c0a1-69e0-4378-930d-5c71a3242d7a · outbound

This paper cites Understanding the Language Model to Solve the Symbolic Multi-Step Reasoning Problem from the Perspective of Buffer Mechanism.

Reinforcement Learning in hyperbolic space for multi-step reasoning Understanding the Language Model to Solve the Symbolic Multi-Step Reasoning Problem from the Perspective of Buffer Mechanism

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:26.585807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:26.585807Z digest=sha256:851e8396a4dc3a0e88977959ea7f0e873e9105fd87b19c2a68c5ed9f880b102d

Observation b41e75c3-aeea-4c77-91fd-e1d67ff4f3eb · outbound

This paper cites Enhancing Multi-Step Reasoning Abilities of Language Models through Direct Q-Function Optimization.

Reinforcement Learning in hyperbolic space for multi-step reasoning Enhancing Multi-Step Reasoning Abilities of Language Models through Direct Q-Function Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:26.666597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:26.666597Z digest=sha256:8f3992c96544a8fc90a8d6b55a58cdab6e9cd82a419922bac75a77aa08da9ab7

Observation de208ac0-a99d-4c4e-9d92-bd2b09d0416c · outbound

This paper cites an unresolved cited work.

Reinforcement Learning in hyperbolic space for multi-step reasoning Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:24:29.852764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:24:26.793498Z digest=sha256:9afc8ad2e9b57056e7583285e792f5c1a7ef9698bb1cba0480e87a6038804aed

Observation f0e5f976-146d-40fe-9e6a-ced587c64e51 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning in hyperbolic space for multi-step reasoning Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:24:29.689007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:24:26.889102Z digest=sha256:5ee0ae0d7a49ca1466020516f0acf42eed8ecf000596dd549dae08535b07b618

Observation efb3d764-7090-4e84-8a54-65a4ec6c0b1c · outbound

This paper cites an unresolved cited work.

Reinforcement Learning in hyperbolic space for multi-step reasoning Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:24:29.547578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:24:26.986934Z digest=sha256:4e863c843dd8d98b595aac8d3ad810de3746462774d075682cc82445579d2b5c

Observation 8cc50101-0e0b-45a8-81f1-4a2740de302c · outbound

This paper cites an unresolved cited work.

Reinforcement Learning in hyperbolic space for multi-step reasoning Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:24:29.386796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:24:27.052694Z digest=sha256:af94ee70da807e4d8501ba55798876bddeab51a480cec648ed052cb567f3252c

Observation ee9628fd-abe5-4c13-92e9-51b449639072 · outbound

This paper cites Poincar\'e GloVe: Hyperbolic Word Embeddings.

Reinforcement Learning in hyperbolic space for multi-step reasoning Poincar\'e GloVe: Hyperbolic Word Embeddings

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:27.151997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:27.151997Z digest=sha256:de74bcc0da33bf27c95577bd62ee5387ad28df81747486d88eba66ca37a0b63d

Observation 86128f63-7e99-4d89-b410-f9c31ad43f95 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning in hyperbolic space for multi-step reasoning Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:24:29.249154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:24:27.220722Z digest=sha256:9172e29c051db72664c4c22c68a5e340c7e2b705ef328bd3b9b19627a69c8856

Observation 19b80823-8346-4964-95f5-fd171b1c823b · outbound

This paper cites TransMLA: Multi-Head Latent Attention Is All You Need.

Reinforcement Learning in hyperbolic space for multi-step reasoning TransMLA: Multi-Head Latent Attention Is All You Need

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:27.310795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:27.310795Z digest=sha256:43ad4ee2e90605e4d3e552d4f7fabce5030644c2ec6492ba8763382d1dbc1846

Observation c910248f-3e95-48f7-b327-122eeb0e90bb · outbound

This paper cites an unresolved cited work.

Reinforcement Learning in hyperbolic space for multi-step reasoning Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:24:29.068005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:24:27.405059Z digest=sha256:3f26e3080f9e0de458cec12d01282bef54f21768e7d85da6202bf731bc83018b

Observation 0fd0795c-86f6-41e1-b087-1f98c1a8ac92 · outbound

This paper cites an unresolved cited work.

Reinforcement Learning in hyperbolic space for multi-step reasoning Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:24:28.881615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:24:27.511247Z digest=sha256:5079bcccd35b82e121a672d78f15488fef18046692e633064e53d51eac2389c7

Observation cf8a0260-50bd-43bf-a72c-03e63185ed2c · outbound

This paper cites an unresolved cited work.

Reinforcement Learning in hyperbolic space for multi-step reasoning Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:24:28.715723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:24:27.637964Z digest=sha256:63016f81787905807928fe192864f9257e23184fbdae1d44653ec683ccc7d7e0

Observation 7e0152b2-7e3e-4348-be31-8bd04f407e02 · outbound

This paper cites Offline Reinforcement Learning for LLM Multi-Step Reasoning.

Reinforcement Learning in hyperbolic space for multi-step reasoning Offline Reinforcement Learning for LLM Multi-Step Reasoning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:27.704354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:27.704354Z digest=sha256:22b703054b897d81821597d871a523d1976c284d88dc6c44ef97a840b7536a9b

Observation 3495c8fb-34ed-46ca-86e3-adfeb0792238 · outbound

This paper cites Transformers in Reinforcement Learning: A Survey.

Reinforcement Learning in hyperbolic space for multi-step reasoning Transformers in Reinforcement Learning: A Survey

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T15:24:27.790125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:24:27.790125Z digest=sha256:882c5a5a18683033a68f40ad4c84c986f3e455d485e7bd244e16923f8327720c

Observation 5e63b8d5-7ed3-4a55-99b9-9c3cde6264ca · outbound

This paper cites an unresolved cited work.

Reinforcement Learning in hyperbolic space for multi-step reasoning Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:24:28.562546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T15:24:27.883078Z digest=sha256:de250a802a6bfbf18b83625448f475d582f7ef6ae1948ebf1c516f54ab3174dc

Pith citing papers

Observation d5b84770-5e8f-45f9-bf6d-9cdc8a10a36d · inbound

HypEHR: Hyperbolic Modeling of Electronic Health Records for Efficient Question Answering cites this paper.

HypEHR: Hyperbolic Modeling of Electronic Health Records for Efficient Question Answering Reinforcement Learning in hyperbolic space for multi-step reasoning

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:54:45.604319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-05-09T23:51:47.724033Z digest=sha256:578b5f8cf17b6a900481725c31049ea933645ce031248ed37bccc732b88134e2