Pith. sign in

Paper Citation Record · LEDGER

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making

As of 15 August 2026, this Paper Citation Record lists 20 of 20 outbound references and 1 inbound Pith citation observation for arXiv:2607.09322.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.09322 v1

Coverage vector

measured 20 of 20 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-13T03:54:11.805387Z

measured 21 of 21 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:18:05.046139Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T04:18:05.490656Z

Reference resolution

20 of 20 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved19
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c2dc38c9-bdc1-4e7b-8961-534e6ede89a9 · outbound

This paper cites Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making Mem0: Building Production-Ready AI Agents with Scalable Long-Term Memory

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:8b355e8c29e96b5275538536a2ec6b9f62626418e28c3bfec417715dab8e1b10

Observation 28d8a0e3-5c05-4c54-9ba4-d1888d6d36a7 · outbound

This paper cites an unresolved cited work.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:0ac5643b352767a17b3607602527d32e87fc05fdb737dda0190af32e02d728b2

Observation 9b1a0c8c-5a39-4331-aa4d-3e96d3ba3be8 · outbound

This paper cites an unresolved cited work.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:0d07b3d04c00e8c2023b1ad0ceb91ebea2e4643e62bbfefa2509e85ecf7aa12c

Observation 71f76bde-a72e-49e4-8407-3fab2c8dcce8 · outbound

This paper cites Journal of medical Internet research 27, e84120 (2025).

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making Journal of medical Internet research 27, e84120 (2025)

Reference 4

Resolution
malformed identifier
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:b07afdafb86b5f42b586fc803bebd33588106ff2a16ad007fb13f8cad228b101

Observation 0bf90271-06f2-4de6-b820-f1e188789564 · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:835de3919c0a2710dc7288f46815c6070b4200f6029c0e0bf5fc95d04c11865e

Observation f0f38a46-759d-4f96-8a0b-d64f15ac5029 · outbound

This paper cites NEJM AI p.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making NEJM AI p

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:fb39607add31794d2ef4974123ffe1c6f88f194a2aed362524a936ccb27db52c

Observation 82dda204-1e89-4c0c-b6a7-cf6f494e0323 · outbound

This paper cites PhysioNet (Oct 2024).

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making PhysioNet (Oct 2024)

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:42ae593090f844f9072ae44ec531d6f67a54370fd3096846f2373978c0a5697a

Observation 9b5a0c37-3660-428a-87d4-636e8b001b28 · outbound

This paper cites The NarrativeQA Reading Comprehension Challenge.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making The NarrativeQA Reading Comprehension Challenge

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:064288d7c0b96bd7306ded1724de73523e2b51b0dc3330c876c940a59d760287

Observation 7d106ae4-3c01-43c2-b565-11a61e6529d8 · outbound

This paper cites Advances in Neural Information Processing Systems 35, 15589–15601 (2022).

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making Advances in Neural Information Processing Systems 35, 15589–15601 (2022)

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:168f3c649d922de60a486be77d7920ffef9c1109ef10f8c65e165704b1ff6dce

Observation ccbd14d6-5ced-4988-890d-f5e5a86af649 · outbound

This paper cites ReflecTool: Towards Reflection-Aware Tool-Augmented Clinical Agents.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making ReflecTool: Towards Reflection-Aware Tool-Augmented Clinical Agents

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:95084428688edbbf25015ee44aac4596b36d1c6ed3792a4a87bae41dce4530cb

Observation 12c92b9d-138e-4f9e-989f-b1932c4b5d59 · outbound

This paper cites Lost in the Middle: How Language Models Use Long Contexts.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making Lost in the Middle: How Language Models Use Long Contexts

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:2ca2524a4d889698d8537e1c779f5b9d444eead8b6dc455a2aec88298ce6f308

Observation ab4aa5d5-7d36-45c2-98df-64383f0f16bf · outbound

This paper cites Evaluating Very Long-Term Conversational Memory of LLM Agents.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making Evaluating Very Long-Term Conversational Memory of LLM Agents

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:5ef976e8cba1c2c4c46790fc1ec16553e4797a2f8a312b228aa8c25ed0b3a7bd

Observation faaa1bf4-56e8-4a62-a837-9eafcb49ecec · outbound

This paper cites https://developers.openai.com/api/docs/models/gpt- 5-mini (2026), accessed 22 Feb 2026.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making https://developers.openai.com/api/docs/models/gpt- 5-mini (2026), accessed 22 Feb 2026

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:6b95ef498373f03957787c9af84720c299ddf35e6f495a2c76eadef93d2121ae

Observation 70e9150b-08d4-491c-b4dd-941c922a4ddb · outbound

This paper cites https://developers.openai.com/api/docs/models/text- embedding-3-small (2026), accessed 22 Feb 2026.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making https://developers.openai.com/api/docs/models/text- embedding-3-small (2026), accessed 22 Feb 2026

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:bd174fe5e90d60beea8a910d78abc08897d3b46c1abb77d6ea72c7b6e1b4cd96

Observation d8dcca31-e913-427f-ab89-399fad0e1c7f · outbound

This paper cites an unresolved cited work.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:994bce3503dc05443816d1b039b65c5a21127c4d5a0913e1a89eb06c7d75ae51

Observation 936a5c28-b295-4068-b876-b2adf1a71291 · outbound

This paper cites https://qwen.ai/blog?id=qwen3 (2024), accessed: 22 Feb 2026.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making https://qwen.ai/blog?id=qwen3 (2024), accessed: 22 Feb 2026

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:9ea1092eff8e668e06d9c9dd8e2a955f417d1b4634fea21021c48b5b1150e910

Observation dc345ca6-188b-488c-81de-58a74508ceb7 · outbound

This paper cites AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:c76f9d962945dc18e386280069967705145b9adb1792d71907d64315374237b0

Observation 570233fe-6d1d-48d5-94fb-843e3950cecf · outbound

This paper cites an unresolved cited work.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:d8d865ee5ae7dbef65becabda8411e9a18abb4924f7d7a411ae22e1563169e12

Observation feaafc4f-d5aa-46e7-a84b-b8207ae98d3f · outbound

This paper cites LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory.

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making LongMemEval: Benchmarking Chat Assistants on Long-Term Interactive Memory

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:c74ca624a7b713a0d3c945fa5c58168d7b5799a3732218bdc27b5f328b5d72a4

Observation 0777cc87-ae87-4d29-bf74-8ca3aa481b63 · outbound

This paper cites In: Conference on Empirical Methods in Natural Language Processing (EMNLP) (2018).

LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making In: Conference on Empirical Methods in Natural Language Processing (EMNLP) (2018)

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-13T03:54:11.805387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T03:54:11.805387Z digest=sha256:a80316f7856ae2c251378d754981b3a223f4826b8f374d330229bf4338e90702

Pith citing papers

Observation 6b953d02-81ef-47eb-b079-2255cc91aa7c · inbound

CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR cites this paper.

CliniCARE-Bench: Clinical Calibrated Audit of Medical Reasoning in EHR LongMedBench: Benchmarking Medical Agents for Long-Horizon Clinical Decision-Making

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:18:05.497141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T04:18:05.046139Z digest=sha256:d8de4dec7959032123453c71a1c041405b393aa23f1b3e9d5cf8e02948fe8c81