Pith. sign in

Paper Citation Record · LEDGER

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction

As of 21 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2507.16718.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.16718 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:06:22.975019Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 38c28270-dfbb-43b8-9abc-c2780e9f88a4 · outbound

This paper cites Qwen2.5-VL Technical Report.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:22.938905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:22.938905Z digest=sha256:4b3ee6d3401546ad84d0767500987c2e15d1dacc1edcab6fb16798315cc4dbe9

Observation 3f2176c7-d82d-4ff1-a70c-a5d5436599ad · outbound

This paper cites Towards Intelligent Speech Assistants in Operating Rooms: A Multimodal Model for Surgical Workflow Analysis.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Towards Intelligent Speech Assistants in Operating Rooms: A Multimodal Model for Surgical Workflow Analysis

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:22.942371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:22.942371Z digest=sha256:97735cf657dd91212e27887c50f45035cc14cb166efd42398637f225b0a2551d

Observation 51299db0-0732-41a4-8e82-c53c0aa15b54 · outbound

This paper cites A spatio- temporal network for video semantic segmentation in surgical videos.International Journal of Computer Assisted Radiology and Surgery , 19(2):375–382, 2024.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction A spatio- temporal network for video semantic segmentation in surgical videos.International Journal of Computer Assisted Radiology and Surgery , 19(2):375–382, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:06:23.119196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:06:22.945334Z digest=sha256:ae9f29a5e23a2a48537aa4754dd84a6dd05af2362e339051314603d5c24d84ff

Observation f854057d-338e-4e6c-9d27-a28e435522a2 · outbound

This paper cites Temporal memory relation network for workflow recognition from surgical video.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Temporal memory relation network for workflow recognition from surgical video

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:06:23.109869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:06:22.948068Z digest=sha256:c196fef0b18b92f0016827416573f7bd4153d87141a37f483c3e8ee72959c9c2

Observation b557b853-3366-48ef-884e-1855b990b60c · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Lisa: Reasoning segmentation via large language model

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:06:23.100640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:06:22.950833Z digest=sha256:223bd0628cbcf28ea5bb3074ffe8b26b536fc73ea3f91bb083295163a78254de

Observation bcda117c-9cef-428e-9288-5287183f452f · outbound

This paper cites Improved baselines with visual instruction tuning.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Improved baselines with visual instruction tuning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:22.953421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:22.953421Z digest=sha256:e0f301034b150a6c19b338855e135bd5e246555a40dcdd3454089f30c14fa87d

Observation e9ee0c61-ecca-434e-b589-1fe233297714 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction SAM 2: Segment Anything in Images and Videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:22.956091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:22.956091Z digest=sha256:a289901c50ed43f354d9792157d2c77b2ba4ccd5aba0df6a18245c97503d5c9f

Observation 0e0199d5-8493-4970-832a-e4eeb65e3214 · outbound

This paper cites Position: Foundation Models Need Digital Twin Representations.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Position: Foundation Models Need Digital Twin Representations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:22.959059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:22.959059Z digest=sha256:bba54ff83687efa552d8c4c63863a91cea9bea4baad2dfbcba67eefb2cc85a9b

Observation 9c38a73e-c019-4a1e-a51a-9686a1255dcc · outbound

This paper cites RVTBench: A Benchmark for Visual Reasoning Tasks.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction RVTBench: A Benchmark for Visual Reasoning Tasks

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:06:23.035823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:06:22.961694Z digest=sha256:9a721c16432ff211e861776d03b34af2266cbb6c5ac9ebfc69e05027a0f479c4

Observation caf26f12-4e97-48c3-b8e7-52893dc495fb · outbound

This paper cites Operating Room Workflow Analysis via Reasoning Segmentation over Digital Twins.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Operating Room Workflow Analysis via Reasoning Segmentation over Digital Twins

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:22.964559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:22.964559Z digest=sha256:d36f8bf79d47421c611818dfaa143734b9bc611e470225feee0e36f0aafb5335

Observation 6507a37d-5882-44da-844d-3acb78c39a8e · outbound

This paper cites Reasoning Segmentation for Images and Videos: A Survey.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Reasoning Segmentation for Images and Videos: A Survey

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:22.967237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:22.967237Z digest=sha256:2124f3c50f5019dee3fb0cfc762272d4521e8dc73333f1a861b8f9cf71b02358

Observation 86b35c74-1a23-4b64-8af0-19d51503112d · outbound

This paper cites MVOR: A Multi-view RGB-D Operating Room Dataset for 2D and 3D Human Pose Estimation.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction MVOR: A Multi-view RGB-D Operating Room Dataset for 2D and 3D Human Pose Estimation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:22.969839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:22.969839Z digest=sha256:50e44a4ed50722d0c77af99aef93869e6be98d10c1dac7c045d9b6f43a7decf3

Observation bda4ed03-9f55-4b56-a48d-4da90aa1cd5f · outbound

This paper cites an unresolved cited work.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:06:23.084249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T15:06:22.972763Z digest=sha256:4943098b620424fad1020a9228e2dfab11a4d790f3435d2ebcd5e91419b86c5e

Observation e6792027-f9b5-4af3-855a-b28b710000d3 · outbound

This paper cites Depth Anything V2.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Depth Anything V2

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:22.975019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:22.975019Z digest=sha256:7b581a194a638604e6c462fdb87d8f6806c1d69bd1f7921a61e75d4caeb60cc3

Pith citing papers

No inbound Pith citation observations are available.