Pith. sign in

Paper Citation Record · LEDGER

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction

As of 20 August 2026, this Paper Citation Record lists 14 of 14 outbound references and 0 inbound Pith citation observations for arXiv:2507.16718.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.16718 v1

Coverage vector

measured 14 of 14 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:06:22.975019Z

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

14 of 14 outbound references displayed

  • verified exact1
  • verified fuzzy3
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 38c28270-dfbb-43b8-9abc-c2780e9f88a4 · outbound

This paper cites Qwen2.5-VL Technical Report.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:22.938905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:22.938905Z digest=sha256:57007cdf672f0fcb85f9b5a1566f497b88b78118016d319a8ebb30530b472dcb

Observation 3f2176c7-d82d-4ff1-a70c-a5d5436599ad · outbound

This paper cites Towards Intelligent Speech Assistants in Operating Rooms: A Multimodal Model for Surgical Workflow Analysis.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Towards Intelligent Speech Assistants in Operating Rooms: A Multimodal Model for Surgical Workflow Analysis

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:22.942371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:22.942371Z digest=sha256:a5d40755051d9417571a9e541dba6a20d2384fbc8860edced7479f244725428a

Observation 51299db0-0732-41a4-8e82-c53c0aa15b54 · outbound

This paper cites A spatio- temporal network for video semantic segmentation in surgical videos.International Journal of Computer Assisted Radiology and Surgery , 19(2):375–382, 2024.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction A spatio- temporal network for video semantic segmentation in surgical videos.International Journal of Computer Assisted Radiology and Surgery , 19(2):375–382, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:06:23.119196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:06:22.945334Z digest=sha256:5590f04328f4c2af940e9ae58ee348e874d4a33b23a4f9cf5304354d9c430ddd

Observation f854057d-338e-4e6c-9d27-a28e435522a2 · outbound

This paper cites Temporal memory relation network for workflow recognition from surgical video.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Temporal memory relation network for workflow recognition from surgical video

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:06:23.109869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:06:22.948068Z digest=sha256:154bc461cdc43fee55cd81d8087464b47eaf7fd8065ab7fd4818f089e95028cc

Observation b557b853-3366-48ef-884e-1855b990b60c · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Lisa: Reasoning segmentation via large language model

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:06:23.100640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:06:22.950833Z digest=sha256:3efef712c279d237985dbf919cd5eb4f5d15e3ffd0c70963af679506237c34a8

Observation bcda117c-9cef-428e-9288-5287183f452f · outbound

This paper cites Improved baselines with visual instruction tuning.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Improved baselines with visual instruction tuning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:22.953421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:22.953421Z digest=sha256:ec3797464321e38f40f983235217e2c26f216a2e4b3db59891ac26840446b6db

Observation e9ee0c61-ecca-434e-b589-1fe233297714 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction SAM 2: Segment Anything in Images and Videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:22.956091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:22.956091Z digest=sha256:512c05f5e0b23e355a1a4d6493c73923ec05aabb2340acbe579e36b00940a567

Observation 0e0199d5-8493-4970-832a-e4eeb65e3214 · outbound

This paper cites Position: Foundation Models Need Digital Twin Representations.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Position: Foundation Models Need Digital Twin Representations

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:22.959059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:22.959059Z digest=sha256:07521296e09f84b8a061c86ddaffa87340daf8f583ca2b2fda1bedd1b3a74b51

Observation 9c38a73e-c019-4a1e-a51a-9686a1255dcc · outbound

This paper cites RVTBench: A Benchmark for Visual Reasoning Tasks.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction RVTBench: A Benchmark for Visual Reasoning Tasks

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:06:23.035823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:06:22.961694Z digest=sha256:b4af57421c6ec71090e117110c3e5da42a0e110d943a49f58bc997d43b5a0b11

Observation caf26f12-4e97-48c3-b8e7-52893dc495fb · outbound

This paper cites Operating Room Workflow Analysis via Reasoning Segmentation over Digital Twins.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Operating Room Workflow Analysis via Reasoning Segmentation over Digital Twins

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:22.964559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:22.964559Z digest=sha256:fc230e3636db9684c5de98b1ec3d812d54fa65f1d5de144f18a53fa05c436137

Observation 6507a37d-5882-44da-844d-3acb78c39a8e · outbound

This paper cites Reasoning Segmentation for Images and Videos: A Survey.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Reasoning Segmentation for Images and Videos: A Survey

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:22.967237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:22.967237Z digest=sha256:5985fe4993d27f673428005b34e1134bb6187a62dd75468564e2009dceb5edf3

Observation 86b35c74-1a23-4b64-8af0-19d51503112d · outbound

This paper cites MVOR: A Multi-view RGB-D Operating Room Dataset for 2D and 3D Human Pose Estimation.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction MVOR: A Multi-view RGB-D Operating Room Dataset for 2D and 3D Human Pose Estimation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:22.969839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:22.969839Z digest=sha256:b6cbccd0f93a2d7bd373869674d5a19da83291c1c326f3c561ba975a17fac06f

Observation bda4ed03-9f55-4b56-a48d-4da90aa1cd5f · outbound

This paper cites an unresolved cited work.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:06:23.084249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T15:06:22.972763Z digest=sha256:20082135d6cf4b6f1d7f92a362581ee0908a2d01fc43b9fde3bcfffe87e4fcf6

Observation e6792027-f9b5-4af3-855a-b28b710000d3 · outbound

This paper cites Depth Anything V2.

Temporally-Constrained Video Reasoning Segmentation and Automated Benchmark Construction Depth Anything V2

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:06:22.975019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:06:22.975019Z digest=sha256:f8fe701f399b011273502ddea0fdfc582573a5f7a6e21203e289988aa1e4fe69

Pith citing papers

No inbound Pith citation observations are available.