Pith. sign in

Paper Citation Record · LEDGER

Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2312.10300.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2312.10300 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T21:11:42.144841Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:38:56.149993Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 4c9ee0e4-a239-41b4-8a4d-371a50d25083 · inbound

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs cites this paper.

Text-to-Edit: Controllable End-to-End Video Ad Creation via Multimodal LLMs Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T21:11:42.144841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T21:11:42.144841Z digest=sha256:7d5ef65f1a594d124da8d678e32fa05ac940599c380ddccba51560609d61ab36

Observation 5eead944-f033-4430-97d5-61140c8c3ef9 · inbound

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking cites this paper.

VideoCap-R1: Enhancing MLLMs for Video Captioning via Structured Thinking Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T11:40:11.807797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:40:11.807797Z digest=sha256:813bfc000766f03ac64e45edbd7c02fa47f2f2d7826798dd6a186a9d782130d5

Observation 3959a487-d5e7-4256-979c-3260fefde00a · inbound

CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models cites this paper.

CausalVQA: A Physically Grounded Causal Reasoning Benchmark for Video Models Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T04:44:59.302001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:44:59.302001Z digest=sha256:27a39bbde4b7f37c91dd44012b7f7cc18af415022b44dd7608166dd2ec0a1721

Observation e5aed2b2-69a5-42fe-87bd-b4dce7efe8ad · inbound

Comparing Learning Paradigms for Egocentric Video Summarization cites this paper.

Comparing Learning Paradigms for Egocentric Video Summarization Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T22:23:41.969901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:23:41.969901Z digest=sha256:105b1c9f5becb94c7de1c9c377dc14a6fec896d5cfbed7fb8bcdc8c947cc3f98

Observation 6f53dc11-111a-4c9e-9ca0-ef9b3f4c1988 · inbound

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding cites this paper.

AuroraLong: Bringing RNNs Back to Efficient Open-Ended Video Understanding Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T20:29:50.482506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:29:50.482506Z digest=sha256:7e2b3d8a6ee7dc0d43441ea481f75dffe005e3d2e3cf7678da872316f0e605ea

Observation 6b864fc9-9403-4a05-a35b-e0262781c528 · inbound

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting cites this paper.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:17:57.824025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:17:57.824025Z digest=sha256:4a8767c1f8b2a8d78044eccd45443135d323afa0c89c4d7506aaad3ceef8e183

Observation 71c663e5-043f-4578-b719-2fe8944040cd · inbound

NoteIt: A System Converting Instructional Videos to Interactable Notes Through Multimodal Video Understanding cites this paper.

NoteIt: A System Converting Instructional Videos to Interactable Notes Through Multimodal Video Understanding Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T18:39:29.914405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:39:29.914405Z digest=sha256:bff3466525c15721a874892118bd79a354c8e583a23af5c38629faa9fcda5929

Observation 072fa698-ddc5-4a3a-878a-c96338c1d383 · inbound

Streaming Video Instruction Tuning cites this paper.

Streaming Video Instruction Tuning Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:48:21.847902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T19:44:11.032898Z digest=sha256:bf6f60ddcf0a38792b270f4dbceacb772709882720531b4b915ef5a7c85c6154

Observation de5a588d-53e0-4229-a971-0412b48b708f · inbound

MCSC-Bench: Multimodal Context-to-Script Creation for Realistic Video Production cites this paper.

MCSC-Bench: Multimodal Context-to-Script Creation for Realistic Video Production Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:13:29.933725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T09:10:31.245928Z digest=sha256:3d804b5816db8119ff55a5e30e93641d98b5febf8ab89b86a03ce889710e25ca

Observation 1b527f97-cdc2-4c25-a66d-72bd07c71120 · inbound

MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation cites this paper.

MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T21:11:19.082732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-08T06:28:42.129881Z digest=sha256:ba16b51e009440db308ab3d6b5f3d3450f3f4aa1aca3a8415a735b1c2d046793

Observation 78b7ed1d-47e0-4726-993f-b16a86d3fa47 · inbound

MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation cites this paper.

MuSS: A Large-Scale Dataset and Cinematic Narrative Benchmark for Multi-Shot Subject-to-Video Generation Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 12

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T00:51:15.120188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T00:50:10.509727Z digest=sha256:693c6416ed2097dd9f40952ff3ed4a846f8db59d18b9c5f6fff9b07568ad082f

Observation afdedb8a-49e6-476f-9ab3-44a88148bfaf · inbound

Harnessing Streaming Video in the Wild cites this paper.

Harnessing Streaming Video in the Wild Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-02T22:37:25.681985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T18:47:55.910417Z digest=sha256:d6d1a3ed6ec7d47fb1bcf91dc06b4ce03545d9d4f7c8731d53b63b0f64445076

Observation cabb546d-af1a-46f3-8abe-4ff690e6b441 · inbound

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams cites this paper.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 92

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:56.151854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:a19a70919393452c6f37a4692637fa51b9042cbef8262a63483562222386320f