Pith. sign in

Paper Citation Record · LEDGER

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions

As of 9 August 2026, this Paper Citation Record lists 9 of 9 outbound references and 4 inbound Pith citation observations for arXiv:2602.08711.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.08711 v3

Coverage vector

measured 9 of 9 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T03:15:11.601779Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T10:02:03.266407Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T17:40:00.898211Z

Reference resolution

9 of 9 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4c776e3f-1171-4350-aae9-9bc360072212 · outbound

This paper cites • Each minute usually contains4–5 segments, but prioritize the video’s logic over strict numbers.

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions • Each minute usually contains4–5 segments, but prioritize the video’s logic over strict numbers

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:11.405612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:11.405612Z digest=sha256:44bd0bac5ab96d4826d27a11283f86172b2b8690129f40a3993cede8a6bf3263

Observation c7a12329-78df-4c02-a3f0-2b912ae77545 · outbound

This paper cites • Includecharacters, actions, objects, emotions, and scene detailswhere relevant.

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions • Includecharacters, actions, objects, emotions, and scene detailswhere relevant

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:11.419923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:11.419923Z digest=sha256:bc9659a8acebb6e2f675126c673c40cbbfeef6aa7393923c96d936a0c8d4c1b0

Observation ebf37857-5121-4598-a9ae-5277186fafd9 · outbound

This paper cites • Timestamp format:minutes:seconds(e.g.,0:00 - 0:11).

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions • Timestamp format:minutes:seconds(e.g.,0:00 - 0:11)

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:11.451377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:11.451377Z digest=sha256:f95fcee6cec63f9393bfb398995c0796dcaff78719f69f9965f604c592c65992

Observation ed901d18-4d82-40c8-a04b-510d88aaa265 · outbound

This paper cites Each GT caption defines the rough boundaries of a segment.

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions Each GT caption defines the rough boundaries of a segment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:11.478148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:11.478148Z digest=sha256:1de0f88b08167d323cdb1ef85b55175d144d256e941b63fbbac52dc031dd1d70

Observation c71922d9-2cfc-4bca-b50e-9b56dc80e979 · outbound

This paper cites Use the video itself toexpand with details: •Characters: actions, gestures, facial expressions, emotions.

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions Use the video itself toexpand with details: •Characters: actions, gestures, facial expressions, emotions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:11.509777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:11.509777Z digest=sha256:b86b1ad8921dbbf7df00c357f8d92606953fa81b96d0b1c7ec34e37e7a33cf67

Observation 163a77b1-b06a-42c1-9fe1-1480feb004d2 · outbound

This paper cites timestamp.

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions timestamp

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:11.541291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:11.541291Z digest=sha256:17ac0aa9dc9a43340adc22bb6c1104585dc87fe5cdcb036072a374dd09843918

Observation a576c31c-d18f-463f-b3fd-044dad84c7e8 · outbound

This paper cites an unresolved cited work.

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:11.569122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:11.569122Z digest=sha256:eba5e4980a662cd328b7be41d98b04d72421c99034dd1b6428c51ee73f703756

Observation 8bbc8a29-9465-4f13-b4cf-9b98fa517a21 · outbound

This paper cites by_dim": {.

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions by_dim": {

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:11.601779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:11.601779Z digest=sha256:cc989870239e3d61a9cce80907c4246df8e9e2deb163fde642ef75046a909a0c

Observation 95218483-f9a3-4ebb-b2e0-62ea836414ee · outbound

This paper cites HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context.

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:11.377213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:11.377213Z digest=sha256:61a877af242a2a144b585505a9a64c680b7f335bd99617d349d59ac265604514

Pith citing papers

Observation 52056ccb-249e-40b0-a81d-75779896f93b · inbound

OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video cites this paper.

OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-03T02:17:09.422974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T15:35:23.842691Z digest=sha256:60ba0d0a6e30428b493c55ffa7519f1734b21ab046ace23b5b8aaa362dd8bbcc

Observation 0a8d3301-f527-4e13-bd04-03d87c9e37d8 · inbound

DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning cites this paper.

DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-03T02:17:09.422974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T17:56:47.621050Z digest=sha256:c13d1a92aca1df37be3f05cc63fc0842842c0636ed4b0f825712f19bec09aa88

Observation ae5e7b81-915d-4acf-ac1f-cc23263dbccd · inbound

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning cites this paper.

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-04T17:40:00.899555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-25T23:29:24.520537Z digest=sha256:d67fc33e83618bee3ae02f3506158361b206b92f72898b468f081285d74ad25b

Observation 7deff422-86ba-4f42-a199-15919089ae58 · inbound

PercepCap: Video Captioner with Structured Spatio-Temporal Perception cites this paper.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.266407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.266407Z digest=sha256:f6f61f9c39d748bcb40802947651a6b8682c56c7508785588125a95d3b49ea74