Pith. sign in

Paper Citation Record · LEDGER

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions

As of 8 August 2026, this Paper Citation Record lists 9 of 9 outbound references and 4 inbound Pith citation observations for arXiv:2602.08711.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.08711 v3

Coverage vector

measured 9 of 9 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T03:15:11.601779Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T10:02:03.266407Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-04T17:40:00.898211Z

Reference resolution

9 of 9 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4c776e3f-1171-4350-aae9-9bc360072212 · outbound

This paper cites • Each minute usually contains4–5 segments, but prioritize the video’s logic over strict numbers.

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions • Each minute usually contains4–5 segments, but prioritize the video’s logic over strict numbers

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:11.405612Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:11.405612Z digest=sha256:9523d164dd7996e6b10389b6abf3cd5dd20712d80bb438899272306080261fbf

Observation c7a12329-78df-4c02-a3f0-2b912ae77545 · outbound

This paper cites • Includecharacters, actions, objects, emotions, and scene detailswhere relevant.

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions • Includecharacters, actions, objects, emotions, and scene detailswhere relevant

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:11.419923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:11.419923Z digest=sha256:8dd11894690c78b08eb683fe960fcd843a34c8939eb2f3e619fd741275a39831

Observation ebf37857-5121-4598-a9ae-5277186fafd9 · outbound

This paper cites • Timestamp format:minutes:seconds(e.g.,0:00 - 0:11).

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions • Timestamp format:minutes:seconds(e.g.,0:00 - 0:11)

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:11.451377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:11.451377Z digest=sha256:ad2c8224d8922f877d10c45c16e45c5e8c42edf9ca26683169780b1eaef56995

Observation ed901d18-4d82-40c8-a04b-510d88aaa265 · outbound

This paper cites Each GT caption defines the rough boundaries of a segment.

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions Each GT caption defines the rough boundaries of a segment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:11.478148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:11.478148Z digest=sha256:7f142e58d4a25d2e8bb887a3305ecacb84d4061d77e1cf43c97816a89647749b

Observation c71922d9-2cfc-4bca-b50e-9b56dc80e979 · outbound

This paper cites Use the video itself toexpand with details: •Characters: actions, gestures, facial expressions, emotions.

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions Use the video itself toexpand with details: •Characters: actions, gestures, facial expressions, emotions

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:11.509777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:11.509777Z digest=sha256:6755eab30a4d9217803d089f03d15843a5c5eb5ddb3cab4aebe4aece1e4e4bed

Observation 163a77b1-b06a-42c1-9fe1-1480feb004d2 · outbound

This paper cites timestamp.

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions timestamp

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:11.541291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:11.541291Z digest=sha256:9cb529db30a284caa1f8b9c656ff072334cffab1b2b1fd6e72d2b02b366454ca

Observation a576c31c-d18f-463f-b3fd-044dad84c7e8 · outbound

This paper cites an unresolved cited work.

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:11.569122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:11.569122Z digest=sha256:b9ed093ac608efa6f9a11194a3ffb89cb536aa98e3b448e411c807d81d320bd2

Observation 8bbc8a29-9465-4f13-b4cf-9b98fa517a21 · outbound

This paper cites by_dim": {.

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions by_dim": {

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:11.601779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:11.601779Z digest=sha256:8b5b5e8ca35c2c90659e8de03edddc572a29992e0d52dfad788f716a8a087682

Observation 95218483-f9a3-4ebb-b2e0-62ea836414ee · outbound

This paper cites HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context.

TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T03:15:11.377213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:15:11.377213Z digest=sha256:6bdc7f94cef30eedbdbddbb57107c359f81735d237b002135cfefb04d3b0c625

Pith citing papers

Observation 52056ccb-249e-40b0-a81d-75779896f93b · inbound

OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video cites this paper.

OmniScript: Towards Audio-Visual Script Generation for Long-Form Cinematic Video TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-07-03T02:17:09.422974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:35:23.842691Z digest=sha256:ed8c6d055d569232d8f6ddd35a960ffa24c94ec62cad79aedd1c29228326deec

Observation 0a8d3301-f527-4e13-bd04-03d87c9e37d8 · inbound

DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning cites this paper.

DiffCap-Bench: A Comprehensive, Challenging, Robust Benchmark for Image Difference Captioning TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-07-03T02:17:09.422974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-08T17:56:47.621050Z digest=sha256:162d2372bf7380970cd7e498104a759fb47682db2705eb01f632d7e958d32247

Observation ae5e7b81-915d-4acf-ac1f-cc23263dbccd · inbound

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning cites this paper.

CineCap: Structured Reasoning with Spatio-Temporal Anchors for Cinematographic Video Captioning TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-04T17:40:00.899555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-25T23:29:24.520537Z digest=sha256:9f7c3cf498f47cbca875c51bf7b6a461974984348f0d3457003c359265175499

Observation 7deff422-86ba-4f42-a199-15919089ae58 · inbound

PercepCap: Video Captioner with Structured Spatio-Temporal Perception cites this paper.

PercepCap: Video Captioner with Structured Spatio-Temporal Perception TimeChat-Captioner: Scripting Multi-Scene Videos with Time-Aware and Structural Audio-Visual Captions

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T10:02:03.266407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:02:03.266407Z digest=sha256:7b8c72c42116377c25fc99e4f43e63711bbb8d1c8c434f3e8a0d7d2d9f2a04a5