Pith. sign in

Paper Citation Record · LEDGER

Unsupervised Transcript-assisted Video Summarization and Highlight Detection

As of 10 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2505.23268.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.23268 v1

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:54:19.856159Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy22
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 255ca5ff-7fb4-4850-be8d-db87a2bb0e1a · outbound

This paper cites Align and attend: Multimodal summarization with dual contrastiv e losses,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection Align and attend: Multimodal summarization with dual contrastiv e losses,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:25.528941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:54:18.096395Z digest=sha256:e6594a794865581193adebb39720f12d67a79955526136e2e44966a2d8d22608

Observation 6e930162-5b9e-46fd-bbd3-55ae43317b2f · outbound

This paper cites Clover: T owards a unified video-language alignment and fusion model,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection Clover: T owards a unified video-language alignment and fusion model,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:25.313406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:54:18.203435Z digest=sha256:fd0a6fdd2047d2eebc3708cc12be99bb1be71203598c94b7ba0c60a7e791d816

Observation 37a123c6-8d8f-45e2-9e3d-cf28bb2b2ff3 · outbound

This paper cites SUSiNet: See, understand and summarize it,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection SUSiNet: See, understand and summarize it,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:25.093946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:54:18.296863Z digest=sha256:7b052ca2061ed18e1b874d28e18588d851e798178d9f45151478911bd77ba158

Observation b95352db-86b7-431a-b4c9-d784454a8592 · outbound

This paper cites UMT: Un ified multi-modal transformers for joint video moment retrieval and highlight detection,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection UMT: Un ified multi-modal transformers for joint video moment retrieval and highlight detection,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:24.873048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:54:18.391846Z digest=sha256:03f526ae882980694a2accefeb0e19bbbb91a2454f4faaee6c0f18508ac143ab

Observation b18d4533-263d-4070-8ab2-c6c5b36c6e65 · outbound

This paper cites CLIP-it! la nguage-guided video summarization,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection CLIP-it! la nguage-guided video summarization,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:24.575792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:54:18.485127Z digest=sha256:cc1af1192ef873a883db018cb6dda251957dfd79e6ce4e417ce9e2a61a0fafbc

Observation eed8a926-2057-4442-8544-0f509613b852 · outbound

This paper cites Creating summaries from user videos,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection Creating summaries from user videos,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:24.373861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:54:18.562836Z digest=sha256:bd901f132a827e5d08d9c21211a8fb22781bdae14e28ab7500630c9e1806d29e

Observation 45204a32-4423-49fc-9ad1-098b3bba9675 · outbound

This paper cites TVSum: Summarizing web videos using titles,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection TVSum: Summarizing web videos using titles,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:24.135808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:54:18.624966Z digest=sha256:8d3a80141fbdd3683c76e3c19e7144a6dc57e1ad7f5e7187a820a5e18e3e362b

Observation e7d54eac-11c2-4a59-a3f9-22799e0ee230 · outbound

This paper cites Deep reinforcement learning for uns upervised video summarization with diversity-representativeness reward,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection Deep reinforcement learning for uns upervised video summarization with diversity-representativeness reward,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:23.948873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:54:18.696700Z digest=sha256:286234669e017c1ffb47cfe7d9f19cb0e7778f73683645fb77165115b12b3fd6

Observation b50f2ff4-c28b-4b04-9487-0885bc7b84bd · outbound

This paper cites Summarizing videos using concentrated attention and considering the un iqueness and diversity of the video frames,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection Summarizing videos using concentrated attention and considering the un iqueness and diversity of the video frames,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:23.704468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:54:18.750814Z digest=sha256:d5e93f1a1e5650481b56456b05b06157396b9b01158a91ae0e4a5728e70a0345

Observation e77ff443-b4d3-477e-aa60-56085cc752ce · outbound

This paper cites TL;DW? Summarizing Instructio nal Videos with Task Relevance and Cross-Modal Saliency,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection TL;DW? Summarizing Instructio nal Videos with Task Relevance and Cross-Modal Saliency,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:23.487700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:54:18.814953Z digest=sha256:7ebd4c67415cc73f4d1dca78fd771066cfdb3634ff84eb48d96c67202090bc5f

Observation 55f3bec8-60de-408a-881d-0d810d202a93 · outbound

This paper cites Rethinkin g spatiotem- poral feature learning: Speed-accuracy trade-offs in vide o classification,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection Rethinkin g spatiotem- poral feature learning: Speed-accuracy trade-offs in vide o classification,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:23.276417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:54:18.880999Z digest=sha256:625dc2141de45381d24a881648bf1a9e4a1661419718d92e0a407592c5fabacb

Observation 9d8e2308-f7ff-452b-bd09-0dece83de2df · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:18.964559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:18.964559Z digest=sha256:008e3e2e1099ad29c6dbd38606062013281e78f8f8f4881c559dc191876542bb

Observation 8f372dae-3bed-4e7b-80a0-5702f1a0bce0 · outbound

This paper cites AREDSUM: Adaptive redundancy-aware iterative sentence ranking for extracti ve document summarization,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection AREDSUM: Adaptive redundancy-aware iterative sentence ranking for extracti ve document summarization,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:23.095873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:54:19.039626Z digest=sha256:721b6bb36a17059f165585ae968d772a6f48496372e8721f14a4c9e2c921019d

Observation cff0ff2b-9299-4c67-88a8-4af7ca625b7f · outbound

This paper cites Attention is all you need,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection Attention is all you need,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:22.903832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:54:19.086216Z digest=sha256:577f9dd0835336eef85b386b4f87c1c165a4b9a66523af874b9ffaeff770207e

Observation 701e8788-8a6e-4869-bdc3-bb00788978c6 · outbound

This paper cites VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection VLMo: Unified Vision-Language Pre-Training with Mixture-of-Modality-Experts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:19.140326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:19.140326Z digest=sha256:7aa43afd5c2ada4c247b0e7d15c16d9d0f332481d2037f5c50ef651d0205d0f6

Observation 5254dac2-47db-4fad-a325-79a852aafcc6 · outbound

This paper cites Simple statistical gradient-followi ng algorithms for connectionist reinforcement learning,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection Simple statistical gradient-followi ng algorithms for connectionist reinforcement learning,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:22.639669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:54:19.191277Z digest=sha256:180e9fe16ea4746694924ac7faafdd1cb28c2c7855122ff05d5fd55dcfed0bf1

Observation a98664f6-bc7e-442c-ae1b-07415d1cbf90 · outbound

This paper cites Adam: A method for stochastic opt imization,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection Adam: A method for stochastic opt imization,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:19.241770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:19.241770Z digest=sha256:263d877bfc1f60e8f142565adffa1486aa5f55fe77c2a010b423fe654f8f8b8c

Observation 58d02cb6-099b-4849-8583-bff9ff7666aa · outbound

This paper cites Abstractive text summarization using sequence-to-seque nce RNNs and beyond,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection Abstractive text summarization using sequence-to-seque nce RNNs and beyond,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:22.405014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:54:19.289591Z digest=sha256:23fe57782b0fa5f926b2fdc3a997d45408011b93e2975e7f474d60ba81f15fe3

Observation 61934291-2a83-4dc6-9eb2-afc055363ec3 · outbound

This paper cites HowTo100M: Learning a text-video embedding by watching hu ndred million narrated video clips,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection HowTo100M: Learning a text-video embedding by watching hu ndred million narrated video clips,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:22.235780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:54:19.335478Z digest=sha256:d31fc07dac8a96e26d6dc370504fa054081a31db200d80dbf1e7e6070c80cb13

Observation a23ffb87-3931-4893-8939-5b73da35ad07 · outbound

This paper cites Mr. HiSum: A large-scale data set for video highlight detection and summarization,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection Mr. HiSum: A large-scale data set for video highlight detection and summarization,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:22.019824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:54:19.400921Z digest=sha256:3424bc211965fdaa93dfa430b3cad7c09acbf131df7d7a32304467933ebd105b

Observation 75c4784f-b26e-4d75-83ae-6f872e9962c4 · outbound

This paper cites Reth inking the evaluation of video summaries,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection Reth inking the evaluation of video summaries,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:21.831846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:54:19.459604Z digest=sha256:28fb6e3337cb3956596e8334d6d5b104d7e953c88eb8acb5551e79bdc0a5aa10

Observation 72018b24-f092-453f-bdfc-639cb90eaf01 · outbound

This paper cites Zwillinger and S.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection Zwillinger and S

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:21.598933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:54:19.537886Z digest=sha256:ff36336532d40b07ce685cc182a6b4bd5d0d509e8cf5475b1be2bb134dd5b97c

Observation e035518a-dc10-473d-bac5-0deeccae4290 · outbound

This paper cites The treatment of ties in ranking problem s,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection The treatment of ties in ranking problem s,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:21.180234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:54:19.632314Z digest=sha256:1bfd436b0e81b2e727c5636977903549376f019ff3aae11ffb14cfd2473f4438

Observation 245fb392-2849-4c05-a7e9-de4e43be91e6 · outbound

This paper cites Cate gory-specific video summarization,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection Cate gory-specific video summarization,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:20.866198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:54:19.724624Z digest=sha256:2562582c28fe774010a1f2a125cb8100bc64ce3886a823003cce1c8ab6f5fccc

Observation eb81ee09-0ff3-4f1d-b697-77588af1849d · outbound

This paper cites WikiHow: A Large Scale Text Summarization Dataset.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection WikiHow: A Large Scale Text Summarization Dataset

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T12:54:19.775314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:54:19.775314Z digest=sha256:440067b197f089801a40706f074ab59d7e8643be93fd02afb379d55d939583e6

Observation 1b59550b-3d56-4209-82fb-6d6801571aac · outbound

This paper cites COG N- IMUSE: a multimodal video database annotated with saliency , events, semantics and emotion with application to summarization,.

Unsupervised Transcript-assisted Video Summarization and Highlight Detection COG N- IMUSE: a multimodal video database annotated with saliency , events, semantics and emotion with application to summarization,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T12:54:20.479050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T12:54:19.856159Z digest=sha256:ae9af2afe12213e70833cc9a9ed3f42f079166a2435131ac10010e5068a53ea4

Pith citing papers

No inbound Pith citation observations are available.