Pith. sign in

Paper Citation Record · LEDGER

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization

As of 9 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 0 inbound Pith citation observations for arXiv:2506.23714.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.23714 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:38:45.547295Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

36 of 36 outbound references displayed

  • verified exact2
  • verified fuzzy16
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1c8d0ab6-eabf-45ec-ba85-c2c154aa2a93 · outbound

This paper cites Apostolidis, E.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Apostolidis, E

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:51.946959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:41.726899Z digest=sha256:ffe2b02a07a955254ad89ee8845e949c90f4a026a2b2e68c0ab7f76574e2ef19

Observation 2eb97bf5-b907-4c40-9252-8890bcc32079 · outbound

This paper cites an unresolved cited work.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:38:51.675229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:41.843334Z digest=sha256:515aed08c75dd70f67e0a055cebcb52885a1068f20da8e738f1d3902bd1c6ff4

Observation a9d4485a-7722-4d00-85f7-da7d4868509a · outbound

This paper cites Tiwari, C.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Tiwari, C

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:51.479100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:41.927525Z digest=sha256:44316259243eed499ea86fa1f4e75dddab4d406529bd7da5f1cf89b5115be4bd

Observation 381d26d5-de0e-4068-917a-b53ec001a0dc · outbound

This paper cites Otani, Y.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Otani, Y

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:51.283393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:42.019442Z digest=sha256:14d3bd5e13464f857644d548c1e07b4660f2b1d57da6a57d6957e89486a1d6e3

Observation cbc56d9b-5ee8-48fc-8314-4b61e76cd7bd · outbound

This paper cites Rochan, L.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Rochan, L

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:51.086223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:42.083343Z digest=sha256:16199a2f9de74b3e07098b807b7a39f8d9c935e6384f0b5545692d022f5f3f5e

Observation 961b9393-b0aa-4f4b-9f6b-1dc8cdf10302 · outbound

This paper cites an unresolved cited work.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:38:50.835567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:42.160765Z digest=sha256:5ec385b7e88ea679a494fb7392473cab1e89dc4a92b692ab47fef5f8c1bab798

Observation 0c249360-1a60-4e4d-8517-c8cd63ef52e5 · outbound

This paper cites an unresolved cited work.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:38:50.613929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:42.271990Z digest=sha256:5a6dbf00ef2f2a1bc61109a5857c9f88346d7c5e342357559ce33d3911eaac72

Observation f5e1c958-ed12-4208-b7d2-94f2b619d2b5 · outbound

This paper cites an unresolved cited work.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Unresolved cited work

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T21:38:42.386654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:38:42.386654Z digest=sha256:fd1d33526a306b925ac0e975a3528838f7144e6ab0e0744a82ea4cf84d400890

Observation 2c3d545d-d505-4b1e-9333-520bc826e5bb · outbound

This paper cites Get To The Point: Summarization with Pointer-Generator Networks.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Get To The Point: Summarization with Pointer-Generator Networks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T21:38:42.501169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:38:42.501169Z digest=sha256:51abd44b9e5ae2bd334809dca3cc6668f70cd1e0894c3eebb3c28ceca411169e

Observation 0d9b036c-f9b5-483c-886a-d230ec05a78b · outbound

This paper cites an unresolved cited work.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:38:50.445452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:42.593261Z digest=sha256:359ba37dfa3bdf8a58dddeea291f4808bba237261915d15e8bdfe4d7cd5ee0ad

Observation a207b5fe-3468-416a-982f-cee44c9e27ea · outbound

This paper cites an unresolved cited work.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:38:50.194173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:42.792890Z digest=sha256:feef87e88c9d94c6c481afca24d571e9646031e368a659fba2273aa5126ca904

Observation a72f83fe-2201-4393-bfcd-80fed780a55d · outbound

This paper cites an unresolved cited work.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:38:50.036229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:42.831226Z digest=sha256:0b5d192529de92cdd1943e1c8abb6effc424721095ea240b5430f4236e903934

Observation c798f581-c284-4204-adee-12e3d0c0733a · outbound

This paper cites an unresolved cited work.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:38:49.782366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:42.910882Z digest=sha256:e1b0d20f1958c42102dcc7669ca04d0737802d31e01adca6a3c85ef77a1bedef

Observation 5a0af3e5-a26c-4430-9d5a-89f6a6ffd8f0 · outbound

This paper cites Zhang, W.-L.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Zhang, W.-L

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:49.528283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:43.006112Z digest=sha256:eaeb5daa31004ceb5e52c0d73e6bcf2d9c8558f61742cdbed2550ccfc520a504

Observation 62b27317-ba1b-4d54-8481-cb8ebca2c983 · outbound

This paper cites Saini, K.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Saini, K

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:49.287084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:43.118186Z digest=sha256:a8d19bc5f209bc4f788e21ca530b63e405275fc7d4580e586f021e2e251d69ab

Observation 886df382-a896-4c47-9f4c-2b180d455ba8 · outbound

This paper cites Evangelopoulos, A.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Evangelopoulos, A

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:49.078689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:43.213503Z digest=sha256:4f3c14c9dc8964fca6496b7760196ce2a84546db312122b540f5e1902b5873ae

Observation fb2656d3-fa37-472f-b2af-f9f51d9aa4df · outbound

This paper cites Multimodal Frame-Scoring Transformer for Video Summarization.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Multimodal Frame-Scoring Transformer for Video Summarization

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:38:45.965102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:43.276662Z digest=sha256:a48d1e996df1d62e026bffcc801a0d77c3c8dafb25fee055f76c0393b76e2658

Observation d39ed602-c3d9-42ff-bcc5-f0bda509258b · outbound

This paper cites an unresolved cited work.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:38:48.856571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:43.444350Z digest=sha256:bf1ebc71f098e1c3e561f9706751778c0d83b4b9af3387c96afc68f60c2cb03e

Observation b1503f76-54af-49d9-b4d3-70b0802bf9d5 · outbound

This paper cites an unresolved cited work.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:38:48.607542Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:43.569978Z digest=sha256:0f375994d055c39e4d0b55cd75143e028413cc14d6e9778450f58cb4825cf5f0

Observation a9f653a7-cd91-4f50-9d39-55e416539141 · outbound

This paper cites Psallidas, P.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Psallidas, P

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:48.323566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:43.730639Z digest=sha256:c5f77881e60972025603f0b932822014fe67ce54ded78701f778a6dbf3fd1c2c

Observation cbc3dd90-8b34-4a14-ac97-60f9b3d3be1a · outbound

This paper cites an unresolved cited work.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:38:48.174758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:43.802627Z digest=sha256:4e670aedef479fa26d90ca564d3c5a9b0b0a8b8ef6d93adfdc2f02fdade6a3a7

Observation 144dfa5c-f4b8-4c54-a3bf-1e44247daee2 · outbound

This paper cites an unresolved cited work.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:38:48.054461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:43.934293Z digest=sha256:d983a4ed6d08f38b76165fe9d936d37f33213c90eff05ca45416ed4504a40744

Observation 8c7936bb-aefd-439f-97de-1c0c8bb55e04 · outbound

This paper cites an unresolved cited work.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:38:47.918043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:44.117750Z digest=sha256:f8db17c56bdec211ab354c532d3e98629dc5bb88d28880f98876564026b0e373

Observation e6bc727e-9258-4621-9e86-eb581face9c4 · outbound

This paper cites Ponce-López, B.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Ponce-López, B

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:47.765723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:44.274131Z digest=sha256:f6b9eb940e9933ec790d7762a9fc158ea9f9dea5451eb24012fd3814b0b36dba

Observation 454caf69-4c5d-4748-8eef-257f5696f746 · outbound

This paper cites Robust Speech Recognition via Large-Scale Weak Supervision.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Robust Speech Recognition via Large-Scale Weak Supervision

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T21:38:44.378445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:38:44.378445Z digest=sha256:da476e99771b8fdf029d70cabd16b1be3d2b9d47e0a9859ad92fefd92648577d

Observation 40129d1c-c0d4-4d33-a40f-6d5796b06ee4 · outbound

This paper cites Honnibal, I.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Honnibal, I

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:47.609541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:44.467892Z digest=sha256:1e8d4a6cae7da9e56a385003bb75bd37888bf65466498b1c4fd49ec475309715

Observation 15ad8d3d-7527-4caa-b676-3655989e9ea0 · outbound

This paper cites McAuliffe, M.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization McAuliffe, M

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:47.442804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:44.581457Z digest=sha256:577d05ec1bca932ebc2e02196940a0850479f071c5f9d7cb1d8d670903cfc2ca

Observation 2b019d12-fe56-4340-aa0d-f2d290cdfe2a · outbound

This paper cites Bradski, The opencv library., Dr.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Bradski, The opencv library., Dr

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:47.182495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:44.656403Z digest=sha256:3f44412bf39c90366f751b60a6faabe9b5f7cc60966d28ed7cb2ca919a3af13f

Observation bdce2f8c-fb21-465d-a7e6-e7799bf975fc · outbound

This paper cites MediaPipe: A Framework for Building Perception Pipelines.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization MediaPipe: A Framework for Building Perception Pipelines

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T21:38:44.780495Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:38:44.780495Z digest=sha256:52ec2755dd1fe24acf7ba10e3e1bca2b7abfe32106196adc55fd3556463e1dfb

Observation 5c5b0eb6-fc0a-4835-9f06-09b9bbe92be0 · outbound

This paper cites Serengil, A.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Serengil, A

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:38:44.868440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:38:44.868440Z digest=sha256:9873a9ccdbbf3929cc528320cae7473121798096e6a3cc1d8f6e1856cce8b1e2

Observation 5bc15467-43d0-45ed-b8e9-6b5c2583ff94 · outbound

This paper cites Kasi, Yet another algorithm for pitch tracking (yaapt), 2002.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Kasi, Yet another algorithm for pitch tracking (yaapt), 2002

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:46.996042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:45.019358Z digest=sha256:ba37c972a2e45a3965c4dd7925c45206acd962a64dd26a9e8d48f52b08368f2f

Observation 65596910-48c2-42f9-aad3-7f074369f7c3 · outbound

This paper cites Eyben, M.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Eyben, M

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:46.771570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:45.121470Z digest=sha256:4a39cfe37f5fa814df58e9a29ebfded8442cefb53a2da6cf784e08239112ae27

Observation a4398015-e2c8-4dc7-b105-5bb6393f2704 · outbound

This paper cites Sparck Jones, A statistical interpretation of term specificity and its application in retrieval, Journal of documentation 28 (1972) 11–21.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Sparck Jones, A statistical interpretation of term specificity and its application in retrieval, Journal of documentation 28 (1972) 11–21

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:46.554911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:45.264167Z digest=sha256:2cd7abb5fd9e7c7b4f3ea41356af5fa9da6348901e5e95cc316ec37c5395d22a

Observation 8296f4a2-12d4-42c6-89b9-ca931e52cdeb · outbound

This paper cites an unresolved cited work.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:38:46.397336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:45.362663Z digest=sha256:da5892fa2b90fbc9805fbf986a2f9b514b7306e6ed618c3e4a07b32542d8b3e0

Observation a69251c5-26de-4f6c-b9ca-fa59d31aa6b9 · outbound

This paper cites Scaling Up Video Summarization Pretraining with Large Language Models.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Scaling Up Video Summarization Pretraining with Large Language Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:38:45.783605Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:45.460292Z digest=sha256:2825cb581b6abae98dcb1c83ef664fe37d8810a92348fd3c9001dbd1dea31cf4

Observation d4c6b550-d87d-4990-bc63-c22d7e5f54d0 · outbound

This paper cites Narasimhan, A.

Towards an Automated Multimodal Approach for Video Summarization: Building a Bridge Between Text, Audio and Facial Cue-Based Summarization Narasimhan, A

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:38:46.223123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T21:38:45.547295Z digest=sha256:fe885758a90b8a85897c1b3f67c43a4531077bddc0e47ad0133d3d37e22d3345

Pith citing papers

No inbound Pith citation observations are available.