Pith. sign in

Paper Citation Record · LEDGER

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting

As of 14 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 1 inbound Pith citation observation for arXiv:2507.16873.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.16873 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:17:58.243690Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T00:55:59.857014Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-20T00:57:54.247973Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact1
  • verified fuzzy40
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 111e67ee-b96a-4a46-9964-b9ba7fe484d4 · outbound

This paper cites GPT-4 Technical Report.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:17:57.694812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:17:57.694812Z digest=sha256:fdb7772d28649aef5a35b2ccec29a5b7b992a1553e31dd2308674fdbadf4a6f1

Observation 14b28a1e-c8e4-472a-b772-576670cb78ee · outbound

This paper cites Video summarization using deep neural networks: A survey.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Video summarization using deep neural networks: A survey

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:18:00.559165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:57.706308Z digest=sha256:24e323319711b0cdcf546bad50f07024aeea94696a75cef1a275cb6b543ba360

Observation 4f93a273-ade6-4024-b885-94c9a4835eff · outbound

This paper cites Towards automated movie trailer generation.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Towards automated movie trailer generation

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:18:00.526219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:57.715747Z digest=sha256:fae9fa85c95bdde33563c72a8182dd26f2d2c21f7ac6cc94a5ac6c3813d2022b

Observation f06a89c7-b8a4-444d-a8e3-b2d4a61b9262 · outbound

This paper cites Scaling up video summarization pretraining with large language models.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Scaling up video summarization pretraining with large language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:18:00.494069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:57.732501Z digest=sha256:ae79041a660c83ac2c290e48b975d9c5f78deb9e6a17cc5269f37efb2e9189b7

Observation 9d250bf7-c4b0-4f42-a71a-659e9ba1c1da · outbound

This paper cites Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability Curvature.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Fast-DetectGPT: Efficient Zero-Shot Detection of Machine-Generated Text via Conditional Probability Curvature

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:17:57.742908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:17:57.742908Z digest=sha256:cb262de95ea917f00e2056cf6fd671dff2145285195e1aabbeaaafffad705bd8

Observation c5e70f11-e328-4a11-9644-a2a5d94b2764 · outbound

This paper cites End-to-end object detection with transformers.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting End-to-end object detection with transformers

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:17:57.754625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:17:57.754625Z digest=sha256:3d5730f7310141407355be7d87238cda53e1fda68ab6df25927cb932b95b3d81

Observation 971a1022-1cae-47db-bb32-8a8d388802f9 · outbound

This paper cites Personalized video summarization by multimodal video understanding.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Personalized video summarization by multimodal video understanding

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:18:00.446030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:57.772859Z digest=sha256:d1c76948a4d6ca2d91804ea4e5b6ec9ec14cc8e739cdbd0fd2099288c8dd6bf9

Observation 2df36854-e189-4c3a-bb60-8d9fccec8895 · outbound

This paper cites Chatbot arena: An open platform for evaluating llms by human preference.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Chatbot arena: An open platform for evaluating llms by human preference

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T15:17:57.780042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:17:57.780042Z digest=sha256:c1e98462bb4c3b5bc58890f4ec1123ca14a6595bed87848efacbbfc842caee73

Observation 2f9c3627-2e56-4a5e-84c9-3cce8d9556ee · outbound

This paper cites Tall: Temporal activity localization via language query.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Tall: Temporal activity localization via language query

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:18:00.397974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:57.794841Z digest=sha256:1483058fa711a383362cbd76952658aa0a1de85997d713b765eaae1da7dc3638

Observation 6f3e2665-8d99-4161-b88d-a1bbf8e26051 · outbound

This paper cites Creating summaries from user videos.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Creating summaries from user videos

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:18:00.361373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:57.805901Z digest=sha256:42a6dfba3b07b2fef2138c25cbad07e681931c9345ddbbb66a53d717c3e65aae

Observation 2e82b1b4-5224-40cd-9210-3ccff6cd019a · outbound

This paper cites Video2gif: Automatic generation of animated gifs from video.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Video2gif: Automatic generation of animated gifs from video

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:18:00.314682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:57.813044Z digest=sha256:ddf72aa835f481129bf6bf1158ceef66bdd2325deb315d46a25564469dcbf138

Observation 6b864fc9-9403-4a05-a35b-e0262781c528 · outbound

This paper cites Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Shot2Story: A New Benchmark for Comprehensive Understanding of Multi-shot Videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T15:17:57.824025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:17:57.824025Z digest=sha256:4a8767c1f8b2a8d78044eccd45443135d323afa0c89c4d7506aaad3ceef8e183

Observation 5b936faf-0f4c-404c-995e-7bdbeb775617 · outbound

This paper cites V2xum-llm: Cross-modal video summarization with temporal prompt instruction tuning.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting V2xum-llm: Cross-modal video summarization with temporal prompt instruction tuning

Reference 13

Resolution
verified exact
raw_fallback, observed 2026-08-06T15:17:58.535719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:57.833243Z digest=sha256:77c3eb00d377476cfb652de4ce5047010a545503a386cbdd2b5eb425656b5762

Observation 178f61e5-5ca4-4ccb-bd6a-6b78c0c7cf7e · outbound

This paper cites Movienet: A holistic dataset for movie understanding.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Movienet: A holistic dataset for movie understanding

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:18:00.274641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:57.850713Z digest=sha256:53819eed0f88e6a096080d61236327482414279b653549c1757f811b679967fd

Observation 100bc01e-0ebb-4d22-a71c-1079ae8835d8 · outbound

This paper cites Video summarization with attention-based encoder--decoder networks.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Video summarization with attention-based encoder--decoder networks

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:18:00.233683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:57.866956Z digest=sha256:89d2e0fbcf03c72e2b2604f775d8a2aab9bc8ccfd6310b1c2f18b635a406fa73

Observation b5e92f8f-3ef2-4c14-a09d-c8b297af2f97 · outbound

This paper cites Mdetr-modulated detection for end-to-end multi-modal understanding.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Mdetr-modulated detection for end-to-end multi-modal understanding

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T15:17:57.878808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:17:57.878808Z digest=sha256:5896399d3d5c1a00575decc8c1080cd4dff92b7066970ab82aba46c054814721

Observation 78658a8a-e7bf-41e2-9f1d-5f2293280a20 · outbound

This paper cites Self-attentive sequential recommendation.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Self-attentive sequential recommendation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:17:57.889332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:17:57.889332Z digest=sha256:076c3a3bcf096e68046ab2a78daee0a0da60d4700920b8147161957570350db1

Observation 34c922f3-384c-43b9-822b-ac4b3f72e2e4 · outbound

This paper cites Tvr: A large-scale dataset for video-subtitle moment retrieval.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Tvr: A large-scale dataset for video-subtitle moment retrieval

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:18:00.099995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:57.894979Z digest=sha256:af3781b39d97a276ee4d2888325ac42f4da9fbdd83fee8a08351511725fb4f5f

Observation 92b3b6c7-adb3-4da3-8832-cbbe029447bf · outbound

This paper cites Detecting moments and highlights in videos via natural language queries.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Detecting moments and highlights in videos via natural language queries

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:17:57.900813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:17:57.900813Z digest=sha256:afb783769463fb454f36667602aa089109b26161de943a59b18ef801dd94e83e

Observation 1186ada7-6db8-4df6-9898-174df4cae153 · outbound

This paper cites HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:17:57.908906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:17:57.908906Z digest=sha256:c78a14c6e8eda8e8b5caf14212c022735196eddcbeec3d4eef342466b22ea35d

Observation 2fffdc20-b80c-4b42-ad59-a56ad77212c1 · outbound

This paper cites Univtg: Towards unified video-language temporal grounding.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Univtg: Towards unified video-language temporal grounding

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:18:00.005787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:57.915019Z digest=sha256:5f8083ed1d7e9545c9778a8d9ee8c09336d0c00586c67d37ef7b038efc1faa23

Observation d74afcc3-c1aa-4ae7-93cd-14992b383a2f · outbound

This paper cites Visual instruction tuning.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Visual instruction tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:17:57.923338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:17:57.923338Z digest=sha256:b008faeb2b8f45c55148441106ecb964b7b31cf9749edc295685c9dadddee394

Observation fbe27845-2ba9-4e8f-afa8-131644f5e416 · outbound

This paper cites Attentive moment retrieval in videos.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Attentive moment retrieval in videos

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:17:59.942980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:57.930475Z digest=sha256:2162ddce5322c1d4dbab222b3c6169b859d49b5d601f068505a25440e97f8063

Observation cb57efb1-20af-4de0-9b32-4c036c771603 · outbound

This paper cites Multi-task deep visual-semantic embedding for video thumbnail selection.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Multi-task deep visual-semantic embedding for video thumbnail selection

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:17:59.916236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:57.938843Z digest=sha256:41bb23cce6c0a9f47fd88b2090609ddd8bad5958fc5fabbe554ef0b9527b56b4

Observation 200f224b-e477-4969-b44a-eda34a757df2 · outbound

This paper cites Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Umt: Unified multi-modal transformers for joint video moment retrieval and highlight detection

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:17:59.894062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:57.950702Z digest=sha256:42e6bf6d5e9dbe047cdc173011b3eb990bc1cf7a75e0718c56b8c2eb5ff15ac6

Observation 67d4abdc-b013-481e-acd5-af011be4bf6e · outbound

This paper cites Debug: A dense bottom-up grounding approach for natural language video localization.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Debug: A dense bottom-up grounding approach for natural language video localization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:17:59.858277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:57.959213Z digest=sha256:25bd612ae586ee30f3900f9c227cb81c12890ef302234ef2871ec1820b8acabf

Observation 14265216-7375-4778-b9c6-da0ad1b99330 · outbound

This paper cites Videoautoarena: An automated arena for evaluating large multimodal models in video analysis through user simulation.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Videoautoarena: An automated arena for evaluating large multimodal models in video analysis through user simulation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:17:59.813663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:57.968763Z digest=sha256:56b28aa8ca5ea08cd46cd9c2be613e6e453f58a7401588130d737c506265c983

Observation 6370adc4-c383-4e9f-aa75-86741df5413c · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T15:17:57.978108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:17:57.978108Z digest=sha256:9accdc173369ada71cfc92da842de5135d3778053585d63d8eab95d11002a4a9

Observation 13ca2bd9-7318-45af-a15f-56f231f4b646 · outbound

This paper cites Detectgpt: Zero-shot machine-generated text detection using probability curvature.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Detectgpt: Zero-shot machine-generated text detection using probability curvature

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:17:59.747306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:57.986161Z digest=sha256:3fd92b3cea041df4ebccbd9caff7d60aacd1155cc6c900a8e996df51686dee26

Observation 12fcf8e1-6451-4820-ac18-dff6ed7fb08f · outbound

This paper cites Query-dependent video representation for moment retrieval and highlight detection.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Query-dependent video representation for moment retrieval and highlight detection

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:17:59.706354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:57.997352Z digest=sha256:081fd363568e43635d3b57fc1653d5ab030dfa5139f5c9a1fd25b84ae84c21ab

Observation cf66e6ea-cd8a-4ed0-b290-2fcea3507307 · outbound

This paper cites Clip-it! language-guided video summarization.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Clip-it! language-guided video summarization

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:17:59.673686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:58.005873Z digest=sha256:3b2b6d3a969be4cf53b8991315809dbea694bd6f713a433788495ceb3a937694

Observation 69e54f2b-58c6-4089-b36f-483ebc85c181 · outbound

This paper cites Sumgraph: Video summarization via recursive graph modeling.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Sumgraph: Video summarization via recursive graph modeling

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:17:59.632181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:58.011349Z digest=sha256:0fcee0eccc8601709493788deba1c81c396372a18c24435df34c7d8e396c9f88

Observation 0fe70fca-d3e8-468f-bda8-39f3cd0bbfde · outbound

This paper cites Mmsum: A dataset for multimodal summarization and thumbnail generation of videos.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Mmsum: A dataset for multimodal summarization and thumbnail generation of videos

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:17:59.587725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:58.017617Z digest=sha256:4d070fe5102040b6ef8dcb860972dc6de9b7a493d797dcebd37bb724e0ec1c93

Observation 612b5a05-1947-4417-b6d6-ed88046ad67c · outbound

This paper cites Learning transferable visual models from natural language supervision.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Learning transferable visual models from natural language supervision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T15:17:58.026255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:17:58.026255Z digest=sha256:94e28481051634648e0fffb1771d29bc92ac9459178522268d35d6185cbd9838

Observation 12a5c193-1394-4a55-9fd9-3744ba6d020d · outbound

This paper cites Bpr: Bayesian personalized ranking from implicit feedback.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Bpr: Bayesian personalized ranking from implicit feedback

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:17:59.508037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:58.034667Z digest=sha256:16fb880a15056da7a09b7745ae61eab244c6be2bce9fc96721ef71b15ad4008f

Observation fbca70f4-3f80-48e4-90c8-484a13ac1b35 · outbound

This paper cites Adaptive video highlight detection by learning from user history.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Adaptive video highlight detection by learning from user history

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:17:59.462359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:58.042738Z digest=sha256:555bf57ab8e573c9eae0b983cceba77c4e175b0e853414411f80dba2faafc917

Observation 778060dd-f944-4ab2-b973-c03a99ae4853 · outbound

This paper cites Query-focused extractive video summarization.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Query-focused extractive video summarization

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:17:59.415123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:58.048754Z digest=sha256:d22d59b3a07ab5e1e0b655ff36bc070435b78e78f09403ee5009b45a1ff9a8ca

Observation 068a8527-2123-4f8e-b6f3-a414943f5124 · outbound

This paper cites Query-focused video summarization: Dataset, evaluation, and a memory network based approach.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Query-focused video summarization: Dataset, evaluation, and a memory network based approach

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:17:59.379996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:58.054036Z digest=sha256:34955e477c73e648fe914ecf18fcb4600fbcdabc5fd963cb4cc73ec7d0202bec

Observation 15bbd598-720d-4cd9-87df-175cf218596e · outbound

This paper cites Tvsum: Summarizing web videos using titles.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Tvsum: Summarizing web videos using titles

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:17:59.342002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:58.061056Z digest=sha256:9f2d7ffdb9b79c84a8cb8d1e243c63e6024b0ca1b7609489f30a41f5ee267da9

Observation fcf77a06-11a7-4cb3-8e6f-e4e37d409487 · outbound

This paper cites To click or not to click: Automatic selection of beautiful thumbnails from videos.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting To click or not to click: Automatic selection of beautiful thumbnails from videos

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:17:59.312544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:58.067085Z digest=sha256:75b69c3e598cbc2e01c56aa3fea5773993c7e398ec83efb68f384189423dd018

Observation a6087fab-8d80-46c5-90a8-546fbb73ef42 · outbound

This paper cites an unresolved cited work.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:17:59.284673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:58.076964Z digest=sha256:b8d8fd69468a130d46b75d9ac8295f54bf843f9e4d513be64e2d4872145f5e98

Observation 1600e304-033b-4d22-a5a0-ccfea5bc1e24 · outbound

This paper cites Tr-detr: Task-reciprocal transformer for joint moment retrieval and highlight detection.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Tr-detr: Task-reciprocal transformer for joint moment retrieval and highlight detection

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:17:59.250454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:58.087219Z digest=sha256:59b41cda03e88997e047fcd1555de5c0e9bdd67d870f4399991430260539816d

Observation 4fedb22b-42da-4fab-a4e1-2b9bc923309d · outbound

This paper cites Ranking domain-specific highlights by analyzing edited videos.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Ranking domain-specific highlights by analyzing edited videos

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:17:59.215237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:58.099601Z digest=sha256:ffe03b65c3006401d1c55e5669fdea37970d25a16a1467cd92116e2f8e6c7b0b

Observation 87556897-de50-49d6-ba8a-993fb85164d3 · outbound

This paper cites Query-adaptive video summarization via quality-aware relevance estimation.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Query-adaptive video summarization via quality-aware relevance estimation

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:17:59.175888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:58.108373Z digest=sha256:321f1fdb39e68b5e48d8939276f0406750ab6efc96eb134eb026cc1422b8a52a

Observation 11843b59-aebc-42d7-a3f9-d711aec94c8c · outbound

This paper cites Videoagent: Long-form video understanding with large language model as agent.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Videoagent: Long-form video understanding with large language model as agent

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T15:17:58.113681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:17:58.113681Z digest=sha256:1ae2933c57d6f461731adf708d6cdc76772ffa3cda76e02266236fb98a0a73b6

Observation 5b6a1701-91b5-4a9e-a64b-4b4a2456f62c · outbound

This paper cites Query-biased self-attentive network for query-focused video summarization.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Query-biased self-attentive network for query-focused video summarization

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:17:59.096946Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:58.119440Z digest=sha256:d8929cecad6cd838ec102bde6be467e74d5167e1603c4c9849e877d2015be784

Observation 4b5ea4cb-bc56-48d0-b451-09a8d639da2e · outbound

This paper cites Convolutional hierarchical attention network for query-focused video summarization.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Convolutional hierarchical attention network for query-focused video summarization

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:17:59.053131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:58.126430Z digest=sha256:f2c765b12faee3eeb542f30b319e59c533574703d10ecd77254325e8de154656

Observation fdfec61e-997d-4624-8fe8-0aaf263c8ae3 · outbound

This paper cites Bridging the gap: A unified video comprehension framework for moment retrieval and highlight detection.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Bridging the gap: A unified video comprehension framework for moment retrieval and highlight detection

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:17:59.017479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:58.135930Z digest=sha256:39a51df9c6603c18d66585f7063681a6d259d58b783ee1aff7e9ca67b9bebbc0

Observation a63772b5-cfef-4cb1-8124-3312178cc852 · outbound

This paper cites Cross-category video highlight detection via set-based learning.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Cross-category video highlight detection via set-based learning

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:17:58.977655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:58.142535Z digest=sha256:5c25675ba6223a1e2b676b0607a7d03c5490661a88ab76a8a4db8a9b7d6fb4a7

Observation 150ced57-45b7-4bec-a253-c45bab977bf5 · outbound

This paper cites Mh-detr: Video moment and highlight detection with cross-modal transformer.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Mh-detr: Video moment and highlight detection with cross-modal transformer

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:17:58.936518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:58.147689Z digest=sha256:110f115d9d02dd3c2042959c934eefc1e5e1b2750accc6a34d383a6221e6baf2

Observation ed1d3f71-31c3-44a2-8ad1-67beec4c60e7 · outbound

This paper cites Highlight detection with pairwise deep ranking for first-person video summarization.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Highlight detection with pairwise deep ranking for first-person video summarization

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:17:58.909446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:58.153170Z digest=sha256:21db85a3912106f1a00895b34135485f351f5e4155af9f73cdf3a8def2c84cf8

Observation 1326983a-f5e9-476d-af92-e90ef741f5d4 · outbound

This paper cites Semantic conditioned dynamic modulation for temporal sentence grounding in videos.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Semantic conditioned dynamic modulation for temporal sentence grounding in videos

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:17:58.872839Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:58.158832Z digest=sha256:af50c15e1b7c098c8fd2b05eaa5e5ec53b3ecc2d57a2ab4212a4bf026796c9c2

Observation bf2b387c-6436-4f76-a3e4-5da413d9b1f2 · outbound

This paper cites Hierarchical video-moment retrieval and step-captioning.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Hierarchical video-moment retrieval and step-captioning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:17:58.836105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:58.165948Z digest=sha256:98bed05da347d04cc62c6b77009fe3c6fa4f29da7ea452b047b54205fdc30b56

Observation 04726dc6-2e68-4748-91af-cb19249712e3 · outbound

This paper cites Moment is important: Language-based video moment retrieval via adversarial learning.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Moment is important: Language-based video moment retrieval via adversarial learning

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:17:58.809922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:58.174479Z digest=sha256:b08f1291ac4e0d1a226026dceb07be3159aae7e10a1cddf02274f4a07b2e4f79

Observation 8d2e8ea1-3eb3-4cce-b704-39e7b0c783fe · outbound

This paper cites Span-based Localizing Network for Natural Language Video Localization.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Span-based Localizing Network for Natural Language Video Localization

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T15:17:58.184421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:17:58.184421Z digest=sha256:000c6cb1a0fedde277763d1db305db9883ca6b8afb8b8fd52f20129c536886df

Observation 3a048f91-13ce-4e15-b481-71f3410b713f · outbound

This paper cites Towards automatic learning of procedures from web instructional videos.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Towards automatic learning of procedures from web instructional videos

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:17:58.779679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-08-06T15:17:58.193137Z digest=sha256:d42daf897a12a85228a14ecb0c172fb5d6159194a58881c0d9804043cb35a072

Observation 8eb3311f-c2fb-43f5-899b-7f6807e76280 · outbound

This paper cites write newline.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting write newline

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T15:17:58.209470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:17:58.209470Z digest=sha256:e4d0e0f97f31732b1592044859bea51482db53ffa05af06f56a621536c64eb98

Observation a8d73be5-de15-45fd-bba3-b628a287f923 · outbound

This paper cites @esa (Ref.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting @esa (Ref

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T15:17:58.220624Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:17:58.220624Z digest=sha256:fc913ec22f29a9d62317d5eeeae7dbd3b14d709f9b0eb174dd3293dc0fa135ac

Observation 221f98c4-a519-4b45-af4a-53619fe51ada · outbound

This paper cites an unresolved cited work.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T15:17:58.235989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:17:58.235989Z digest=sha256:6c84d040a8ca137616a438a0f5b506079d8d55e865cc33b439ab9589d1539903

Observation ffbfffb5-013f-445d-90c2-45a901683c77 · outbound

This paper cites an unresolved cited work.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T15:17:58.243690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:17:58.243690Z digest=sha256:f4bc0af50568857d27057826656f79767a31eb6997cf379ebf5d85c81c827a2e

Pith citing papers

Observation 07909adf-c8b8-4550-8409-6fd6794033a1 · inbound

Will It Go Viral? Grounding Micro-Video Popularity Prediction on the Open Web cites this paper.

Will It Go Viral? Grounding Micro-Video Popularity Prediction on the Open Web HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-20T00:57:54.249989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-20T00:55:59.857014Z digest=sha256:229873a6881389d2340de94602afdb4926f6c0b17a62ce66981d4357b02b8064