Pith. sign in

Paper Citation Record · LEDGER

HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

As of 21 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2005.00200.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2005.00200 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T12:19:36.470416Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:38:56.173895Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation f4092bb5-b83a-4515-add3-d7bca8a114e5 · inbound

Flamingo: a Visual Language Model for Few-Shot Learning cites this paper.

Flamingo: a Visual Language Model for Few-Shot Learning HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-12T04:22:30.162304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T04:22:30.008355Z digest=sha256:b47a62e0e92fbd13d6c947a6939114ba84192d38cae08811f217a1ff0f26493a

Observation 1e4841d5-236e-42b0-a3c7-faf4e1082a84 · inbound

LVBench: An Extreme Long Video Understanding Benchmark cites this paper.

LVBench: An Extreme Long Video Understanding Benchmark HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-19T11:55:30.191460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T11:55:30.048525Z digest=sha256:b68540785f0b5f627ce4c175bbb4069759b2965a0f639752997a0fee79531da0

Observation e12795ea-5475-42ed-b2fc-9ccfdb18bf82 · inbound

Do Language Models Understand Time? cites this paper.

Do Language Models Understand Time? HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-11T12:47:17.503698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:47:17.503698Z digest=sha256:9eea4ea884671db656b07de63f2274d39a0ff32936f7b4073fa68fa7e7e072b9

Observation 6f3045a2-0cff-4fa4-a05e-d11e361ffebb · inbound

Efficient Vision Language Model Fine-tuning for Text-based Person Anomaly Search cites this paper.

Efficient Vision Language Model Fine-tuning for Text-based Person Anomaly Search HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-09T05:30:25.474010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T05:30:25.474010Z digest=sha256:11ba32e8e7ac2c10b8bc62d68df63e7d7636da4fca2bda6ff6c6863277f2510f

Observation d4736d93-9ddd-4276-a815-ba364627e2bb · inbound

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions cites this paper.

Stitch-a-Demo: Video Demonstrations from Multistep Descriptions HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-23T00:32:18.256615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-23T00:30:55.729900Z digest=sha256:d5e6e4f36179420ada9214bbde11fd80dc3931bf15a47c399fd64a405854f417

Observation 43398dab-a72d-4441-b1e2-55da1f3df5d2 · inbound

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding cites this paper.

PerceptionLM: Open-Access Data and Models for Detailed Visual Understanding HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T12:19:36.470416Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T12:19:36.470416Z digest=sha256:8165faab8842e126ec753fd722427fa7ac6d30c105021407eb4323ff2d72c126

Observation d47c1bce-e1cd-4041-b537-6e78ce544380 · inbound

HierSum: A Global and Local Attention Mechanism for Video Summarization cites this paper.

HierSum: A Global and Local Attention Mechanism for Video Summarization HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T10:16:14.317594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:16:14.317594Z digest=sha256:bf17b0a95eb2cca5dd95b02806b275e8be14a8f0ca1b5e7291051826a379bf6f

Observation 365af228-59ce-46e0-ae1c-7e91497c43c6 · inbound

GeoMM: On Geodesic Perspective for Multi-modal Learning cites this paper.

GeoMM: On Geodesic Perspective for Multi-modal Learning HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.611677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.611677Z digest=sha256:b54c9a777da6083a26a7df5360f1f82f6ff412733c461ce70daddb8d9dca6ed5

Observation da815215-73cc-48dd-b574-c1ff35b1a32b · inbound

Generalizing vision-language models to novel domains: A comprehensive survey cites this paper.

Generalizing vision-language models to novel domains: A comprehensive survey HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 241

Resolution
unresolved
no resolver link, observed 2026-08-06T23:21:01.327799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:21:01.327799Z digest=sha256:f58629de36120540713adb2c3c810c66c690b284130c2539064671e5d2d1f338

Observation d9e9770b-45bc-4ef1-95c7-74c14caa103e · inbound

THYME: Temporal Hierarchical-Cyclic Interactivity Modeling for Video Scene Graphs in Aerial Footage cites this paper.

THYME: Temporal Hierarchical-Cyclic Interactivity Modeling for Video Scene Graphs in Aerial Footage HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:07:16.832602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:07:16.832602Z digest=sha256:17e32e97a9d73981925389902360414142ae6e8c438e3fc45b376414088f302c

Observation 0ca66b65-6b30-4633-b844-1cc8f35be18f · inbound

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition cites this paper.

From Coarse to Nuanced: Cross-Modal Alignment of Fine-Grained Linguistic Cues and Visual Salient Regions for Dynamic Emotion Recognition HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T17:08:54.214011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:08:54.214011Z digest=sha256:2c3d2039d07b8b0f06e158a50ca6a8e5404200fe884825facdedc353575302d2

Observation 1186ada7-6db8-4df6-9898-174df4cae153 · inbound

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting cites this paper.

HIPPO-Video: Simulating Watch Histories with Large Language Models for Personalized Video Highlighting HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:17:57.908906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T15:17:57.908906Z digest=sha256:8e77847291d0007fae380f7f595264ab1f02cf01a47f5290fb2151bdaf2459e6

Observation 3254beaf-9463-4b14-8c2b-92733863bfeb · inbound

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding cites this paper.

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T13:01:58.914260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T13:01:58.914260Z digest=sha256:5940bd8908f814c19b81c7bd61cf24c99d5a3f7184abc545dd9271af9b03e652

Observation 12cd106b-3b83-45b6-ab5a-bb008a2f4481 · inbound

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding cites this paper.

NarrativeTrack: Evaluating Entity-Centric Reasoning for Narrative Understanding HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-04T06:36:33.848152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T06:36:33.848152Z digest=sha256:14de4c6eb0a1d2d08477b6ebbe7b735bb3a36ef5db57048ba43a145e12cb705d

Observation fe6319d5-aa72-41db-8405-77fa52d251b9 · inbound

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams cites this paper.

LiveStarPro: Proactive Streaming Video Understanding with Hierarchical Memory for Long-Horizon Streams HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:38:56.175406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-06-27T01:12:46.295455Z digest=sha256:f239183a1c78dd5181b8a31967cf555e216115ab7643564895debeed0c9a284e

Observation 490c03fc-1224-4d82-beb0-35458f2dc982 · inbound

Video-Text Temporal Localization via Multi-Scale Convolution and Dynamic Routing cites this paper.

Video-Text Temporal Localization via Multi-Scale Convolution and Dynamic Routing HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-11T08:55:44.783824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-11T08:55:44.783824Z digest=sha256:7fedb1a93555ae4789dfc999f94c4fb3a0847aab7c615e02a4cb60ff712ce28e