Pith. sign in

Paper Citation Record · LEDGER

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization

As of 9 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 1 inbound Pith citation observation for arXiv:2506.14356.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14356 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:26:05.265658Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T21:29:27.063028Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T21:35:04.650377Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact3
  • verified fuzzy34
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3b3039c5-4f63-43b2-96a1-d4c4fc357c19 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Quo vadis, action recognition? a new model and the kinetics dataset,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.550525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:25:59.431559Z digest=sha256:17f02fbe225f416e924670a2ba769a7b4c03bb8fcaead9a6b837acd735126297

Observation 3b05827e-9bcb-4504-b918-91617a462caa · outbound

This paper cites Video summarization through reinforcement learning with a 3d spatio- temporal u-net,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Video summarization through reinforcement learning with a 3d spatio- temporal u-net,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.535398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:25:59.542659Z digest=sha256:00c55f26f2a4560c933a323abbac677590b63d23ad5bb11a522191d2981d41ac

Observation b2e2b62d-3ff3-42f4-b510-693cd4de7b11 · outbound

This paper cites Deep attention network for egocentric action recognition,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Deep attention network for egocentric action recognition,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.521628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:25:59.700863Z digest=sha256:d25e3cbe8fe867b03948b0c81c80649f665151258d811e88a04986127e65a313

Observation 416ae5dc-9eea-4ae5-8c37-38a17c72361d · outbound

This paper cites Training a Large Video Model on a Single Machine in a Day.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Training a Large Video Model on a Single Machine in a Day

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:26:05.951232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:25:59.818943Z digest=sha256:5c7869ef8717bbdece31abd937e120180084366c5b59069b8ccf57ee9c1a42c9

Observation c5d6b164-6f58-4f03-a66e-b3565b0bde8d · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:25:59.990074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:25:59.990074Z digest=sha256:97b5d90898e11f19d94442cfd05f983ebe6e2ed665978debcd5d3776811918f2

Observation 0c6c2b82-1304-4e98-8301-7975bacc3446 · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:00.161920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:00.161920Z digest=sha256:27a4eb624927ad9eb81886055294ffdd4cbdaa07d22d46501e46ba9bd9ae8d9c

Observation a2d5163b-29bd-4182-9040-32cfece62edd · outbound

This paper cites EgoVLPv2: Egocentric video-language pre-training with fusion in the backbone,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization EgoVLPv2: Egocentric video-language pre-training with fusion in the backbone,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.507959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:00.334603Z digest=sha256:a83bbf04f29fc03c1d0dec7884851e90e375ba83d5130fc70b4ec0abe8ce4423

Observation 82d49f24-f819-402c-81c9-472ae702a1bd · outbound

This paper cites Egocentric video-language pretraining,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Egocentric video-language pretraining,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.495279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:00.477538Z digest=sha256:f9f71ceb54fd7f955674a7a3d48d06de378a398a2fe2e2c224dd4f21ece8b74b

Observation 144cc8dd-9c57-48e0-8412-8e9a04012dba · outbound

This paper cites Improving semantic video retrieval models by training with a relevance-aware online mining strategy,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Improving semantic video retrieval models by training with a relevance-aware online mining strategy,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.481056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:00.576434Z digest=sha256:4957b1e38dc10e38d24ae265b6fb4af5b144be79cb2298ad8471528afe462695

Observation 04357bdc-9f02-452b-ae9e-54018f9be793 · outbound

This paper cites Learning video representations from large language models,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Learning video representations from large language models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.467688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:00.716989Z digest=sha256:fdaec862f57f6f7be40411003c154f713025b2a1018699313dbdd14eac4235a8

Observation f9ec9a56-fbec-4170-bf4d-b6170093b1b6 · outbound

This paper cites EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:00.852320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:00.852320Z digest=sha256:21313f026dad9ff58d4e4e67d4d59ef14a22ed82da5a258315bb34f12df79c43

Observation 93e8b51b-8004-400e-9a20-d4c80012e913 · outbound

This paper cites Videomae: Masked autoen- coders are data-efficient learners for self-supervised video pre-training,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Videomae: Masked autoen- coders are data-efficient learners for self-supervised video pre-training,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.453114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:01.017157Z digest=sha256:053968e5137ac7fbcac79c95e962fe6ac702a2f7eaa7bfcb684d5b4beb64ba6c

Observation 3aa1e723-af16-4d5f-8aa8-4345f828d67c · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:01.224115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:01.224115Z digest=sha256:a9078a069869fd52b2342a3a9728fd2c3d3d3074003ce71b8bade8fa44af8103

Observation 4cda262c-684a-4f56-9402-0d6eabd98073 · outbound

This paper cites Internvid: A large-scale video-text dataset for multimodal understanding and generation,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Internvid: A large-scale video-text dataset for multimodal understanding and generation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.437990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:01.393406Z digest=sha256:12b3b2995ef0e39bcacd99bf5c07e71ee71c08947a28a705165f32b6bf053ac2

Observation 69ab1970-654d-4c41-a237-ef7ac463104c · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:01.534507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:01.534507Z digest=sha256:59bc8735c7bae4fd2612307d8ba10e470b88091adfe380f2f014538857efcbf2

Observation 7fbc78e4-1606-4303-b25e-ece1889a162a · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:01.692337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:01.692337Z digest=sha256:d925401ba7c5c98c98fbd0694d8df87ba049facdd329881e2edef219eba224c4

Observation 7176760f-561d-43fe-bc91-9b9f7c99fb15 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:01.868488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:01.868488Z digest=sha256:645037cfb598dc94f1927d9a84fc89d82cfb953c5e621b94d0c8740dd8766783

Observation 0c2c0630-67ea-4f02-96fb-a8e0cd32a142 · outbound

This paper cites EVA-02: A Visual Representation for Neon Genesis.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization EVA-02: A Visual Representation for Neon Genesis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:01.964467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:01.964467Z digest=sha256:6553a298e9050e7b07ea8ad9c207a2853948adc3b951efe2d0cedd79e0766aa5

Observation 45c01c0e-2c0c-4446-ae29-20e8d6fed1c2 · outbound

This paper cites Multi- similarity loss with general pair weighting for deep metric learning,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Multi- similarity loss with general pair weighting for deep metric learning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.424959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:02.137317Z digest=sha256:65db9cf00fa001f22d791396d38270213648636713fada0982f5167031529850

Observation ab8d33a3-8344-4b78-99e6-7feda72268a3 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Ego4d: Around the world in 3,000 hours of egocentric video,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.410562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:02.208447Z digest=sha256:e476e7a66fb1ebb40de45c7daae839fd8f39f2e1c63b608d50c0f732eda33a93

Observation 7aed2aee-2f71-4fe6-bcdd-308d381bc970 · outbound

This paper cites Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens- 100,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens- 100,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.218130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:02.280538Z digest=sha256:b91f80abf0ddd4281935ac13a4156ba92d498e64c61793baa2637a8c339e3e2d

Observation 7ba327eb-6879-4ede-adc7-5b81d3a058b7 · outbound

This paper cites Scaling egocentric vision: The epic-kitchens dataset,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Scaling egocentric vision: The epic-kitchens dataset,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:09.831272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:02.395824Z digest=sha256:ce9066157fe4ef130796b87cc2d9ca8afbe1fb9c31ce21c37a5c203b88b34ef0

Observation f8f59c61-0302-4d25-aa0e-c79e170886c6 · outbound

This paper cites Charades-Ego: A Large-Scale Dataset of Paired Third and First Person Videos.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Charades-Ego: A Large-Scale Dataset of Paired Third and First Person Videos

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:02.474819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:02.474819Z digest=sha256:f240141df29d2954057c9183ea39405d565c770744898344026ec11d71099960

Observation 7c7b3501-3185-4670-a377-336c7b19a095 · outbound

This paper cites Long short-term memory,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Long short-term memory,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:02.579192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:02.579192Z digest=sha256:8277cdd54e70155cd5069a3d7d26988b251783a2698fe8ec9c6f0aa3414ffef9

Observation 0f86341f-085c-4b6c-9496-678cb767245d · outbound

This paper cites Is space-time attention all you need for video understanding?.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Is space-time attention all you need for video understanding?

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:09.719279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:02.688759Z digest=sha256:521a3b002aaa2301c726389a728ba4808c09cf35003880665afbbdf56bc18a43

Observation 8a3e82d7-031b-4fea-b2b7-9e35cac358e4 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Frozen in time: A joint video and image encoder for end-to-end retrieval,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:02.803661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:02.803661Z digest=sha256:0726a5b563f6b3a453220336144d36f34ef4044e4cd32e969112b89f7a9a3d08

Observation df4130a7-8ddd-434b-a195-251c34c4c424 · outbound

This paper cites VideoMAE V2: Scaling video masked autoencoders with dual masking,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization VideoMAE V2: Scaling video masked autoencoders with dual masking,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:09.436074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:02.910107Z digest=sha256:9461c104bb74b29a7b692df25fa987f65fbbecc6ba4f647a2eaef414148f3ab1

Observation 47013c40-a70d-4d51-b498-12e7e211c017 · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Flamingo: a visual language model for few-shot learning,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:09.120647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:02.983354Z digest=sha256:3af52852eda8f5d96d1ceab974d4d4d49da1dc453f74caaf459ba2aa6c694583

Observation 46fab9ae-e11f-4030-a325-dc9b400980ad · outbound

This paper cites Roformer: En- hanced transformer with rotary position embedding,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Roformer: En- hanced transformer with rotary position embedding,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:08.908407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:03.094336Z digest=sha256:e062a0030ce926b543e0b4ea8a03c18bc07acd0aed67dcea53b6fe4525f93e09

Observation d3b604c5-7e92-45b2-bcb2-12430e34096d · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:03.203928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:03.203928Z digest=sha256:037fda000a70ea95b2ef352391ec52b28bfe78f51572258d0fa329c7fb2ef89d

Observation 72d0da19-90bd-49a7-bad9-1dce2fcd7bc2 · outbound

This paper cites VideoRoPE: What Makes for Good Video Rotary Position Embedding?.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:03.315947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:03.315947Z digest=sha256:c14e01032e3e07be4eda2bfd6c2b87f8952eabcf9f4214339a3bce20e484cfc6

Observation 1db5d195-264f-4332-b8ae-69d2eef1df4c · outbound

This paper cites Supervised contrastive learn- ing,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Supervised contrastive learn- ing,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:08.615395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:03.397899Z digest=sha256:eec9792adf6fa2bb925ee05b0294d3b96dbe23bb054839e03f8eaa1961bb9661

Observation a6dcc2cf-ea91-4430-a219-604ca220f721 · outbound

This paper cites Parameter-free deep multi-modal clustering with reliable contrastive learning,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Parameter-free deep multi-modal clustering with reliable contrastive learning,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:08.308703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:03.518745Z digest=sha256:fdb0ad8ba910a6d57360a9cff3bda6d9319c5c2f209d54a812c748e0d92fd2ef

Observation 34495b16-3787-4229-aac2-8c1013cfe167 · outbound

This paper cites Cross-modal contrastive learning network for few-shot action recognition,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Cross-modal contrastive learning network for few-shot action recognition,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:08.040114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:03.600867Z digest=sha256:125bd789f1c55f0fa9343cf45a8e93283577d0fc435d9b5327fe2ea0e69066da

Observation 2423d424-3315-4b21-a30c-f5c861e37277 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Representation Learning with Contrastive Predictive Coding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:03.695009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:03.695009Z digest=sha256:393bbe753fa3c81dc9c114f09581065141f3f6f1dc221a6bfe9e1825cb0b18bf

Observation 30bce6e3-7aa2-4791-a3b7-8242b8c22b93 · outbound

This paper cites End-to-end learning of visual representations from uncurated instructional videos,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization End-to-end learning of visual representations from uncurated instructional videos,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:07.827552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:03.805158Z digest=sha256:35710d1b22a3179df6d7a772dc1ff57251242de31ddbd1c7f6938e885ed97e69

Observation dd7dbc9d-7868-4a71-ae29-fec6e005ca2f · outbound

This paper cites Facenet: A unified embed- ding for face recognition and clustering,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Facenet: A unified embed- ding for face recognition and clustering,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:07.728889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:03.911202Z digest=sha256:f469b58c1c1eb1116ea39c2674cd9545bf9f72244835d1d2318c88c2c1eea494

Observation 5858fc87-7a64-427c-8795-10f7672704a3 · outbound

This paper cites Circle loss: A unified perspective of pair similarity optimization,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Circle loss: A unified perspective of pair similarity optimization,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:07.605657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:03.992358Z digest=sha256:80ba3c088aec8704c6c892ebd683846e9b37182a28c3cf55e87dd8441e68220e

Observation 49ed26cd-60af-4674-93e2-0eb323b8581d · outbound

This paper cites Relevance-based margin for contrastively-trained video retrieval mod- els,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Relevance-based margin for contrastively-trained video retrieval mod- els,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:07.497252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:04.100339Z digest=sha256:0ea6a26c35e003c4318f5326a3100f81d02a717b7a77672969544d0c27b165fc

Observation 677b1418-3e49-410a-a9de-4e81b4db830a · outbound

This paper cites Masked video distillation: Rethinking masked feature modeling for self-supervised video representation learning,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Masked video distillation: Rethinking masked feature modeling for self-supervised video representation learning,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:07.350471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:04.178813Z digest=sha256:a11bbadf93675c17b10036f07f7daf6ecb7fbaa9fe9c8e788f1a2938a39e92a7

Observation 2f8337f0-1222-4470-8bbf-3b7e10e8222a · outbound

This paper cites Fine-grained action retrieval through multiple parts-of-speech embeddings,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Fine-grained action retrieval through multiple parts-of-speech embeddings,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:07.212680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:04.268357Z digest=sha256:56e989b45b63c5c18b25796bc489afb488953316188a5415f05503731895edc5

Observation aa116c9e-61d8-4509-b642-0b598643aea5 · outbound

This paper cites On semantic similarity in video retrieval,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization On semantic similarity in video retrieval,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:07.075302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:04.332494Z digest=sha256:9c3a94571c243760191905091e00b6ad1fd1e665dd66ebad8fd6702b50d7c814

Observation 431349d9-3713-4553-afbb-2d7961ed415d · outbound

This paper cites Egocentric Video-Language Pretraining @ EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2022.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Egocentric Video-Language Pretraining @ EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2022

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:26:05.630065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:04.405236Z digest=sha256:acce677c65c6edabdc0211312368322379720ef4b3917cb43a4398b625e3271b

Observation d653bc96-b238-49b9-bc71-1b6c966c223c · outbound

This paper cites Collecting highly parallel data for paraphrase evaluation,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Collecting highly parallel data for paraphrase evaluation,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:06.939788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:04.491269Z digest=sha256:ccf0be9f9db5d1ff3e2ffd2ba446b580da5adb2e7fb5ad4f7973dfa002a1f2f6

Observation f784b997-ab31-4fbf-b6d6-9a2195580094 · outbound

This paper cites Epic-fusion: Audio-visual temporal binding for egocentric action recognition,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Epic-fusion: Audio-visual temporal binding for egocentric action recognition,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:06.834335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:04.577115Z digest=sha256:d09116deee863fcebb7eeff27b4ca667ccde0a3be59e64a2d009548f3664083a

Observation 1b72421b-4411-4f93-bd51-689cb203ecfb · outbound

This paper cites Language models are unsupervised multitask learners,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Language models are unsupervised multitask learners,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:04.664578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:04.664578Z digest=sha256:748b0d364dd37bd3b23a76729d52ee3b930c19c5a7480140e661bd10664d3129

Observation 7a5978af-2eed-4efd-9175-fb775f39d593 · outbound

This paper cites HowTo100M: Learning a text-video embedding by watching hundred million narrated video clips,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization HowTo100M: Learning a text-video embedding by watching hundred million narrated video clips,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:06.732224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:04.750789Z digest=sha256:7bb80a192cf605439a9af86f6ad104835d713d5a188dcb4a9a1e38ce31f8951c

Observation cc0c3d76-bd7f-44da-b907-3baf74be4d47 · outbound

This paper cites Rethinking spatiotem- poral feature learning: Speed-accuracy trade-offs in video classification,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Rethinking spatiotem- poral feature learning: Speed-accuracy trade-offs in video classification,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:06.594967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:04.815632Z digest=sha256:96ef35277ef026d1910e90b06ab6e136e244c9a67714f5361a3e083a77203020

Observation b475f87a-9c8c-4317-b1b7-9243d14519ca · outbound

This paper cites Learning transferable visual models from natural language supervi- sion,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Learning transferable visual models from natural language supervi- sion,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:06.422516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:04.888141Z digest=sha256:9074f103dbd87797ad8c68ec6b9c8596e1769a03ff621a22187be21f0a9c3b9b

Observation e9c8fe45-b40d-420a-b7cb-4a383d5660b0 · outbound

This paper cites Hiervl: Learning hierarchical video-language embeddings,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Hiervl: Learning hierarchical video-language embeddings,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:06.261608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:04.957092Z digest=sha256:d73e22f1d95c0bf31085da5700e8a9d932e488eb445928142fef2e8749e5d55f

Observation c85f5577-4f51-4526-b62d-8f395bb174e9 · outbound

This paper cites SViTT-Ego: A Sparse Video-Text Transformer for Egocentric Video.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization SViTT-Ego: A Sparse Video-Text Transformer for Egocentric Video

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:26:05.459079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:05.023665Z digest=sha256:beddbec837f2c34f8c28352ea021880dc577ec00fa7fab42b91cc3729236635a

Observation 5ecc0798-5e90-4dfe-858a-40557e35fb32 · outbound

This paper cites Decoupled weight decay regularization.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Decoupled weight decay regularization

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:06.125884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T00:26:05.111484Z digest=sha256:3b6ab33311f007d22ab070b85a17e0b08816470919f05e6b414b3f081349fa79

Observation 4be5b5db-e73a-4a44-8640-2a6d06e242af · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:05.197274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:05.197274Z digest=sha256:741a6886e6ab43641674f7e32659f912ef6d0c873a0f81cb35898e64ca571994

Observation 784e0a32-bc55-4749-9325-564a43c2f4c0 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:05.265658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:05.265658Z digest=sha256:16ef450c1537c082392a443bf43a3cb8910eb3e61879388c2fc5d93175964717

Pith citing papers

Observation f36432fc-d3bc-424a-92ef-4a6a6846fb18 · inbound

EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding cites this paper.

EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:35:04.652174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T21:29:27.063028Z digest=sha256:a6cc803bc676b5af04845aae966102ad46afdcb14170f00c70bc6358a916a5d6