Pith. sign in

Paper Citation Record · LEDGER

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization

As of 18 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 1 inbound Pith citation observation for arXiv:2506.14356.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.14356 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:26:05.265658Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T21:29:27.063028Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T21:35:04.650377Z

Reference resolution

54 of 54 outbound references displayed

  • verified exact3
  • verified fuzzy34
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3b3039c5-4f63-43b2-96a1-d4c4fc357c19 · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Quo vadis, action recognition? a new model and the kinetics dataset,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.550525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:25:59.431559Z digest=sha256:69720485de4f51ae431a247d62520ca565ec3696250a357dd823e4df7c0aff6b

Observation 3b05827e-9bcb-4504-b918-91617a462caa · outbound

This paper cites Video summarization through reinforcement learning with a 3d spatio- temporal u-net,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Video summarization through reinforcement learning with a 3d spatio- temporal u-net,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.535398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:25:59.542659Z digest=sha256:88fdc4f77abbd6023e18558671ce69a930864e0874c6be4bf68f95405e80650d

Observation b2e2b62d-3ff3-42f4-b510-693cd4de7b11 · outbound

This paper cites Deep attention network for egocentric action recognition,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Deep attention network for egocentric action recognition,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.521628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:25:59.700863Z digest=sha256:0c4e9f38724a999e47e5ee2631b2d6fabea98db8b1398110e7d6fe11f0afc018

Observation 416ae5dc-9eea-4ae5-8c37-38a17c72361d · outbound

This paper cites Training a Large Video Model on a Single Machine in a Day.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Training a Large Video Model on a Single Machine in a Day

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:26:05.951232Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:25:59.818943Z digest=sha256:0c7e04f7a5c1efd091f98c764125f1b0f1fd2b907452f6d1f2a464e89708685f

Observation c5d6b164-6f58-4f03-a66e-b3565b0bde8d · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T00:25:59.990074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:25:59.990074Z digest=sha256:79813043d047d43a7199f33d24bc5477fc878418024444a89cd7350c5dad2b1e

Observation 0c6c2b82-1304-4e98-8301-7975bacc3446 · outbound

This paper cites InternVideo2: Scaling Foundation Models for Multimodal Video Understanding.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization InternVideo2: Scaling Foundation Models for Multimodal Video Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:00.161920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:00.161920Z digest=sha256:ad5050bd163bcc44a565de05ed3fa7db174909cd39c00206f05c8f8b6573e9ff

Observation a2d5163b-29bd-4182-9040-32cfece62edd · outbound

This paper cites EgoVLPv2: Egocentric video-language pre-training with fusion in the backbone,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization EgoVLPv2: Egocentric video-language pre-training with fusion in the backbone,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.507959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:00.334603Z digest=sha256:b5feaee31350a8ec14fece049ba15a78ccfa13a242c306e1159f68ab4c078786

Observation 82d49f24-f819-402c-81c9-472ae702a1bd · outbound

This paper cites Egocentric video-language pretraining,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Egocentric video-language pretraining,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.495279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:00.477538Z digest=sha256:d33ef0213612dfe1fbac236d3f779cff3c778e648c36ae3077e1087a363b45d2

Observation 144cc8dd-9c57-48e0-8412-8e9a04012dba · outbound

This paper cites Improving semantic video retrieval models by training with a relevance-aware online mining strategy,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Improving semantic video retrieval models by training with a relevance-aware online mining strategy,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.481056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:00.576434Z digest=sha256:3df8ef529e31356a4223f5315fa9f59ce01c289ebb02817c1861e123922641d3

Observation 04357bdc-9f02-452b-ae9e-54018f9be793 · outbound

This paper cites Learning video representations from large language models,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Learning video representations from large language models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.467688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:00.716989Z digest=sha256:a934a65ed4d5142b54fceb139f6c2f733677da1f705963c3a903fa4b5f0875f7

Observation f9ec9a56-fbec-4170-bf4d-b6170093b1b6 · outbound

This paper cites EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization EgoVideo: Exploring Egocentric Foundation Model and Downstream Adaptation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:00.852320Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:00.852320Z digest=sha256:613f05ad410cc36e357258ed97afbbf6a7c16ad823e1cad60dcab33dabcf73b3

Observation 93e8b51b-8004-400e-9a20-d4c80012e913 · outbound

This paper cites Videomae: Masked autoen- coders are data-efficient learners for self-supervised video pre-training,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Videomae: Masked autoen- coders are data-efficient learners for self-supervised video pre-training,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.453114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:01.017157Z digest=sha256:e4f4d443091c15ec8441c191ae633416ba07b503f06897c057f748c374756845

Observation 3aa1e723-af16-4d5f-8aa8-4345f828d67c · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:01.224115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:01.224115Z digest=sha256:170e4f0e06fb6505b62fe2bfb7423e831b5b465fae2df6384e61e7876ba1ca19

Observation 4cda262c-684a-4f56-9402-0d6eabd98073 · outbound

This paper cites Internvid: A large-scale video-text dataset for multimodal understanding and generation,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Internvid: A large-scale video-text dataset for multimodal understanding and generation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.437990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:01.393406Z digest=sha256:ea38e796614f9e377bc97a4fcc78a7a9f99194fb2178a59dc062d496522c9628

Observation 69ab1970-654d-4c41-a237-ef7ac463104c · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:01.534507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:01.534507Z digest=sha256:d32852c7ef3ed1662a1f42d9a57ef67ddd180e8e60dd719806fe93b3d87488c6

Observation 7fbc78e4-1606-4303-b25e-ece1889a162a · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:01.692337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:01.692337Z digest=sha256:d3907c6d35be3ffb1040ed1d8514bd0981af94a6dc0c99e6f45244d937656e31

Observation 7176760f-561d-43fe-bc91-9b9f7c99fb15 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:01.868488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:01.868488Z digest=sha256:afd9d9f31575919c194595bd34008b86b75bc42601f038967faeb942555ae303

Observation 0c2c0630-67ea-4f02-96fb-a8e0cd32a142 · outbound

This paper cites EVA-02: A Visual Representation for Neon Genesis.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization EVA-02: A Visual Representation for Neon Genesis

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:01.964467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:01.964467Z digest=sha256:ceeda4c8849536f424cafbaa0581983298269990373195a3e4f55d9f6dddf21d

Observation 45c01c0e-2c0c-4446-ae29-20e8d6fed1c2 · outbound

This paper cites Multi- similarity loss with general pair weighting for deep metric learning,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Multi- similarity loss with general pair weighting for deep metric learning,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.424959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:02.137317Z digest=sha256:a8f3a32e476797ceb53ef7b933ebdd3be099a8ca94cd96bc4b2733a480fc3fbd

Observation ab8d33a3-8344-4b78-99e6-7feda72268a3 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Ego4d: Around the world in 3,000 hours of egocentric video,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.410562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:02.208447Z digest=sha256:67b099d152ecb23c8b27d8cd1f642cfff0be5c141cd8992260d239116efc3569

Observation 7aed2aee-2f71-4fe6-bcdd-308d381bc970 · outbound

This paper cites Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens- 100,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Rescaling egocentric vision: Collection, pipeline and challenges for epic-kitchens- 100,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:10.218130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:02.280538Z digest=sha256:c09683391347d12ae0f4870cc8aedc00768fa6f8632ca570ca19c98daeda03e5

Observation 7ba327eb-6879-4ede-adc7-5b81d3a058b7 · outbound

This paper cites Scaling egocentric vision: The epic-kitchens dataset,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Scaling egocentric vision: The epic-kitchens dataset,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:09.831272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:02.395824Z digest=sha256:d597e9eabcbd0f3d200b1fa6fa017eaa0d864b595c2d3c77c27eb4a48010cf53

Observation f8f59c61-0302-4d25-aa0e-c79e170886c6 · outbound

This paper cites Charades-Ego: A Large-Scale Dataset of Paired Third and First Person Videos.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Charades-Ego: A Large-Scale Dataset of Paired Third and First Person Videos

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:02.474819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:02.474819Z digest=sha256:be87692f925d8a78e94647a58c744bbba829e8064ae9d254e7624b39bc3f68d8

Observation 7c7b3501-3185-4670-a377-336c7b19a095 · outbound

This paper cites Long short-term memory,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Long short-term memory,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:02.579192Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:02.579192Z digest=sha256:e5c3bfd84d85358047b5621f5e27203ad4c7f0c637daccbd0cef2cd48e659aaf

Observation 0f86341f-085c-4b6c-9496-678cb767245d · outbound

This paper cites Is space-time attention all you need for video understanding?.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Is space-time attention all you need for video understanding?

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:09.719279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:02.688759Z digest=sha256:a354892c6a7abcdf4b8a531e4964b3b01de56337c4888e57539c30e0457dc08b

Observation 8a3e82d7-031b-4fea-b2b7-9e35cac358e4 · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Frozen in time: A joint video and image encoder for end-to-end retrieval,

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:02.803661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:02.803661Z digest=sha256:645bfec39785034426477730712e596df35721bcfcd07ee100e3f80863585edd

Observation df4130a7-8ddd-434b-a195-251c34c4c424 · outbound

This paper cites VideoMAE V2: Scaling video masked autoencoders with dual masking,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization VideoMAE V2: Scaling video masked autoencoders with dual masking,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:09.436074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:02.910107Z digest=sha256:653385980a48ae17e593a172fae0be659eb8c86f1e07d8741febc353e2d6208b

Observation 47013c40-a70d-4d51-b498-12e7e211c017 · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Flamingo: a visual language model for few-shot learning,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:09.120647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:02.983354Z digest=sha256:fe6b6db6a92939bf0d1651cac8e9f0fb797b3dbb1abb12aebb3884333d841f98

Observation 46fab9ae-e11f-4030-a325-dc9b400980ad · outbound

This paper cites Roformer: En- hanced transformer with rotary position embedding,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Roformer: En- hanced transformer with rotary position embedding,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:08.908407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:03.094336Z digest=sha256:2f343f289b3652dfeb1970a55c0027fc2378457196a8bc3015b4b31fcf0d8c2d

Observation d3b604c5-7e92-45b2-bcb2-12430e34096d · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:03.203928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:03.203928Z digest=sha256:6359d651499e27e9240ef03a12aa247f758543fdc47134e1d53db9e5110489a9

Observation 72d0da19-90bd-49a7-bad9-1dce2fcd7bc2 · outbound

This paper cites VideoRoPE: What Makes for Good Video Rotary Position Embedding?.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization VideoRoPE: What Makes for Good Video Rotary Position Embedding?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:03.315947Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:03.315947Z digest=sha256:254a6d71bc615e14ef051734aedd90217e44bf132d5915ee4c9111ff404c0495

Observation 1db5d195-264f-4332-b8ae-69d2eef1df4c · outbound

This paper cites Supervised contrastive learn- ing,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Supervised contrastive learn- ing,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:08.615395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:03.397899Z digest=sha256:b61d098598439aa8585812757756d44aa018ded7f037476036ec143a882971eb

Observation a6dcc2cf-ea91-4430-a219-604ca220f721 · outbound

This paper cites Parameter-free deep multi-modal clustering with reliable contrastive learning,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Parameter-free deep multi-modal clustering with reliable contrastive learning,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:08.308703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:03.518745Z digest=sha256:563941e7f45e16186544e04c51dc5e4e25820aad8f007ec4aefa4f82554f9c10

Observation 34495b16-3787-4229-aac2-8c1013cfe167 · outbound

This paper cites Cross-modal contrastive learning network for few-shot action recognition,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Cross-modal contrastive learning network for few-shot action recognition,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:08.040114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:03.600867Z digest=sha256:5c5702f63e0cc6d92e16623fc2c80fd912795010ae2671cdee4411328bdc0cc7

Observation 2423d424-3315-4b21-a30c-f5c861e37277 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Representation Learning with Contrastive Predictive Coding

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:03.695009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:03.695009Z digest=sha256:e8fe35c48b658a3866daed8affd754c15b4fdb6cc0c7998c2695cc5e6e6f8449

Observation 30bce6e3-7aa2-4791-a3b7-8242b8c22b93 · outbound

This paper cites End-to-end learning of visual representations from uncurated instructional videos,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization End-to-end learning of visual representations from uncurated instructional videos,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:07.827552Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:03.805158Z digest=sha256:8c00c21e318fcc43273808e33a40b08ff943ec9976fc54c02fb27d6bd1ca6f2b

Observation dd7dbc9d-7868-4a71-ae29-fec6e005ca2f · outbound

This paper cites Facenet: A unified embed- ding for face recognition and clustering,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Facenet: A unified embed- ding for face recognition and clustering,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:07.728889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:03.911202Z digest=sha256:f06e578e2fd43cc5ddeb17da4a91b94046e33145a59c4314c48b085482c28866

Observation 5858fc87-7a64-427c-8795-10f7672704a3 · outbound

This paper cites Circle loss: A unified perspective of pair similarity optimization,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Circle loss: A unified perspective of pair similarity optimization,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:07.605657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:03.992358Z digest=sha256:177219acf0b13b2fdc1da41aafe7818f2ba1848601163529abbe494a69e24bff

Observation 49ed26cd-60af-4674-93e2-0eb323b8581d · outbound

This paper cites Relevance-based margin for contrastively-trained video retrieval mod- els,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Relevance-based margin for contrastively-trained video retrieval mod- els,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:07.497252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:04.100339Z digest=sha256:e7efc4e999d912da142072d660c38971f843a2da149a795a0b126ae3a3ae289e

Observation 677b1418-3e49-410a-a9de-4e81b4db830a · outbound

This paper cites Masked video distillation: Rethinking masked feature modeling for self-supervised video representation learning,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Masked video distillation: Rethinking masked feature modeling for self-supervised video representation learning,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:07.350471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:04.178813Z digest=sha256:723d6c94e6d68312116e28b453240c276de073e6019d46be3ab0c600156d2aca

Observation 2f8337f0-1222-4470-8bbf-3b7e10e8222a · outbound

This paper cites Fine-grained action retrieval through multiple parts-of-speech embeddings,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Fine-grained action retrieval through multiple parts-of-speech embeddings,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:07.212680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:04.268357Z digest=sha256:5512c3770a2a2a1898c7ec61ef1e3c1e3639cd8477d5b341bdf33a5602ed8f7d

Observation aa116c9e-61d8-4509-b642-0b598643aea5 · outbound

This paper cites On semantic similarity in video retrieval,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization On semantic similarity in video retrieval,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:07.075302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:04.332494Z digest=sha256:436e201d095a7464f6fa07267ff0d23f441597431edf3a24d8b86c93b584ea64

Observation 431349d9-3713-4553-afbb-2d7961ed415d · outbound

This paper cites Egocentric Video-Language Pretraining @ EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2022.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Egocentric Video-Language Pretraining @ EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge 2022

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:26:05.630065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:04.405236Z digest=sha256:14929e822f2853e4940a374bec1425f33cb1390f7d040d9ebbc37bbca7930470

Observation d653bc96-b238-49b9-bc71-1b6c966c223c · outbound

This paper cites Collecting highly parallel data for paraphrase evaluation,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Collecting highly parallel data for paraphrase evaluation,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:06.939788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:04.491269Z digest=sha256:0a18d78debba1213f882ee8b67e3818c1a765a93d7e0b961b1b1ef6bab77381a

Observation f784b997-ab31-4fbf-b6d6-9a2195580094 · outbound

This paper cites Epic-fusion: Audio-visual temporal binding for egocentric action recognition,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Epic-fusion: Audio-visual temporal binding for egocentric action recognition,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:06.834335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:04.577115Z digest=sha256:a7a8af0937e385834c412397cda2123b78d72a756df71c6734de3bef8bd1c6c7

Observation 1b72421b-4411-4f93-bd51-689cb203ecfb · outbound

This paper cites Language models are unsupervised multitask learners,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Language models are unsupervised multitask learners,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:04.664578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:04.664578Z digest=sha256:867e7f5787765d98a41cf5eee980f627abe139eb17e150c800122451090206b7

Observation 7a5978af-2eed-4efd-9175-fb775f39d593 · outbound

This paper cites HowTo100M: Learning a text-video embedding by watching hundred million narrated video clips,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization HowTo100M: Learning a text-video embedding by watching hundred million narrated video clips,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:06.732224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:04.750789Z digest=sha256:42bb7db68a75c7b4ef1d55488445264ab3c71f55d76eb228da8f0a38a2256c13

Observation cc0c3d76-bd7f-44da-b907-3baf74be4d47 · outbound

This paper cites Rethinking spatiotem- poral feature learning: Speed-accuracy trade-offs in video classification,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Rethinking spatiotem- poral feature learning: Speed-accuracy trade-offs in video classification,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:06.594967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:04.815632Z digest=sha256:0ba25a08fc83301beaede2382a7f8f596cc82e0140063db4f31a385e2e4a028b

Observation b475f87a-9c8c-4317-b1b7-9243d14519ca · outbound

This paper cites Learning transferable visual models from natural language supervi- sion,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Learning transferable visual models from natural language supervi- sion,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:06.422516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:04.888141Z digest=sha256:9fb6e0259b57f5b9625ed96300a122b9b1915647d50100d5edff98d665577e9d

Observation e9c8fe45-b40d-420a-b7cb-4a383d5660b0 · outbound

This paper cites Hiervl: Learning hierarchical video-language embeddings,.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Hiervl: Learning hierarchical video-language embeddings,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:06.261608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:04.957092Z digest=sha256:b77843fb7d5477fb9e9fa63592b7dddf5f7653d1cba89f667396da520751f6e0

Observation c85f5577-4f51-4526-b62d-8f395bb174e9 · outbound

This paper cites SViTT-Ego: A Sparse Video-Text Transformer for Egocentric Video.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization SViTT-Ego: A Sparse Video-Text Transformer for Egocentric Video

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-08-07T00:26:05.459079Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:05.023665Z digest=sha256:8238ec5232c377f67ed340af1d8e51bbf2da9800f541558699955dd0cb55884a

Observation 5ecc0798-5e90-4dfe-858a-40557e35fb32 · outbound

This paper cites Decoupled weight decay regularization.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization Decoupled weight decay regularization

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:26:06.125884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T00:26:05.111484Z digest=sha256:187b2595ce975603a9a2b03b7ed3a8fb0977bb6316d23ed2390af23fb1336eee

Observation 4be5b5db-e73a-4a44-8640-2a6d06e242af · outbound

This paper cites DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization DistilBERT, a distilled version of BERT: smaller, faster, cheaper and lighter

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:05.197274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:05.197274Z digest=sha256:869a8260a2846d09a1652f0ef608a910e3da4a39fb17fefe2d43c066bf443e0d

Observation 784e0a32-bc55-4749-9325-564a43c2f4c0 · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T00:26:05.265658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:26:05.265658Z digest=sha256:8249b7b0894c152b8d5f2f6dc18a18a5e03d94eb8e031d245dcd28a64eb2a061

Pith citing papers

Observation f36432fc-d3bc-424a-92ef-4a6a6846fb18 · inbound

EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding cites this paper.

EARL: Towards a Unified Analysis-Guided Reinforcement Learning Framework for Egocentric Interaction Reasoning and Pixel Grounding EVA02-AT: Egocentric Video-Language Understanding with Spatial-Temporal Rotary Positional Embeddings and Symmetric Optimization

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-06-30T21:35:04.652174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-30T21:29:27.063028Z digest=sha256:1659041bf1d04ccc6cd30f9108a790d6b5c82dd068ab38ee95a444fe55db0a89