Pith. sign in

Paper Citation Record · LEDGER

Principles of Visual Tokens for Efficient Video Understanding

As of 15 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2411.13626.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13626 v2

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T16:38:09.895402Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact3
  • verified fuzzy32
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5a99891b-2e49-4489-8005-51c9dc53d51a · outbound

This paper cites Vivit: A video vi- sion transformer.

Principles of Visual Tokens for Efficient Video Understanding Vivit: A video vi- sion transformer

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:11.107604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.538275Z digest=sha256:e7c471970398a215e7ed4ef7a61a207cfd1993e35db4b5f5a101447020ac60be

Observation c9348b78-6a67-4449-bb64-609cc30b4912 · outbound

This paper cites Is Space-Time Attention All You Need for Video Understanding?.

Principles of Visual Tokens for Efficient Video Understanding Is Space-Time Attention All You Need for Video Understanding?

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.546390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.546390Z digest=sha256:3a0b9ad732297d2fba7b4a79c33db0b69cba734f2d93680c2673074e29bdfd62

Observation 49cabee7-aefe-4090-8476-4ad7f9bdc92f · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, page 4, 2021.

Principles of Visual Tokens for Efficient Video Understanding Is space-time attention all you need for video understanding? In ICML, page 4, 2021

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.554368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.554368Z digest=sha256:c202f243cad24f42460f05c17f2bfa27db191a4adb157fb90680cdb3b0c768f0

Observation 81fd44e9-98bc-428e-9fa7-c3e18c494187 · outbound

This paper cites Token merging: Your ViT but faster.

Principles of Visual Tokens for Efficient Video Understanding Token merging: Your ViT but faster

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:11.075015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.561593Z digest=sha256:adba54c2777f9c00d8b0392e912a1f9e456e4cb2eb2ce18d9e97d56f4d51e813

Observation 42b7f1f4-3f60-4fac-a742-73a4fe5dc6ef · outbound

This paper cites Revisiting the” video” in video-language understanding.

Principles of Visual Tokens for Efficient Video Understanding Revisiting the” video” in video-language understanding

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:11.052252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.568188Z digest=sha256:4fd03c1f5a6daa6d8d16409e8887c1bfaac7496857223e1928f13d5eddc6d8e5

Observation 9651a4fb-f220-45f0-8ef8-75e05b8d0472 · outbound

This paper cites Space-time mixing attention for video transformer.

Principles of Visual Tokens for Efficient Video Understanding Space-time mixing attention for video transformer

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:11.031143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.584966Z digest=sha256:b211bc77b30c78fa4d2e0b707eb28f2d59b06d933bc8bb54d0ed5ff7cf25a3be

Observation fabe6ce3-449a-4b92-b86a-857a4caa921e · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

Principles of Visual Tokens for Efficient Video Understanding Activitynet: A large-scale video benchmark for human activity understanding

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.595147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.595147Z digest=sha256:fc3e938c3638d6a2c5cc2208b68827011a25a5ad5afc05ecd785bcc3b7168add

Observation e9612554-9565-471e-920a-d940a518b9df · outbound

This paper cites Quo vadis, action recognition? a new model and the kinetics dataset.

Principles of Visual Tokens for Efficient Video Understanding Quo vadis, action recognition? a new model and the kinetics dataset

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:11.001220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.609017Z digest=sha256:6def6c408f416b19cfa9b815831e5bed40e570d3287c2dda624c727bcbf91445

Observation a852fa36-f5f0-458f-bc78-b053c27ab8f5 · outbound

This paper cites An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models, 2024.

Principles of Visual Tokens for Efficient Video Understanding An image is worth 1/2 tokens after layer 2: Plug-and-play inference acceleration for large vision-language models, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.981335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.615828Z digest=sha256:3b2db01c8112ccf02ec447aeaf688eda2d70de206866fabc05aaf03835fa0e75

Observation 25b26d8c-8dc3-44cb-bad3-b381833ec000 · outbound

This paper cites an unresolved cited work.

Principles of Visual Tokens for Efficient Video Understanding Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:38:10.960743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.622145Z digest=sha256:b5f4d66e47a8ea8addb949d0db10058284dd3409d2fd8f5c15867eedd015f97a

Observation 47cf182b-374d-48fe-93a6-9457ea518760 · outbound

This paper cites Prune spatio-temporal tokens by semantic-aware temporal accumulation.

Principles of Visual Tokens for Efficient Video Understanding Prune spatio-temporal tokens by semantic-aware temporal accumulation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.941052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.629775Z digest=sha256:1e969b84ec948988211734e0622370161afb2bd1e992fa0ecc91c4142afcc269

Observation ff94fd5d-3e5c-40e1-83de-4cce007a8738 · outbound

This paper cites An image is worth 16x16 words: Trans- formers for image recognition at scale.

Principles of Visual Tokens for Efficient Video Understanding An image is worth 16x16 words: Trans- formers for image recognition at scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.635930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.635930Z digest=sha256:8d6056a2c048afba6c746657ac13db8b095d815b82266f6416f6ca1a4607a710

Observation 5b345fdf-3da9-4331-a4cc-dd588465739c · outbound

This paper cites Multiscale vision transformers.

Principles of Visual Tokens for Efficient Video Understanding Multiscale vision transformers

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.910305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.643481Z digest=sha256:922102b9f35095b5f9c2cde83a81ca4399bb989eb428c133fb5ca45b7edb077f

Observation 48f3a647-1921-41d7-8c5e-859bf389090c · outbound

This paper cites X3d: Expanding architectures for efficient video recognition.

Principles of Visual Tokens for Efficient Video Understanding X3d: Expanding architectures for efficient video recognition

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.890972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.649118Z digest=sha256:a242df4ec7c404e7fbca610e80b46eba8fc6e31cb605376a56081af69875c338

Observation 0bd73605-a781-43b2-ac83-0abdbf5e436a · outbound

This paper cites Slowfast networks for video recognition.

Principles of Visual Tokens for Efficient Video Understanding Slowfast networks for video recognition

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.658326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.658326Z digest=sha256:68408462b2605e6a4a2d921c24087373c26858ca0f4eb0f29241990ab7e9e42b

Observation a7ea3224-d89b-494d-a436-a9a769696a5a · outbound

This paper cites Efficient video transformers via spatial-temporal token merging for action recognition.

Principles of Visual Tokens for Efficient Video Understanding Efficient video transformers via spatial-temporal token merging for action recognition

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.851935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.664529Z digest=sha256:b576b84c8be4fe1a89a81a254779015c123eed6542955ae48795e92ecf327dda

Observation ec917cda-c365-402b-a165-4e86a39f60ab · outbound

This paper cites Smart frame selection for action recognition.

Principles of Visual Tokens for Efficient Video Understanding Smart frame selection for action recognition

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.670309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.670309Z digest=sha256:f4f85344a0e1450bd9535f86d6da3f05924dbdbbf416b116cc2eb1e8e2502149

Observation a3eeda23-3d4a-4dea-9fd3-c64001b2df71 · outbound

This paper cites Watt For What: Rethinking Deep Learning's Energy-Performance Relationship.

Principles of Visual Tokens for Efficient Video Understanding Watt For What: Rethinking Deep Learning's Energy-Performance Relationship

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.676924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.676924Z digest=sha256:a47e1622d6a58475c2ffe024f53f1815873110bd3b598fe57cd6d35205ae54d6

Observation bb986a7b-b58f-402c-a25d-1fc5628338a2 · outbound

This paper cites Optimizing factorized encoder models: Time and memory reduction for scalable and efficient action recognition.

Principles of Visual Tokens for Efficient Video Understanding Optimizing factorized encoder models: Time and memory reduction for scalable and efficient action recognition

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.812740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.684927Z digest=sha256:5f68fae6799f9cf5c6d5792d663b9bda97ae3f95d44ae8ba74b79cebfa9e102c

Observation 2edfc8e9-b2ce-4bce-a2eb-53498af96b52 · outbound

This paper cites something something.

Principles of Visual Tokens for Efficient Video Understanding something something

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.787929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.691332Z digest=sha256:03ea3ac1568465bcb0faf6cb9785c89796075ae372b9da1a7656cf074c9f74ce

Observation c0a9017a-ad20-4001-8630-4e55eef1602e · outbound

This paper cites Ava: A video dataset of spatio-temporally localized atomic visual actions.

Principles of Visual Tokens for Efficient Video Understanding Ava: A video dataset of spatio-temporally localized atomic visual actions

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.697107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.697107Z digest=sha256:81dc9f59afea2d63ec83447e8e6f7ad55cb3c4fc38d386bb2679e73cb5566d0e

Observation adfa02c1-ed9d-4c8d-90f7-a8970e545584 · outbound

This paper cites an unresolved cited work.

Principles of Visual Tokens for Efficient Video Understanding Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-12T16:38:10.753008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.702748Z digest=sha256:d6889f1ae80b0026a65c3dc054309326c8437a1e0c85f97624749510257ac0d1

Observation cfa34fc8-2a72-4cc2-8ae8-ac1bf3f41028 · outbound

This paper cites LookupViT: Compressing visual information to a limited number of tokens.

Principles of Visual Tokens for Efficient Video Understanding LookupViT: Compressing visual information to a limited number of tokens

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-08-12T16:38:10.119977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.707821Z digest=sha256:79943797f14a730dd64b8b99e54912ac2bdf873799a2c3e0ddfdb95c4e924dd2

Observation ecbcfb01-bc9f-473c-92a5-b1b7cdfe0ec4 · outbound

This paper cites Revisiting token pruning for object detection and instance segmentation.

Principles of Visual Tokens for Efficient Video Understanding Revisiting token pruning for object detection and instance segmentation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.714750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.714750Z digest=sha256:8a1041d74e502b67936e837d4f030ee8ba4d90ae310076e6f308d2d7e8dd943f

Observation 602013d9-90ab-49bc-873f-38010dd65c0a · outbound

This paper cites Swin transformer: 9 Hierarchical vision transformer using shifted windows.

Principles of Visual Tokens for Efficient Video Understanding Swin transformer: 9 Hierarchical vision transformer using shifted windows

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.718586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.720206Z digest=sha256:c0860c5f78991bf6b138abdb27baa88f8a278456a8b0872f021d1f50706f67b2

Observation 1340edae-8416-4d56-8865-ae5b6f5ede89 · outbound

This paper cites Video swin transformer.

Principles of Visual Tokens for Efficient Video Understanding Video swin transformer

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.695342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.725399Z digest=sha256:88c1cf6e90c75d43c5ffc5c3ccde00b77a2833112467e433bdaae00e5951d145

Observation 59cbb268-d70c-42cd-975b-dc68b59c9c99 · outbound

This paper cites Video transformer network.

Principles of Visual Tokens for Efficient Video Understanding Video transformer network

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.671573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.732078Z digest=sha256:1db73a9fdfb07a3c6cdbd4e9985871c4403a4a00426029246db6b90d167e00d9

Observation 6bb2d98a-286b-487b-8c46-1544f41c56a2 · outbound

This paper cites Expanding language-image pretrained models for gen- eral video recognition.

Principles of Visual Tokens for Efficient Video Understanding Expanding language-image pretrained models for gen- eral video recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.737194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.737194Z digest=sha256:5379abd1a65e43c57e10ac4a372fb1cf63e8183a3dcca258f8cc342597d9d888

Observation d0797f54-8feb-45a7-8fdf-0511431ca87b · outbound

This paper cites St-adapter: Parameter-efficient image-to-video transfer learning.

Principles of Visual Tokens for Efficient Video Understanding St-adapter: Parameter-efficient image-to-video transfer learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.744700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.744700Z digest=sha256:ad8dcd579308ab5eabc668c954dd51b5a0379cff14361090260cf4b55ea75a26

Observation a513123d-c351-407e-95d4-8c80f29e1fa2 · outbound

This paper cites K-centered patch sampling for efficient video recognition.

Principles of Visual Tokens for Efficient Video Understanding K-centered patch sampling for efficient video recognition

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.624875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.750699Z digest=sha256:894c668df3c67a7d81476e5b206900ca12a68f86d461544535520adb0f4c76b7

Observation 34fbc06d-8cd9-44a6-8a19-3d578edff604 · outbound

This paper cites Asano, Is- han Misra Florian Metze, Christoph Feichtenhofer, Andrea Vedaldi, and Jo ˜ao F.

Principles of Visual Tokens for Efficient Video Understanding Asano, Is- han Misra Florian Metze, Christoph Feichtenhofer, Andrea Vedaldi, and Jo ˜ao F

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.602725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.756350Z digest=sha256:2e3e020053fc3f6444bec15fa2f746ca9c679eb14db9181afd59ea3498a44a03

Observation d5be688a-fbbc-48ed-862b-48795aed311f · outbound

This paper cites So, Maud Texier, and Jeff Dean.

Principles of Visual Tokens for Efficient Video Understanding So, Maud Texier, and Jeff Dean

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.578776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.761549Z digest=sha256:b5d4b884cf6b882e655efd9a6ab00b946a0dc501ffce2b3fbd4ce0669966b6d9

Observation 89f85019-3d55-4d11-a005-a117a8eef86b · outbound

This paper cites How does the primate brain combine generative and discriminative computations in vision?.

Principles of Visual Tokens for Efficient Video Understanding How does the primate brain combine generative and discriminative computations in vision?

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-12T16:38:10.086545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.767230Z digest=sha256:e9a2da2516f073e2365ce87a4b75b0fa33c9c5c8aa5c0a24a22569fd1f7bfcb0

Observation afecf784-5a76-4a59-b954-f984930b3757 · outbound

This paper cites Dynamicvit: Efficient vision transformers with dynamic token sparsification.

Principles of Visual Tokens for Efficient Video Understanding Dynamicvit: Efficient vision transformers with dynamic token sparsification

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.773955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.773955Z digest=sha256:b7eddbb1bcef300039bdbc0f2effa5a71279958702d06db85ea4d785776efa4e

Observation c6e191aa-4563-456b-9d38-200a2507f493 · outbound

This paper cites TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?.

Principles of Visual Tokens for Efficient Video Understanding TokenLearner: What Can 8 Learned Tokens Do for Images and Videos?

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.781891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.781891Z digest=sha256:b815bc4cd4c3da34acb8b723857ec2951685de789dc7a1f7450049fdb6c8c95f

Observation c64056be-71b2-4f73-a879-28f300e8441f · outbound

This paper cites Selvaraju, Abhishek Das, Ramakrishna Vedantam, Michael Cogswell, Devi Parikh, and Dhruv Ba- tra.

Principles of Visual Tokens for Efficient Video Understanding Selvaraju, Abhishek Das, Ramakrishna Vedantam, Michael Cogswell, Devi Parikh, and Dhruv Ba- tra

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.531612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.788899Z digest=sha256:9f360cc2a468a1ebb28dfdab2e85ed6d86a1c8b08725f05072893630e7d9a6bb

Observation 253eb58c-3783-48b2-8a40-03be771481f6 · outbound

This paper cites Only time can tell: Discovering temporal data for temporal modeling.

Principles of Visual Tokens for Efficient Video Understanding Only time can tell: Discovering temporal data for temporal modeling

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.507034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.795722Z digest=sha256:cd134abe146a06550447ba5d04749fe39c928754099e916a61f7645cc0c2daca

Observation d826674c-9cc0-4f29-8b22-2b105315db20 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Principles of Visual Tokens for Efficient Video Understanding UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.801167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.801167Z digest=sha256:ab27c98a1dabda19978620fb64f1456ac08db13cf75158a5aa00fefa4cfcd471

Observation 74e6b19c-f94e-4ecb-9e89-7e5ed0031801 · outbound

This paper cites VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training.

Principles of Visual Tokens for Efficient Video Understanding VideoMAE: Masked Autoencoders are Data-Efficient Learners for Self-Supervised Video Pre-Training

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.806583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.806583Z digest=sha256:59f4525d060eb858fd58392681c4bfd51cf62a67eb27e6392cd9482bda4956c0

Observation db30876b-61c8-4e75-900b-c2afe0611f5f · outbound

This paper cites Training data-efficient image transformers & distillation through at- tention.

Principles of Visual Tokens for Efficient Video Understanding Training data-efficient image transformers & distillation through at- tention

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.812884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.812884Z digest=sha256:acbd693729e00efa419cbf23d4c058e9dd6a8f7df3d132211eeca10437fed91c

Observation 42879cc2-add6-46c6-ae1a-77b6447084ba · outbound

This paper cites Attention is all you need.

Principles of Visual Tokens for Efficient Video Understanding Attention is all you need

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.463112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.818358Z digest=sha256:89ddac6112b7ef0a43278f187038cab59c3a2d81d5033ebfbcdb2dcae8547d86

Observation 723c88fb-5569-4fa7-b021-9037e90ce588 · outbound

This paper cites Efficient video transformers with spatial- temporal token selection.

Principles of Visual Tokens for Efficient Video Understanding Efficient video transformers with spatial- temporal token selection

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.430041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.823324Z digest=sha256:b2257148b58ea69ea8e22dc2f5b2c64ba2775f59bb90a1d139ee01e9d452c600

Observation dffe762e-8325-4ff8-993b-c2834b93bc77 · outbound

This paper cites Actionclip: Adapting language-image pretrained models for video action recognition.

Principles of Visual Tokens for Efficient Video Understanding Actionclip: Adapting language-image pretrained models for video action recognition

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.405228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.829073Z digest=sha256:bccc861b73aed57fbbb0c096d200c79c3b55f7eee664e2063050892eb5a5e22c

Observation 17e7c3c5-5d9f-40c0-8763-a95ada95aafc · outbound

This paper cites Vila: Efficient video-language alignment for video question answering.

Principles of Visual Tokens for Efficient Video Understanding Vila: Efficient video-language alignment for video question answering

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.378103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.833947Z digest=sha256:f28eccf5bc621858fe72a9c421faf0541319b741ca8ff60228ef57a99a3795fd

Observation 84b220ef-6547-4687-ac53-421d9391f825 · outbound

This paper cites Video-focalnets: Spatio-temporal focal modu- lation for video action recognition.

Principles of Visual Tokens for Efficient Video Understanding Video-focalnets: Spatio-temporal focal modu- lation for video action recognition

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.351817Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.841034Z digest=sha256:ecfcceee4bacba7589444f1b5866b1e7ae6a97848b62c6b2f18293717f72667d

Observation 665eb6e4-ab17-43a3-b74d-afbf74dd7766 · outbound

This paper cites Manmatha, Alex Smola, and Philipp Kr¨ahenb¨uhl.

Principles of Visual Tokens for Efficient Video Understanding Manmatha, Alex Smola, and Philipp Kr¨ahenb¨uhl

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.325086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.847893Z digest=sha256:be5366f494e77768419d41212e98c3fba61620a6e89c12311dff6d66552f3935

Observation 7b6373a5-13ef-4083-b0ba-fab7f0b5b684 · outbound

This paper cites Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition.

Principles of Visual Tokens for Efficient Video Understanding Memvit: Memory-augmented multiscale vision transformer for efficient long-term video recognition

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.306355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.853463Z digest=sha256:6a74070f2374c67e044a7e1f6c038d3ee2e43baa3b743316671415841eab88a5

Observation 79bb0272-bd07-4e33-93ce-2f1be485168c · outbound

This paper cites Can i trust your answer? visually grounded video question answering.

Principles of Visual Tokens for Efficient Video Understanding Can i trust your answer? visually grounded video question answering

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T16:38:09.859186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T16:38:09.859186Z digest=sha256:1f3c8f74d99ad8fdab92e995b4fcaebb33f4464b8164d42b08afd43e5bae3e1b

Observation d5201b10-1d8b-4aa0-a730-2072546fec75 · outbound

This paper cites Aim: Adapting image models for effi- cient video action recognition.

Principles of Visual Tokens for Efficient Video Understanding Aim: Adapting image models for effi- cient video action recognition

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.271496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.865766Z digest=sha256:92eda0cb4304ec6bc369f391ba728d7f1012f192b9ac076f38e20572585c51c3

Observation 7768a25e-84b1-41fe-a56b-15c65398e260 · outbound

This paper cites A-vit: Adap- tive tokens for efficient vision transformer.

Principles of Visual Tokens for Efficient Video Understanding A-vit: Adap- tive tokens for efficient vision transformer

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.253123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.872069Z digest=sha256:8dd4872f6ac6f5bf18a88a4160dc89e18da38608a5e620ca31d47562fde956e1

Observation a7ba7476-5d2f-4e15-8c85-82037d865916 · outbound

This paper cites Self-chained image-language model for video localization and question answering.

Principles of Visual Tokens for Efficient Video Understanding Self-chained image-language model for video localization and question answering

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.235125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.877725Z digest=sha256:e46d9f51af2e81faee2b3859944afa0dc85942d5c8286593de9d68f3eee4a8df

Observation 6f62fb58-6e13-44f0-b29d-15c2176a3f41 · outbound

This paper cites Pyramid feature attention net- work for saliency detection.

Principles of Visual Tokens for Efficient Video Understanding Pyramid feature attention net- work for saliency detection

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.217473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.883184Z digest=sha256:cb3104aaf9d07d1acccfe4c1d9e36265253837ac8bbb3241d859340340e715fc

Observation c2c53224-c922-44fb-bbfc-34ab9845defc · outbound

This paper cites How can objects help action recognition? 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2353–2362, 2023.

Principles of Visual Tokens for Efficient Video Understanding How can objects help action recognition? 2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 2353–2362, 2023

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T16:38:10.197334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.889971Z digest=sha256:8cda11c1316cafb845fb94c15bd77f937d4a29556f4da34af17befa679fb3026

Observation 95ba41bf-3a83-443d-8bac-6c62795b021e · outbound

This paper cites ECO: Efficient Convolutional Network for Online Video Understanding.

Principles of Visual Tokens for Efficient Video Understanding ECO: Efficient Convolutional Network for Online Video Understanding

Reference 54

Resolution
verified exact
local_arxiv, observed 2026-08-12T16:38:09.954883Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-12T16:38:09.895402Z digest=sha256:2aa3f4310b020ed1ee6928578d3723c7e826b5b837f1bf3f6c86cbfba76339e4

Pith citing papers

No inbound Pith citation observations are available.