Pith. sign in

Paper Citation Record · LEDGER

Multi-Token Enhancing for Vision Representation Learning

As of 14 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2411.15787.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.15787 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:59:36.446502Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact1
  • verified fuzzy52
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c0c18bca-5231-4214-b82b-0e51ceb643c7 · outbound

This paper cites Asano, Christian Rupprecht, and Andrea Vedaldi.

Multi-Token Enhancing for Vision Representation Learning Asano, Christian Rupprecht, and Andrea Vedaldi

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.518865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.162993Z digest=sha256:752eaef4ff0aa44fa756fa2023a6463c6286865c674fcba95821eba90f92e470

Observation fb9c483a-0a2e-4aa0-95d1-237f9ae77cf2 · outbound

This paper cites BEit: BERT pre-training of image transformers.

Multi-Token Enhancing for Vision Representation Learning BEit: BERT pre-training of image transformers

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.507105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.167722Z digest=sha256:1a0ba7bca7dd228b772494a74303b20cc76d9440b12e63fbdb3472abaa22e09b

Observation 81363a80-e5b7-4c4c-93aa-0416cae652cb · outbound

This paper cites Cascade r-cnn: Delving into high quality object detection.

Multi-Token Enhancing for Vision Representation Learning Cascade r-cnn: Delving into high quality object detection

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.493756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.171857Z digest=sha256:11124f835aec5287152198ca03b7cae8cd77bccf356d22e8dec07d03e738c9ff

Observation 17b47991-8b28-4e0b-9141-b7f273e49974 · outbound

This paper cites Unsupervised learn- ing of visual features by contrasting cluster assignments.

Multi-Token Enhancing for Vision Representation Learning Unsupervised learn- ing of visual features by contrasting cluster assignments

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.482258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.175892Z digest=sha256:e843d2a16bff9f3d9c8483e6927da99c4d0ec8f72831d32302b28f43c7db25f5

Observation 0eb08d6c-8a62-4bb8-87af-39508d7990a7 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

Multi-Token Enhancing for Vision Representation Learning Emerg- ing properties in self-supervised vision transformers

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.472053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.180754Z digest=sha256:16b23e8a2084b181e0aa127d893eb0cdfdeeb328f7057f9fc1511c21aaa5f6b1

Observation e5b8be40-8107-4031-972d-4d6a660ce06f · outbound

This paper cites Mixed autoencoder for self-supervised visual representation learning.

Multi-Token Enhancing for Vision Representation Learning Mixed autoencoder for self-supervised visual representation learning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.459500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.185131Z digest=sha256:df72dd5412fefda1da270727f15a9911666ae7948c679e9f472f5b573b17e7f8

Observation bae4412f-d772-4e0f-80d8-54654691c16e · outbound

This paper cites A simple framework for contrastive learning of visual representations.

Multi-Token Enhancing for Vision Representation Learning A simple framework for contrastive learning of visual representations

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.446793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.189448Z digest=sha256:a7c19fd18897fdbd2a505ce044929112a93774f0747c47d7a6d58342e21f7b6f

Observation 4987cec4-5022-4483-86b2-723da00bcfe6 · outbound

This paper cites Exploring simple siamese rep- resentation learning.

Multi-Token Enhancing for Vision Representation Learning Exploring simple siamese rep- resentation learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.434514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.193352Z digest=sha256:56ec424dc58165c4b0d8e6e1ca4c515036c3980e6c97b160eb0738d609611141

Observation 7b450849-3c64-41c3-b76d-bf7cd4ad6571 · outbound

This paper cites An empiri- cal study of training self-supervised vision transformers.

Multi-Token Enhancing for Vision Representation Learning An empiri- cal study of training self-supervised vision transformers

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.423402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.197964Z digest=sha256:cb8e29597ce07b98bb96c0120c0a8a9545c6c1673f4fa8336cc3e1dc6e5d2f39

Observation 1c7a8955-0b96-42f2-b548-25581cf233fe · outbound

This paper cites Convit: Improving vision transformers with soft convolutional inductive biases.

Multi-Token Enhancing for Vision Representation Learning Convit: Improving vision transformers with soft convolutional inductive biases

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.410405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.202247Z digest=sha256:0088cd35571b03f1769a05b5ea2d334f82a255ef7cf0778968b523226de39dc1

Observation d5345f7c-f55e-466c-8da6-f9b683ef5111 · outbound

This paper cites Ensemble methods in machine learn- ing.

Multi-Token Enhancing for Vision Representation Learning Ensemble methods in machine learn- ing

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.397794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.205999Z digest=sha256:e99657eab6ad592152cff373f56e367b8709de1a9e9daed62a3781fc9c350a1f

Observation 7c14abbd-a1ad-4f6c-9577-b8b7051979d7 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Multi-Token Enhancing for Vision Representation Learning An image is worth 16x16 words: Transformers for image recognition at scale

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.385594Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.210402Z digest=sha256:aa315ac1131aded87a749dcd2cc160e9990e082e53dbb1bd094d3a1d9bdc4728

Observation 59e7332b-e982-491d-9d09-8a3b37c71d86 · outbound

This paper cites Whitening for self-supervised representation learning.

Multi-Token Enhancing for Vision Representation Learning Whitening for self-supervised representation learning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.371780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.215043Z digest=sha256:a2d6cd54b15473a7ef5aca2929d5390cc7c5d1899ca1887b0de83ebf38350d34

Observation f324fea3-b88c-4cbd-95d9-b320c9086d2e · outbound

This paper cites Seed: Self-supervised dis- tillation for visual representation.

Multi-Token Enhancing for Vision Representation Learning Seed: Self-supervised dis- tillation for visual representation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.358369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.218838Z digest=sha256:23f6060c400f7c70aec8fd0871a496c3096fd3a4a4626fd9efae1005409c614c

Observation f8a93134-6acc-4218-9dec-85847ba44a80 · outbound

This paper cites Evolved part masking for self-supervised learning.

Multi-Token Enhancing for Vision Representation Learning Evolved part masking for self-supervised learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.346320Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.222715Z digest=sha256:01d6b43cc375b36dc7cde1f5ed2d84948f743f39aa079c055b381d62aab34291

Observation 1df1351e-b688-472e-8008-58ad63218e85 · outbound

This paper cites Ganaie, Minghui Hu, A.K.

Multi-Token Enhancing for Vision Representation Learning Ganaie, Minghui Hu, A.K

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.330615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.227572Z digest=sha256:57f1a83f26a911dadf293492d241f10367b0ebfffd6dcb1167b6e0e29cd6cd7d

Observation cc5b256c-db5d-4777-a2d4-2980d658d30e · outbound

This paper cites Large-scale un- supervised semantic segmentation.

Multi-Token Enhancing for Vision Representation Learning Large-scale un- supervised semantic segmentation

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.316715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.232442Z digest=sha256:257a821a5ca39c413a61253114efe6a809356ba6e16a58fd2f1d0fd9cd38ab48

Observation 87d73ed9-4a22-4f77-a574-d3764fe50431 · outbound

This paper cites Richemond, Elena Buchatskaya, Carl Doersch, Bernardo ´Avila Pires, Zhaohan Guo, Moham- mad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, R´emi Munos, and Michal Valko.

Multi-Token Enhancing for Vision Representation Learning Richemond, Elena Buchatskaya, Carl Doersch, Bernardo ´Avila Pires, Zhaohan Guo, Moham- mad Gheshlaghi Azar, Bilal Piot, Koray Kavukcuoglu, R´emi Munos, and Michal Valko

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.245865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.236321Z digest=sha256:6d008a6054fbea88251a3c39a2f1d44d8891928c08f7175a6ac4bcbc9501564f

Observation 0007e657-5c59-4277-9ac2-129af3be0f57 · outbound

This paper cites Visual Attention Network.

Multi-Token Enhancing for Vision Representation Learning Visual Attention Network

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T13:59:36.240463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:59:36.240463Z digest=sha256:dae326caff8b41855e4dc24e04c701d9a0b1b66a2938a26b67f3140d09c40488

Observation 202ee373-4dee-40f8-b622-5c150b8e821b · outbound

This paper cites Hansen and P.

Multi-Token Enhancing for Vision Representation Learning Hansen and P

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.177669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.245057Z digest=sha256:f3eafff6b960ffa094b09624001ff09e6d5b3c00960cb896a81004dbbcde20b2

Observation ec2b9bb3-731d-42c2-8a6a-304c77232e7a · outbound

This paper cites Training independent subnetworks for robust prediction.

Multi-Token Enhancing for Vision Representation Learning Training independent subnetworks for robust prediction

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.165226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.249656Z digest=sha256:95c6f677ef255796e9014e96f8efad22cd7045cacee81864c35e15fa3ef05367

Observation 6f0a9f0a-c6a9-420a-b565-579b2f1db22f · outbound

This paper cites Deep residual learning for image recognition.

Multi-Token Enhancing for Vision Representation Learning Deep residual learning for image recognition

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T13:59:36.253645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:59:36.253645Z digest=sha256:1d4063ff105e620c50cafa0234b5f8681859f72b1bc7d2d600348588f85da66a

Observation 7307fe2b-58c4-4e12-be64-f897cce54176 · outbound

This paper cites Momentum contrast for unsupervised visual rep- resentation learning.

Multi-Token Enhancing for Vision Representation Learning Momentum contrast for unsupervised visual rep- resentation learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.144942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.257724Z digest=sha256:e5efe1a178245d14533ef3a2a2652ef018014a19a6a83dfd71ea17f246fe21a7

Observation fec79ee1-7fa5-4f2b-9c1b-320cee2d22b1 · outbound

This paper cites Masked autoencoders are scalable vision learners.

Multi-Token Enhancing for Vision Representation Learning Masked autoencoders are scalable vision learners

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.132864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.261961Z digest=sha256:deac29dc496ce9b1027c9609bd09836d9046f8733cb6827deaa885e320bfe3d6

Observation ae5b7a43-a9ee-4420-8f01-4c65f282ca8a · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Multi-Token Enhancing for Vision Representation Learning Distilling the Knowledge in a Neural Network

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T13:59:36.265938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:59:36.265938Z digest=sha256:e710d0578fc53ca30e56e380c97509097b45c7edeabc4f724d008baa66858552

Observation f91d5ffd-e3db-4dee-ac1b-1705c7950631 · outbound

This paper cites Conv2Former: A Simple Transformer-Style ConvNet for Visual Recognition.

Multi-Token Enhancing for Vision Representation Learning Conv2Former: A Simple Transformer-Style ConvNet for Visual Recognition

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T13:59:36.270510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:59:36.270510Z digest=sha256:df8c88ac61c1acf3d7dae500f298775209eed52d146fc6adaaf01ee4f6318fbd

Observation 2c8310df-939a-47f3-ba8c-54e875f15300 · outbound

This paper cites Averaging Weights Leads to Wider Optima and Better Generalization.

Multi-Token Enhancing for Vision Representation Learning Averaging Weights Leads to Wider Optima and Better Generalization

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T13:59:36.274517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:59:36.274517Z digest=sha256:e57cac3491b6097c529cb189af3161de85a52a59f9d3425c4777dcc77df2b7a6

Observation 008a7f77-e186-44e7-8e9b-595500a28f5c · outbound

This paper cites Similarity of neural network representa- tions revisited.

Multi-Token Enhancing for Vision Representation Learning Similarity of neural network representa- tions revisited

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.120725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.279665Z digest=sha256:f5b446121476402e2aafb8477b2d2894b35b91f3e019138740bf6d53d0e273da

Observation 6b864517-9da5-4755-9925-88bcbaf7e41c · outbound

This paper cites Fractalnet: Ultra-deep neural networks without residuals.

Multi-Token Enhancing for Vision Representation Learning Fractalnet: Ultra-deep neural networks without residuals

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.108217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.283495Z digest=sha256:8b63b438b94cb43686240c79e6a7ada06599d09d3530325ddd340dbecb6aa26e

Observation cf0c168f-5f8b-47c9-b486-193d522a2456 · outbound

This paper cites Why M Heads are Better than One: Training a Diverse Ensemble of Deep Networks.

Multi-Token Enhancing for Vision Representation Learning Why M Heads are Better than One: Training a Diverse Ensemble of Deep Networks

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T13:59:36.287012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:59:36.287012Z digest=sha256:e66dfc0ca410521c058042d5f19e89759fd4e600643bee1bad63137cfa3d90c7

Observation 49c894ea-0fc9-4cbd-be3b-b64fc9434d9c · outbound

This paper cites Sere: Exploring feature self-relation for self-supervised trans- former.

Multi-Token Enhancing for Vision Representation Learning Sere: Exploring feature self-relation for self-supervised trans- former

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.096601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.291248Z digest=sha256:2aba311a8652cf36f3326a33d792f04aa325d95e236292e2a29334641c48dda0

Observation 06f74ad0-5d71-4f4f-8297-eb3141445aee · outbound

This paper cites Enhancing representa- tions through heterogeneous self-supervised learning.

Multi-Token Enhancing for Vision Representation Learning Enhancing representa- tions through heterogeneous self-supervised learning

Reference 32

Resolution
verified exact
raw_fallback, observed 2026-08-12T13:59:36.630796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.294927Z digest=sha256:b24f777acea1ef74226865f250ec0d01b2443bdf1ad34f4d0d6d3572d91f30fc

Observation 7de05d14-dcb8-4fab-be4f-7cea56036150 · outbound

This paper cites Microsoft coco: Common objects in context.

Multi-Token Enhancing for Vision Representation Learning Microsoft coco: Common objects in context

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.084032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.298601Z digest=sha256:052bf826ff36531e1ed25176f6173b46e69f5ec4071b317d91dfc3aebe0eb9de

Observation 1e928978-8500-4391-873d-2a7a81ef28c1 · outbound

This paper cites Swin trans- former: Hierarchical vision transformer using shifted win- dows.

Multi-Token Enhancing for Vision Representation Learning Swin trans- former: Hierarchical vision transformer using shifted win- dows

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.071789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.302744Z digest=sha256:0c136f04e81086b5df88fe36c8807de3839438b0712ee4012852551228294b6b

Observation cd04de15-3d21-4fa2-bd93-99c50f2f8caa · outbound

This paper cites A convnet for the 2020s.

Multi-Token Enhancing for Vision Representation Learning A convnet for the 2020s

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.058450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.306788Z digest=sha256:6a2c8553382345f2071f55914a37d1338c479c9882ac09c5dc26c3b20baaf7f0

Observation fe97687c-bdf2-4f79-a15a-b052e55841bd · outbound

This paper cites Decoupled weight decay regularization.

Multi-Token Enhancing for Vision Representation Learning Decoupled weight decay regularization

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.044230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.310991Z digest=sha256:289ed47c9c1a7b4ee0cd6ed119cd832431191f2baf7f6affbad86b5d2f92bbb7

Observation 81876898-bc47-4c23-9244-b503f4ac24ce · outbound

This paper cites Representation uncertainty in self-supervised learning as variational inference.

Multi-Token Enhancing for Vision Representation Learning Representation uncertainty in self-supervised learning as variational inference

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.030120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.314763Z digest=sha256:01669e28902af9161329e1373d01ac4379ab5df005f79fbd941a440cf96c423c

Observation 474571af-d00b-4c85-9f9e-7093000d1393 · outbound

This paper cites Simreg: Regression as a sim- ple yet effective tool for self-supervised knowledge distilla- tion.

Multi-Token Enhancing for Vision Representation Learning Simreg: Regression as a sim- ple yet effective tool for self-supervised knowledge distilla- tion

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:37.017322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.318518Z digest=sha256:a5bcc067a0642207ae84e17964599319d13851f82fdddc77eef9606297a26316

Observation 58b883a7-a95f-485c-965d-8b9157c8bdb5 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Multi-Token Enhancing for Vision Representation Learning Representation Learning with Contrastive Predictive Coding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T13:59:36.323425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:59:36.323425Z digest=sha256:e8934121d11259f9b08d535d038b3ab8e31f38a063a295d481bb50995c02e7d2

Observation e92f2de2-92a5-4f0b-90d8-4b6528726863 · outbound

This paper cites Imagenet large scale visual recognition challenge.

Multi-Token Enhancing for Vision Representation Learning Imagenet large scale visual recognition challenge

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T13:59:36.327999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:59:36.327999Z digest=sha256:1df7a99f3e0ba9c9ee38dc7445ad3439a7d3039746d06890ab5e93411183c1b2

Observation 30511b61-22e1-4182-bfb2-190f132ed2ca · outbound

This paper cites Learning common rationale to improve self-supervised rep- resentation for fine-grained visual recognition problems.

Multi-Token Enhancing for Vision Representation Learning Learning common rationale to improve self-supervised rep- resentation for fine-grained visual recognition problems

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.995993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.332762Z digest=sha256:fc86af9a8adbaa6dc42cb7759a37b77e333a7f56d348ca213ef1113c0954fae0

Observation 6c11bba4-cc9b-4d59-b780-d51ed92533c1 · outbound

This paper cites an unresolved cited work.

Multi-Token Enhancing for Vision Representation Learning Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-12T13:59:36.984006Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.337204Z digest=sha256:4f4dcc81136753e8eb8d8532f98033a28c35fe7b531ddef32f24bcd121cee20e

Observation 1397c6fb-8ffa-482b-9707-6664ac7e44ae · outbound

This paper cites Multi- mode online knowledge distillation for self-supervised visual representation learning.

Multi-Token Enhancing for Vision Representation Learning Multi- mode online knowledge distillation for self-supervised visual representation learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.971471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.341166Z digest=sha256:71b7dd2fb2a4479d225f905860dbc7ad626d66ff8aeafb27d103688d6213a558

Observation 39d4be1e-013c-4f9a-ac70-79b84717ee06 · outbound

This paper cites Semantics-consistent feature search for self-supervised visual representation learning.

Multi-Token Enhancing for Vision Representation Learning Semantics-consistent feature search for self-supervised visual representation learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.958797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.345473Z digest=sha256:6a3711f652259c0f6fa155d3a83349b750ee4528eb478805d711f46fc7d81932

Observation 01ba14a2-f08d-40cd-be01-5faafd0dab63 · outbound

This paper cites Dropout: A simple way to prevent neural networks from overfitting.

Multi-Token Enhancing for Vision Representation Learning Dropout: A simple way to prevent neural networks from overfitting

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.946527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.349178Z digest=sha256:cd745c316bd0404294d5ef4b7df28c099982cd7c576501fdce780ce1ce2d54fd

Observation 0ffb2b9b-743a-455d-9305-745c44fb2bd1 · outbound

This paper cites Siamese image modeling for self-supervised vision represen- tation learning.

Multi-Token Enhancing for Vision Representation Learning Siamese image modeling for self-supervised vision represen- tation learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.934827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.353103Z digest=sha256:d368f121bdbdf621d8903b59d9199d3e42d41e1278652358219c0a474fca7708

Observation 5160aa41-7918-4047-808c-b87342537b04 · outbound

This paper cites Un- derstanding self-supervised learning dynamics without con- trastive pairs.

Multi-Token Enhancing for Vision Representation Learning Un- derstanding self-supervised learning dynamics without con- trastive pairs

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.921604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.357880Z digest=sha256:b2bbfa8e5f6537f30daf978a6c4e0a3fa4a105e696d85ce14699f4fef7440720

Observation d4306c5c-5350-40be-a1a1-e9d2f647d87e · outbound

This paper cites The inaturalist species classification and de- tection dataset.

Multi-Token Enhancing for Vision Representation Learning The inaturalist species classification and de- tection dataset

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.908604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.362071Z digest=sha256:2fd6eb7656b480ebdfaba1677d5c99b598de7e0863fc0dc43a5f73591604e79b

Observation fd3cc78a-db4f-4275-9230-f92fded43f8e · outbound

This paper cites Pyramid vision transformer: A versatile backbone for dense prediction without convolutions.

Multi-Token Enhancing for Vision Representation Learning Pyramid vision transformer: A versatile backbone for dense prediction without convolutions

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T13:59:36.365916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:59:36.365916Z digest=sha256:5a75e5ee35a51777c1abc4a01040576150ed71d5cfcadb25aaa52e0e54fec5ef

Observation c21843a7-1cc4-4fb6-b6d0-a03cbfb301fd · outbound

This paper cites Dense contrastive learning for self-supervised visual pre-training.

Multi-Token Enhancing for Vision Representation Learning Dense contrastive learning for self-supervised visual pre-training

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.888226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.370044Z digest=sha256:a22598dcd71e6df45cedb2c98d49b537b98dae7b423b413d273a55325965ec42

Observation 0b6c0f23-d613-4c04-a131-7ab4fed5a8bf · outbound

This paper cites Masked Feature Prediction for Self-Supervised Visual Pre-Training.

Multi-Token Enhancing for Vision Representation Learning Masked Feature Prediction for Self-Supervised Visual Pre-Training

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T13:59:36.374197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:59:36.374197Z digest=sha256:bcb13b017967d50cbcfba1eb3ddc5c09021ea0e8c67e7f4391c2bccaa30a9940

Observation 8f63cfc6-8ea1-4a29-8c9c-fbd761f90e77 · outbound

This paper cites Batchensemble: an alternative approach to efficient ensemble and lifelong learning.

Multi-Token Enhancing for Vision Representation Learning Batchensemble: an alternative approach to efficient ensemble and lifelong learning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.877196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.379007Z digest=sha256:a57dde76f657f6e66baf8d35a28a51f8f56eeb6789700543fae08087acfed714

Observation fd8cabfb-f044-4e51-ae64-d8cf6949b0e4 · outbound

This paper cites Con- vnext v2: Co-designing and scaling convnets with masked autoencoders.

Multi-Token Enhancing for Vision Representation Learning Con- vnext v2: Co-designing and scaling convnets with masked autoencoders

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.865107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.383605Z digest=sha256:0e6632d3af1c82da0b9c83c0c5d3a190e8580366591b9e3bdeb75976b6553a0c

Observation 3f8ca473-1b5c-482f-ad48-83bae95b1767 · outbound

This paper cites Cvt: Introducing con- volutions to vision transformers.

Multi-Token Enhancing for Vision Representation Learning Cvt: Introducing con- volutions to vision transformers

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.851238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.387576Z digest=sha256:d6878c258878164e1698606e5ec532b97205df4169b2110452b05f6cf4ebc805

Observation d1933b00-fa25-46b6-9444-27b306a7c594 · outbound

This paper cites P2T: Pyramid pooling transformer for scene understanding.

Multi-Token Enhancing for Vision Representation Learning P2T: Pyramid pooling transformer for scene understanding

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.837973Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.391428Z digest=sha256:b80a6c34655a879c3881685ddac768b28897ae24690c8be26a5f755bd4997e8c

Observation 85ef4907-c63c-486a-81d0-55e8f7d0a51e · outbound

This paper cites Unified perceptual parsing for scene understand- ing.

Multi-Token Enhancing for Vision Representation Learning Unified perceptual parsing for scene understand- ing

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.824794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.395599Z digest=sha256:807de00d30a568e84d5396370ef410c0fca9a857708a004ea7bd3a263383977d

Observation b3a14ab8-cde0-4d9a-b440-18f5bb7d8396 · outbound

This paper cites Detco: Unsu- pervised contrastive learning for object detection.

Multi-Token Enhancing for Vision Representation Learning Detco: Unsu- pervised contrastive learning for object detection

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.813107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.399528Z digest=sha256:da7782b54605653a162e3d12842290816ae702e309075dd6ee4b9db1901a333e

Observation e9c44f02-2f1e-4160-ba2c-222945df8fcb · outbound

This paper cites Self-Supervised Learning with Swin Transformers.

Multi-Token Enhancing for Vision Representation Learning Self-Supervised Learning with Swin Transformers

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T13:59:36.404002Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:59:36.404002Z digest=sha256:1497a2f50e80a9032846bbca4fa8cd0aafd464ee61ba05a8e95ee225e422eed0

Observation 586afe26-a476-4a99-b125-432adefc9de2 · outbound

This paper cites Propagate yourself: Exploring pixel-level consistency for unsupervised visual representation learning.

Multi-Token Enhancing for Vision Representation Learning Propagate yourself: Exploring pixel-level consistency for unsupervised visual representation learning

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.801568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.408074Z digest=sha256:c14db2ce75b15e63d3474874e6518593e04bcae3285b8a691304117e785941dc

Observation 2262c200-fa96-41d5-9fda-4b9c8c84af89 · outbound

This paper cites Simmim: A simple framework for masked image modeling.

Multi-Token Enhancing for Vision Representation Learning Simmim: A simple framework for masked image modeling

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T13:59:36.412172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:59:36.412172Z digest=sha256:dd4b88eaab86316ae1fca05f9799b502b6abcc34d504a9ee06fe9022b00fd57d

Observation 6aab7386-a01d-4f1e-a02d-c4deaeeaa838 · outbound

This paper cites Bag of instances aggregation boosts self-supervised distillation.

Multi-Token Enhancing for Vision Representation Learning Bag of instances aggregation boosts self-supervised distillation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.782035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.416426Z digest=sha256:bf81839a9897c912ba43b8cf8e1821df352e70884af8bff59ecd67d5f830803d

Observation d73d5289-4ec1-4238-b808-479c9960f467 · outbound

This paper cites Joint unsuper- vised learning of deep representations and image clusters.

Multi-Token Enhancing for Vision Representation Learning Joint unsuper- vised learning of deep representations and image clusters

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T13:59:36.420129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:59:36.420129Z digest=sha256:464de2569f83fd38a25d1ed15b74e1e91a0dfcf63157389b26b53890dfa074a5

Observation b611e978-ca13-4602-a3a5-b61a99a31d0e · outbound

This paper cites Decoupled contrastive learning.

Multi-Token Enhancing for Vision Representation Learning Decoupled contrastive learning

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.759788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.424067Z digest=sha256:71cc1cf50b4f14fd8f0f12e64f69a965e060d50ac70836aa82ebb1e5b6b11727

Observation da304176-fd51-44a5-8855-0610b6db12f6 · outbound

This paper cites Online deep clustering for unsupervised representation learning.

Multi-Token Enhancing for Vision Representation Learning Online deep clustering for unsupervised representation learning

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.745952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.428108Z digest=sha256:f478d37bb6d5a887350b699bee5da81c8a0320b1ad4fe2d9de49709f730828cb

Observation d104ba81-a887-4d37-9d54-85f9fd059ddd · outbound

This paper cites Scene parsing through ade20k dataset.

Multi-Token Enhancing for Vision Representation Learning Scene parsing through ade20k dataset

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.732631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.432014Z digest=sha256:b5b77fc5ef1f1ef39906eb61747e6511eba5678a1bbe52d48340ee0448539978

Observation 59d1a300-238d-4504-bd43-6c4df54e38fb · outbound

This paper cites ibot: Image bert pre-training with online tokenizer.

Multi-Token Enhancing for Vision Representation Learning ibot: Image bert pre-training with online tokenizer

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.719476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.436927Z digest=sha256:811ea322a9833ea1b445b8ddbd99585a1d179aa4de62fc3ee583df012299ce49

Observation 3209ecb5-4892-4949-a2d5-7eb2ad445a7b · outbound

This paper cites Mugs: A Multi-Granular Self-Supervised Learning Framework.

Multi-Token Enhancing for Vision Representation Learning Mugs: A Multi-Granular Self-Supervised Learning Framework

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T13:59:36.441292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:59:36.441292Z digest=sha256:b43b4e7533e7aacf5f08c357d3f703873cba9fe7b31b54d95dbf23d97691f579

Observation a1d707bd-7ebc-4ab5-baa6-8d2b9186b074 · outbound

This paper cites Multi-label self- supervised learning with scene images.

Multi-Token Enhancing for Vision Representation Learning Multi-label self- supervised learning with scene images

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:59:36.707319Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T13:59:36.446502Z digest=sha256:f0470acbe5565d42fc3b872f66c5825914d20c4ac1049699aa256cfe8febb116

Pith citing papers

No inbound Pith citation observations are available.