Pith. sign in

Paper Citation Record · LEDGER

Vision encoders should be image size agnostic and task driven

As of 22 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 1 inbound Pith citation observation for arXiv:2508.16317.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.16317 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:27:27.457813Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T00:18:06.989944Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T00:19:46.830591Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0c8306b9-7663-4232-8c24-e05fde8a6198 · outbound

This paper cites Multiple Object Recognition with Visual Attention.

Vision encoders should be image size agnostic and task driven Multiple Object Recognition with Visual Attention

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.307944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.307944Z digest=sha256:62ebe2dad8a5ccfa11f0518e8089ea193ae2d8f647237a6f70be02ca5ded1a46

Observation 927a8ded-f759-4f72-8949-2ea34848252d · outbound

This paper cites Recurrent memory transformer.Advances in Neural Information Processing Systems, 35:11079–11091, 2022.

Vision encoders should be image size agnostic and task driven Recurrent memory transformer.Advances in Neural Information Processing Systems, 35:11079–11091, 2022

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.312353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.312353Z digest=sha256:c85ab526b840683f6e93135313e98a496c4c253a7130f65e2c758246a7228c96

Observation 76f8050a-e11c-4d45-bdec-1cb159a25abc · outbound

This paper cites Unsupervised Foveal Vision Neural Networks with Top-Down Attention.

Vision encoders should be image size agnostic and task driven Unsupervised Foveal Vision Neural Networks with Top-Down Attention

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.315322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.315322Z digest=sha256:4fe07fd8ba802c3ad9130abfd89fd05a93e443cee8b5114692aee570cff59124

Observation e6dfe1e8-1524-4205-b161-7a7b20bedb88 · outbound

This paper cites End-to-end object detection with transformers.

Vision encoders should be image size agnostic and task driven End-to-end object detection with transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.319481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.319481Z digest=sha256:af110e22dcffb1116d64ee90a9aae9d6de106d7d0cfcd7c96d2e95540230449e

Observation b3aaa650-34d9-4903-a1ff-157b3166428c · outbound

This paper cites Emerging properties in self-supervised vision transformers.

Vision encoders should be image size agnostic and task driven Emerging properties in self-supervised vision transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.323004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.323004Z digest=sha256:a3057575c91dca5bb6631fa2a286b828c7ee50d86b4a3dc5a9932597f09436dc

Observation 8fd0df24-9575-4499-9c72-ffe64e97bda5 · outbound

This paper cites Crossvit: Cross-attention multi- scale vision transformer for image classification.

Vision encoders should be image size agnostic and task driven Crossvit: Cross-attention multi- scale vision transformer for image classification

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.325912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.325912Z digest=sha256:263140cbde8f5e7d273696b01d9c1a88cf74bd9bd6c89ad8f571ecee6f544e14

Observation f9068d13-2f5d-46f2-b32b-c763bb70f79b · outbound

This paper cites Twins: Revisiting the design of spatial attention in vision transformers.

Vision encoders should be image size agnostic and task driven Twins: Revisiting the design of spatial attention in vision transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.329215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.329215Z digest=sha256:dae553dff7ef1431b3b2f62a3a1576e496b38692decaf3d703ec73dc05b5b0ef

Observation 8cd95916-f42c-4cbf-adbc-918a8140f988 · outbound

This paper cites Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context.

Vision encoders should be image size agnostic and task driven Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.331915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.331915Z digest=sha256:f871ded4fbc33650bb5677da7844242509ec80097d5db7ae99b97c8fcac3b478

Observation 19075cf4-b951-4568-9071-75f6726d3fe8 · outbound

This paper cites Vision Transformers Need Registers.

Vision encoders should be image size agnostic and task driven Vision Transformers Need Registers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.335478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.335478Z digest=sha256:68dcc0ab7d486649703c16f9b521c98d6a87531c830e595881bdd9dbafde5707

Observation 6ca2e5bf-ce41-48f1-b0f7-7ad4fdc9a6dc · outbound

This paper cites Imagenet: A large- scale hierarchical image database.

Vision encoders should be image size agnostic and task driven Imagenet: A large- scale hierarchical image database

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.338506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.338506Z digest=sha256:f18fadca3636853040da8329619cad3aee55ba972afa3b2c81e73d533edaab21

Observation c1dce6c8-e76b-45a0-bffe-f61e0e53d208 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

Vision encoders should be image size agnostic and task driven Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.341479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.341479Z digest=sha256:bef46c83c582b56a73dfea0ed666391b0b263d574334538f8ca886f7898d5547

Observation aaf57278-b173-4168-b37f-92f2a593a554 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Vision encoders should be image size agnostic and task driven An image is worth 16x16 words: Transformers for image recognition at scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.344472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.344472Z digest=sha256:584599865843e51f1b03bde81cb10df662696e82202dd0b339a5c03d18d487d9

Observation cd9a9ecc-e463-4993-813d-b69c6f6f83d1 · outbound

This paper cites Multiscale vision transformers.

Vision encoders should be image size agnostic and task driven Multiscale vision transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.347313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.347313Z digest=sha256:cdc488eb334f876b3e15c121e8a06db1be3c17d522bd4ebee19312329081155c

Observation 9d6dada1-ce50-4a24-a437-ad3b1b399603 · outbound

This paper cites Levit: a vision transformer in convnet’s clothing for faster inference.

Vision encoders should be image size agnostic and task driven Levit: a vision transformer in convnet’s clothing for faster inference

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.350398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.350398Z digest=sha256:302de2c976409465df8d524e2d43e7b61f091ef0a7feaec6e79acf4134cdbf52

Observation c15b22e6-0c54-4e8e-aadd-8a96b0755bca · outbound

This paper cites GMAT: Global Memory Augmentation for Transformers.

Vision encoders should be image size agnostic and task driven GMAT: Global Memory Augmentation for Transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.353476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.353476Z digest=sha256:7d19093c96a922f69e248324772a0249f67bac48aaa25e46f4e52505322528ec

Observation 71b3601a-fb5c-4bde-9fdc-39f402495d78 · outbound

This paper cites Deep residual learning for image recognition.

Vision encoders should be image size agnostic and task driven Deep residual learning for image recognition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.356848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.356848Z digest=sha256:ce89daf3fe9a40875ff84591c47b67268269077ad9772eeb7e78ad6a7fb4835e

Observation 1fce41d9-bef1-4e0b-b1a0-a33155f5f907 · outbound

This paper cites Rethinking spatial dimensions of vision transformers.

Vision encoders should be image size agnostic and task driven Rethinking spatial dimensions of vision transformers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.359634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.359634Z digest=sha256:d84841d3f0683de46e6ea287401274d18d92b9af7bb7060c938938ea59572259

Observation af1c7ca2-575d-46dc-bc05-1b8b7a613231 · outbound

This paper cites Long short-term memory.

Vision encoders should be image size agnostic and task driven Long short-term memory

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.362955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.362955Z digest=sha256:e134565397f626b26d0bc6b50529184aeb3b0da820243aa719094b9b61f73a2e

Observation 4b1816c1-d9c0-45b2-bfd1-e280f749adea · outbound

This paper cites FoveaTer: Foveated Transformer for Image Classification.

Vision encoders should be image size agnostic and task driven FoveaTer: Foveated Transformer for Image Classification

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.365982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.365982Z digest=sha256:732d70115b70b849ae6a07911a1f8b6b6e6e8cd8914cc5d4447822a77a0769d1

Observation 98bb2074-9493-4d4c-8edb-ec024923197e · outbound

This paper cites Learning multiple layers of features from tiny images.

Vision encoders should be image size agnostic and task driven Learning multiple layers of features from tiny images

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.369577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.369577Z digest=sha256:553845e739d946b0f73146961f7df7216082d495d698c326448708217928ee3e

Observation 87a21bc2-22e8-4708-a579-c6890c405928 · outbound

This paper cites Imagenet classification with deep convolutional neural networks.

Vision encoders should be image size agnostic and task driven Imagenet classification with deep convolutional neural networks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.372280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.372280Z digest=sha256:e52cb9c01494ee47170e4586a5bc2539228a833abe0a5b789a4f5bdee164d402

Observation b21999c8-807e-4d34-afef-ecaf1ad56060 · outbound

This paper cites The shape of ai to come! Talk presented at the AI Action Summit 2025, February.

Vision encoders should be image size agnostic and task driven The shape of ai to come! Talk presented at the AI Action Summit 2025, February

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.375259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.375259Z digest=sha256:14aecd597dbe4d27fea4774a2738a5ccbbb9a4a439b676c6b3141f681069908d

Observation 874d4f48-181a-4bc9-991c-b30b33576cae · outbound

This paper cites Gradient-based learning applied to document recognition.

Vision encoders should be image size agnostic and task driven Gradient-based learning applied to document recognition

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.381369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.381369Z digest=sha256:05cd828a50868259c5611f1b679c4476c2e33ec440ea2dd19be65b1e1af62283

Observation 42c87409-889a-47eb-92b7-833f2d6179c3 · outbound

This paper cites Swin transformer v2: Scaling up capacity and resolution.

Vision encoders should be image size agnostic and task driven Swin transformer v2: Scaling up capacity and resolution

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.384475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.384475Z digest=sha256:521f21f49b6af6a10248789a4f53503668e1723c221bd46ea1ca4615822f2168

Observation 1797c047-40c3-4b18-84ee-8f7d92039e23 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Vision encoders should be image size agnostic and task driven Swin transformer: Hierarchical vision transformer using shifted windows

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.387249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.387249Z digest=sha256:8319566ef6938b9e89dcc498208c553474e78e4f3180265f855c0bf86c41b5b9

Observation e91a171c-f011-456f-81fa-4bff63703a76 · outbound

This paper cites Decoupled weight decay regularization.

Vision encoders should be image size agnostic and task driven Decoupled weight decay regularization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.390154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.390154Z digest=sha256:8424e9c763461516c1cc207cb281afa188c5fb316ddf08a82917962e5ba8a9c7

Observation 5c347d50-b1ec-403d-81f6-f15313cf9843 · outbound

This paper cites Biologically inspired deep learning model for efficient foveal-peripheral vision.

Vision encoders should be image size agnostic and task driven Biologically inspired deep learning model for efficient foveal-peripheral vision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.393193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.393193Z digest=sha256:db4686c9e738cbe894c464afdba6572f14f3122db69033622fd91663e1543044

Observation 9dd77574-90b7-4942-a527-3ca615b0a47b · outbound

This paper cites Implicit-zoo: A large-scale dataset of neural implicit functions for 2d images and 3d scenes, 2024.

Vision encoders should be image size agnostic and task driven Implicit-zoo: A large-scale dataset of neural implicit functions for 2d images and 3d scenes, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.396652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.396652Z digest=sha256:f835fc75c916cca7c08893afc21fcdc7a2f67a8dd65d414c287b5e3cc16e9629

Observation 1bc456a7-46b0-4cc3-bf0e-17a4f878922e · outbound

This paper cites Recurrent models of visual attention.

Vision encoders should be image size agnostic and task driven Recurrent models of visual attention

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.399521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.399521Z digest=sha256:1f12a9f4ceae700c53529e1ec9e23c869d9c35d5b12521cf9c854b3af7e2f1a1

Observation b48966fb-0c63-4651-83cb-0f33d89b9957 · outbound

This paper cites A focused backpropagation algorithm for temporal pattern recognition.

Vision encoders should be image size agnostic and task driven A focused backpropagation algorithm for temporal pattern recognition

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.402827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.402827Z digest=sha256:272a531178668eef7c7c0448de66cd8e427cb3e7e7be3a51ac50b96c2e32a26b

Observation 2435bb80-3ff5-4bbc-97c1-b56ba8a7efc9 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Vision encoders should be image size agnostic and task driven DINOv2: Learning Robust Visual Features without Supervision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.405864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.405864Z digest=sha256:181504aa4c2b4d00fd5f3a1daa2497ef319037222721c75c8cfb11c0e197cc3c

Observation 48d4fc6e-9e8f-40bd-acd3-ceb99b965205 · outbound

This paper cites Compressive Transformers for Long-Range Sequence Modelling.

Vision encoders should be image size agnostic and task driven Compressive Transformers for Long-Range Sequence Modelling

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.408925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.408925Z digest=sha256:8f0b7f63fc68f8266a4bfe0c7801463a09cfcade0d486ded07f295f3f9b74458

Observation 989c24ae-3e9a-4302-a5c8-f4ab83a36fac · outbound

This paper cites The utility driven dynamic error propagation network, volume 11.

Vision encoders should be image size agnostic and task driven The utility driven dynamic error propagation network, volume 11

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.412060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.412060Z digest=sha256:d348ad445eea6697c4475cd095e28a237ac30dd53d9c4177444e1f2080a11377

Observation 4a252d4e-56bf-4346-8abd-e2766b715858 · outbound

This paper cites Learning to generate artificial fovea trajectories for target detection.

Vision encoders should be image size agnostic and task driven Learning to generate artificial fovea trajectories for target detection

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.415649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.415649Z digest=sha256:092911c67a25d5baa7fba531fddeda9d0c7e866a5c75aa1b33d2b3e909339409

Observation 5179d160-0a2f-4669-b194-7892ebda47ac · outbound

This paper cites Proximal Policy Optimization Algorithms.

Vision encoders should be image size agnostic and task driven Proximal Policy Optimization Algorithms

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.418789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.418789Z digest=sha256:215b5b333d5199a7b376e209086e7749df38f0f63729674eb4f1c610cf82b05a

Observation b9a4db0a-45d7-440e-81c6-7c63814bfb5b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Vision encoders should be image size agnostic and task driven DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.421857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.421857Z digest=sha256:f26332aea427e673a580939a9bc7b0725b2fbb35cc77448e963e06546f061c7e

Observation f2f43293-0f23-40df-8516-3ec868624b0f · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Vision encoders should be image size agnostic and task driven Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.424852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.424852Z digest=sha256:8721571a426fd1286258d189c2912a7676f6a2d0902e19370b80da7c940f2af8

Observation a741a6e9-efd4-4ca6-92eb-4c377c255371 · outbound

This paper cites Going deeper with convolutions.

Vision encoders should be image size agnostic and task driven Going deeper with convolutions

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.427895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.427895Z digest=sha256:fe4041882f074670c0d313d0808b202ee8b13f937d396347e9dedd0d96b6109c

Observation b9b68948-deba-4e4a-8d6b-1cc4eccc121c · outbound

This paper cites Training data-efficient image transformers & distillation through attention.

Vision encoders should be image size agnostic and task driven Training data-efficient image transformers & distillation through attention

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.430683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.430683Z digest=sha256:6b06d8e5b9c4aa838b854f263142e89add79250838d90146ecfdba54fd8c3b90

Observation 06fa8854-08a9-4831-9663-404560e1f09a · outbound

This paper cites Fixing the train-test resolution discrepancy.

Vision encoders should be image size agnostic and task driven Fixing the train-test resolution discrepancy

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.433377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.433377Z digest=sha256:6e5d5ce728475c7dfb99230a8bbaddd95441131ce414ee1bfb52e89a67f27311

Observation c3ac5e17-72f8-48c0-ada2-10a8eee68f8a · outbound

This paper cites Attention is all you need.

Vision encoders should be image size agnostic and task driven Attention is all you need

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.436126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.436126Z digest=sha256:cdc699a79eeb63441ffd188c1fcf49ea6727b24e240661cea4937a9a71c512ff

Observation bf4ad697-0d87-4c00-881b-3044eda8f240 · outbound

This paper cites Pyramid vision transformer: A versatile backbone for dense prediction without convolutions.

Vision encoders should be image size agnostic and task driven Pyramid vision transformer: A versatile backbone for dense prediction without convolutions

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.438932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.438932Z digest=sha256:1cb565ff782ed23bdfd1a97b137d37a4b839af5c3d9fbcd97d58055423c7469c

Observation 165d1bf8-f90b-421c-a49b-418a0760a1ee · outbound

This paper cites Generalization of backpropagation with application to a recurrent gas market model.

Vision encoders should be image size agnostic and task driven Generalization of backpropagation with application to a recurrent gas market model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.442174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.442174Z digest=sha256:7d1fde31ab4e18dc7cec6edf5d10cec317edf11b3b79ee12a16cd2204517327c

Observation 3b8af0ec-c7a8-458f-97a2-51a01d46ca96 · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist reinforce- ment learning.

Vision encoders should be image size agnostic and task driven Simple statistical gradient-following algorithms for connectionist reinforce- ment learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.445077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.445077Z digest=sha256:2fae0d32f35cf24b798c2395df39dea75fc0d5c2bf0002f0dbb0f5ebb41d7d17

Observation 3f04c197-1320-4127-b732-0d3446cc14c1 · outbound

This paper cites Memformer: A Memory-Augmented Transformer for Sequence Modeling.

Vision encoders should be image size agnostic and task driven Memformer: A Memory-Augmented Transformer for Sequence Modeling

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.447770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.447770Z digest=sha256:e4778c78fd4aaf73e922f9f7f5063b8a96963eadfa2c86c1eca8a76d72efca3c

Observation 905ebcd7-5c79-456f-aecd-eada393113de · outbound

This paper cites Neural fields in visual computing and beyond.

Vision encoders should be image size agnostic and task driven Neural fields in visual computing and beyond

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.450958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.450958Z digest=sha256:432bde6367862aca5328b989639f658f26779cad3fcb83f98fb870578046b684

Observation afac8872-93a0-4c4c-8016-5053ab299a07 · outbound

This paper cites Focal Self-attention for Local-Global Interactions in Vision Transformers.

Vision encoders should be image size agnostic and task driven Focal Self-attention for Local-Global Interactions in Vision Transformers

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.454271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.454271Z digest=sha256:760fcea47372b634c49ff023c0731eb087f849fc0229ff473b877469149326cd

Observation 8ffb39f8-4169-4983-87c7-b5a4406cbcc1 · outbound

This paper cites mixup: Beyond Empirical Risk Minimization.

Vision encoders should be image size agnostic and task driven mixup: Beyond Empirical Risk Minimization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.457813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.457813Z digest=sha256:d18b4d20acbb2e2401e7065f82a2cb853a260d38292bf3b5aa62c91583d745c2

Observation 587bb495-b9dd-44c7-bd01-1131514e119e · outbound

This paper cites an unresolved cited work.

Vision encoders should be image size agnostic and task driven Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.378175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.378175Z digest=sha256:b40fdc3125b81acbef2eecfd9480da8531ef033d5abc7bc836c99546275f1efb

Pith citing papers

Observation 5adc7cf5-6218-46ef-94a0-ab5649bd6a69 · inbound

Self-supervised pretraining for an iterative image size agnostic vision transformer cites this paper.

Self-supervised pretraining for an iterative image size agnostic vision transformer Vision encoders should be image size agnostic and task driven

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:19:46.831894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T00:18:06.989944Z digest=sha256:75ffa6e8d1f5fa6e408cdd610a63275a6d13201072ab7ea0df49549aeee666f9