Pith. sign in

Paper Citation Record · LEDGER

Vision encoders should be image size agnostic and task driven

As of 7 August 2026, this Paper Citation Record lists 49 of 49 outbound references and 1 inbound Pith citation observation for arXiv:2508.16317.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.16317 v1

Coverage vector

measured 49 of 49 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T17:27:27.457813Z

measured 50 of 50 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T00:18:06.989944Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T00:19:46.830591Z

Reference resolution

49 of 49 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved49
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 0c8306b9-7663-4232-8c24-e05fde8a6198 · outbound

This paper cites Multiple Object Recognition with Visual Attention.

Vision encoders should be image size agnostic and task driven Multiple Object Recognition with Visual Attention

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.307944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.307944Z digest=sha256:30d4e5ee5eec8a732c7fe4aa002fd3876b76b2fa1d7aefeaa6b6999e23a014bd

Observation 927a8ded-f759-4f72-8949-2ea34848252d · outbound

This paper cites Recurrent memory transformer.Advances in Neural Information Processing Systems, 35:11079–11091, 2022.

Vision encoders should be image size agnostic and task driven Recurrent memory transformer.Advances in Neural Information Processing Systems, 35:11079–11091, 2022

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.312353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.312353Z digest=sha256:25dbfebc71818ad45c0ed232f7758d12b3a96a06bdd7f5624bf58a5c54f46758

Observation 76f8050a-e11c-4d45-bdec-1cb159a25abc · outbound

This paper cites Unsupervised Foveal Vision Neural Networks with Top-Down Attention.

Vision encoders should be image size agnostic and task driven Unsupervised Foveal Vision Neural Networks with Top-Down Attention

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.315322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.315322Z digest=sha256:d26570c68b1db2cef54e24c0e312473c99f82399ca8e0eac065325a6bec25fcb

Observation e6dfe1e8-1524-4205-b161-7a7b20bedb88 · outbound

This paper cites End-to-end object detection with transformers.

Vision encoders should be image size agnostic and task driven End-to-end object detection with transformers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.319481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.319481Z digest=sha256:16ab1237a92cd3de02120b4c293c8d7039f40ab35e75a677fbd16647627fad16

Observation b3aaa650-34d9-4903-a1ff-157b3166428c · outbound

This paper cites Emerging properties in self-supervised vision transformers.

Vision encoders should be image size agnostic and task driven Emerging properties in self-supervised vision transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.323004Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.323004Z digest=sha256:d4a393bef5c7e1c7bf293b5956f7e31536af80473f379b4e7c8f5c11cf0b3437

Observation 8fd0df24-9575-4499-9c72-ffe64e97bda5 · outbound

This paper cites Crossvit: Cross-attention multi- scale vision transformer for image classification.

Vision encoders should be image size agnostic and task driven Crossvit: Cross-attention multi- scale vision transformer for image classification

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.325912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.325912Z digest=sha256:18602a620f0cb178eb665319f757a793384909acc46a4793a552710e31ff10a6

Observation f9068d13-2f5d-46f2-b32b-c763bb70f79b · outbound

This paper cites Twins: Revisiting the design of spatial attention in vision transformers.

Vision encoders should be image size agnostic and task driven Twins: Revisiting the design of spatial attention in vision transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.329215Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.329215Z digest=sha256:6a06bdcda0fb6c6a476ae540e82aa10fa529f4df1455e4630e092a1263bb04c7

Observation 8cd95916-f42c-4cbf-adbc-918a8140f988 · outbound

This paper cites Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context.

Vision encoders should be image size agnostic and task driven Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.331915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.331915Z digest=sha256:587b1820dcb7f5dbe74156fcf9e2df8db7ca29c9c9a7e8478e7502a3d8087a55

Observation 19075cf4-b951-4568-9071-75f6726d3fe8 · outbound

This paper cites Vision Transformers Need Registers.

Vision encoders should be image size agnostic and task driven Vision Transformers Need Registers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.335478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.335478Z digest=sha256:b3b47d61817502b26592b844bf8943c1e552cc5e0b1fda5753782327f9134c3b

Observation 6ca2e5bf-ce41-48f1-b0f7-7ad4fdc9a6dc · outbound

This paper cites Imagenet: A large- scale hierarchical image database.

Vision encoders should be image size agnostic and task driven Imagenet: A large- scale hierarchical image database

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.338506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.338506Z digest=sha256:ac9b037d4aebd2258620e8238927f600abbdc7974dd1987e9a753858954437ce

Observation c1dce6c8-e76b-45a0-bffe-f61e0e53d208 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding.

Vision encoders should be image size agnostic and task driven Bert: Pre-training of deep bidirectional transformers for language understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.341479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.341479Z digest=sha256:4f36bb072ed97181e69884dea1e069f005d4e5ae6b1e7055f5043307415438fd

Observation aaf57278-b173-4168-b37f-92f2a593a554 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Vision encoders should be image size agnostic and task driven An image is worth 16x16 words: Transformers for image recognition at scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.344472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.344472Z digest=sha256:6541fb49b670dd49bb8a44d066f27d067e949ee03bc0097a2dcbde26682401c5

Observation cd9a9ecc-e463-4993-813d-b69c6f6f83d1 · outbound

This paper cites Multiscale vision transformers.

Vision encoders should be image size agnostic and task driven Multiscale vision transformers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.347313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.347313Z digest=sha256:c3d1d069a8df71efb4a23f902de0850ef6c9ec345117e82ac3f535993ec71c4c

Observation 9d6dada1-ce50-4a24-a437-ad3b1b399603 · outbound

This paper cites Levit: a vision transformer in convnet’s clothing for faster inference.

Vision encoders should be image size agnostic and task driven Levit: a vision transformer in convnet’s clothing for faster inference

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.350398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.350398Z digest=sha256:7b45979fd8aeaedd54c0ac039fbe8f3b121f5b33f8e62c0d07eec4ecfafbd2ec

Observation c15b22e6-0c54-4e8e-aadd-8a96b0755bca · outbound

This paper cites GMAT: Global Memory Augmentation for Transformers.

Vision encoders should be image size agnostic and task driven GMAT: Global Memory Augmentation for Transformers

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.353476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.353476Z digest=sha256:1fe6fee3d0df886a4606d0d8c877bf1c08ccefa2bfc997c5de6b56580c876f9c

Observation 71b3601a-fb5c-4bde-9fdc-39f402495d78 · outbound

This paper cites Deep residual learning for image recognition.

Vision encoders should be image size agnostic and task driven Deep residual learning for image recognition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.356848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.356848Z digest=sha256:f7c08599128cb1ddf3c07bb77e2eb4b6e40bbbd62e2145a882011d156d1fe1a0

Observation 1fce41d9-bef1-4e0b-b1a0-a33155f5f907 · outbound

This paper cites Rethinking spatial dimensions of vision transformers.

Vision encoders should be image size agnostic and task driven Rethinking spatial dimensions of vision transformers

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.359634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.359634Z digest=sha256:ca97057f705ea07cccea57ab5763a3ea7684f79dd8c64babfdbaf4ca9d96ee44

Observation af1c7ca2-575d-46dc-bc05-1b8b7a613231 · outbound

This paper cites Long short-term memory.

Vision encoders should be image size agnostic and task driven Long short-term memory

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.362955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.362955Z digest=sha256:ccc4842f894411611cc7cb83cb87518fc2dedef2cadaa0347678a85a5992d566

Observation 4b1816c1-d9c0-45b2-bfd1-e280f749adea · outbound

This paper cites FoveaTer: Foveated Transformer for Image Classification.

Vision encoders should be image size agnostic and task driven FoveaTer: Foveated Transformer for Image Classification

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.365982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.365982Z digest=sha256:108bad3aae790182a16f493229e3d0ceeb571e56e63044e3ae4bc604e5d7f724

Observation 98bb2074-9493-4d4c-8edb-ec024923197e · outbound

This paper cites Learning multiple layers of features from tiny images.

Vision encoders should be image size agnostic and task driven Learning multiple layers of features from tiny images

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.369577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.369577Z digest=sha256:4337e87550c9bb331f6726a128ff89832fc79d17008d60a4c6c088c50d5acdc1

Observation 87a21bc2-22e8-4708-a579-c6890c405928 · outbound

This paper cites Imagenet classification with deep convolutional neural networks.

Vision encoders should be image size agnostic and task driven Imagenet classification with deep convolutional neural networks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.372280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.372280Z digest=sha256:74d12585df96bbe3354b6b20b348870b07aa2d8bb2253f97f322943d5df97f7f

Observation b21999c8-807e-4d34-afef-ecaf1ad56060 · outbound

This paper cites The shape of ai to come! Talk presented at the AI Action Summit 2025, February.

Vision encoders should be image size agnostic and task driven The shape of ai to come! Talk presented at the AI Action Summit 2025, February

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.375259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.375259Z digest=sha256:50530a0ddc98a880bc79d338d8dc0c4f75b41d5881ff4641134f553422d51da4

Observation 874d4f48-181a-4bc9-991c-b30b33576cae · outbound

This paper cites Gradient-based learning applied to document recognition.

Vision encoders should be image size agnostic and task driven Gradient-based learning applied to document recognition

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.381369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.381369Z digest=sha256:1c5279eb400ebb9526b35d2ce687988921bcc5271ec1f901516733edc31ebb1b

Observation 42c87409-889a-47eb-92b7-833f2d6179c3 · outbound

This paper cites Swin transformer v2: Scaling up capacity and resolution.

Vision encoders should be image size agnostic and task driven Swin transformer v2: Scaling up capacity and resolution

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.384475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.384475Z digest=sha256:062b893e6f881cd9c9f06d1b710a275c8ae58cdba49acde7d28e0445efc62a93

Observation 1797c047-40c3-4b18-84ee-8f7d92039e23 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Vision encoders should be image size agnostic and task driven Swin transformer: Hierarchical vision transformer using shifted windows

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.387249Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.387249Z digest=sha256:a1a3ce29fa461d19c8216e282a4dba17ebe35451119cbd93f5c512ec876c28b8

Observation e91a171c-f011-456f-81fa-4bff63703a76 · outbound

This paper cites Decoupled weight decay regularization.

Vision encoders should be image size agnostic and task driven Decoupled weight decay regularization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.390154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.390154Z digest=sha256:b758f418243e983abcb120df12630ecc46046b6403bba7776f5d8c1e423ce0a4

Observation 5c347d50-b1ec-403d-81f6-f15313cf9843 · outbound

This paper cites Biologically inspired deep learning model for efficient foveal-peripheral vision.

Vision encoders should be image size agnostic and task driven Biologically inspired deep learning model for efficient foveal-peripheral vision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.393193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.393193Z digest=sha256:7dd79c3ac2dd3da333c0528714fe782c8de76876dfbf829de86a555d62ca34d4

Observation 9dd77574-90b7-4942-a527-3ca615b0a47b · outbound

This paper cites Implicit-zoo: A large-scale dataset of neural implicit functions for 2d images and 3d scenes, 2024.

Vision encoders should be image size agnostic and task driven Implicit-zoo: A large-scale dataset of neural implicit functions for 2d images and 3d scenes, 2024

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.396652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.396652Z digest=sha256:5451a861d598ea0cd0ae2d1fa8eb7f78719a4341030a113ca6b8e595d62763fc

Observation 1bc456a7-46b0-4cc3-bf0e-17a4f878922e · outbound

This paper cites Recurrent models of visual attention.

Vision encoders should be image size agnostic and task driven Recurrent models of visual attention

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.399521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.399521Z digest=sha256:0c7164bd14b283d8e9d66eaed042374f8a6132c92c8c7cb318bf8b29f0cfac64

Observation b48966fb-0c63-4651-83cb-0f33d89b9957 · outbound

This paper cites A focused backpropagation algorithm for temporal pattern recognition.

Vision encoders should be image size agnostic and task driven A focused backpropagation algorithm for temporal pattern recognition

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.402827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.402827Z digest=sha256:1707ac89a23ed7c8e8244f718107f4c33906a3572a528fb73f113ee4b4d5032a

Observation 2435bb80-3ff5-4bbc-97c1-b56ba8a7efc9 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Vision encoders should be image size agnostic and task driven DINOv2: Learning Robust Visual Features without Supervision

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.405864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.405864Z digest=sha256:79c68dfe7219e5bde81a333b7568515a68731f9b88d5e7732913a3d15108466f

Observation 48d4fc6e-9e8f-40bd-acd3-ceb99b965205 · outbound

This paper cites Compressive Transformers for Long-Range Sequence Modelling.

Vision encoders should be image size agnostic and task driven Compressive Transformers for Long-Range Sequence Modelling

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.408925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.408925Z digest=sha256:57dd8c3920f61fb0c55f0330177646310ad41db6319064554ee8fb505543f9f5

Observation 989c24ae-3e9a-4302-a5c8-f4ab83a36fac · outbound

This paper cites The utility driven dynamic error propagation network, volume 11.

Vision encoders should be image size agnostic and task driven The utility driven dynamic error propagation network, volume 11

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.412060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.412060Z digest=sha256:4b70dc23d98de906cdbdf57c995ffb5b0ec56533f554e411854f050fc533cddb

Observation 4a252d4e-56bf-4346-8abd-e2766b715858 · outbound

This paper cites Learning to generate artificial fovea trajectories for target detection.

Vision encoders should be image size agnostic and task driven Learning to generate artificial fovea trajectories for target detection

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.415649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.415649Z digest=sha256:9290be54073c231811e81a6a25de6bbb9d5c3ea8c984ab28bf0222b516f6f71d

Observation 5179d160-0a2f-4669-b194-7892ebda47ac · outbound

This paper cites Proximal Policy Optimization Algorithms.

Vision encoders should be image size agnostic and task driven Proximal Policy Optimization Algorithms

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.418789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.418789Z digest=sha256:bab8144182e07f16f0bc0f262e428b42016d148c311d6acfc2b55cd55b22b6cf

Observation b9a4db0a-45d7-440e-81c6-7c63814bfb5b · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Vision encoders should be image size agnostic and task driven DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.421857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.421857Z digest=sha256:e45e0dddc49c1c54fcef304c35f43655ff94db13ae20e0cb7306b3cedaa4edba

Observation f2f43293-0f23-40df-8516-3ec868624b0f · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Vision encoders should be image size agnostic and task driven Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.424852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.424852Z digest=sha256:c3ba8fb6742b3f971ca9e1988dfe83f061560cfbef9f8b17e4880d660a4a7810

Observation a741a6e9-efd4-4ca6-92eb-4c377c255371 · outbound

This paper cites Going deeper with convolutions.

Vision encoders should be image size agnostic and task driven Going deeper with convolutions

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.427895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.427895Z digest=sha256:ca06f66082b8e7996e42f59730536aad717eea9a387666040e15540234cf7e53

Observation b9b68948-deba-4e4a-8d6b-1cc4eccc121c · outbound

This paper cites Training data-efficient image transformers & distillation through attention.

Vision encoders should be image size agnostic and task driven Training data-efficient image transformers & distillation through attention

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.430683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.430683Z digest=sha256:2744913568147205c487d34391fdc8456dd9afba98ccb2176219dc3dcc464def

Observation 06fa8854-08a9-4831-9663-404560e1f09a · outbound

This paper cites Fixing the train-test resolution discrepancy.

Vision encoders should be image size agnostic and task driven Fixing the train-test resolution discrepancy

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.433377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.433377Z digest=sha256:1c63b6e614bc42b1af9e0aaba9ac01923e0d9c0a19381d12c39c5422901793e3

Observation c3ac5e17-72f8-48c0-ada2-10a8eee68f8a · outbound

This paper cites Attention is all you need.

Vision encoders should be image size agnostic and task driven Attention is all you need

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.436126Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.436126Z digest=sha256:3d3c6b077c577ea76469208db0a040f3e14e8aff4cdd8075352f61f60d542b78

Observation bf4ad697-0d87-4c00-881b-3044eda8f240 · outbound

This paper cites Pyramid vision transformer: A versatile backbone for dense prediction without convolutions.

Vision encoders should be image size agnostic and task driven Pyramid vision transformer: A versatile backbone for dense prediction without convolutions

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.438932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.438932Z digest=sha256:837b6bb19ce4c94db8ac4b5e72fa9b49330fb2e1d2d862d3b5631fde09eb5b3a

Observation 165d1bf8-f90b-421c-a49b-418a0760a1ee · outbound

This paper cites Generalization of backpropagation with application to a recurrent gas market model.

Vision encoders should be image size agnostic and task driven Generalization of backpropagation with application to a recurrent gas market model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.442174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.442174Z digest=sha256:ed40c656d91ca5e344368b506bc6eb9ccd4319cd2ababcaadf4b842a9cdc58f1

Observation 3b8af0ec-c7a8-458f-97a2-51a01d46ca96 · outbound

This paper cites Simple statistical gradient-following algorithms for connectionist reinforce- ment learning.

Vision encoders should be image size agnostic and task driven Simple statistical gradient-following algorithms for connectionist reinforce- ment learning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.445077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.445077Z digest=sha256:ed1dd1ae34c2e2aa2ea24a81546578d2d8067a1b640c89bab762d9a93e960fb2

Observation 3f04c197-1320-4127-b732-0d3446cc14c1 · outbound

This paper cites Memformer: A Memory-Augmented Transformer for Sequence Modeling.

Vision encoders should be image size agnostic and task driven Memformer: A Memory-Augmented Transformer for Sequence Modeling

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.447770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.447770Z digest=sha256:90e6f23472e6c8ee46ca8c7031ac1a50322fcc0c49946b0513899c1636871568

Observation 905ebcd7-5c79-456f-aecd-eada393113de · outbound

This paper cites Neural fields in visual computing and beyond.

Vision encoders should be image size agnostic and task driven Neural fields in visual computing and beyond

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.450958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.450958Z digest=sha256:3c39cfdf40b414feb4c2c8cd1e36681951b55aef864b8dba87f1584c18ffc87b

Observation afac8872-93a0-4c4c-8016-5053ab299a07 · outbound

This paper cites Focal Self-attention for Local-Global Interactions in Vision Transformers.

Vision encoders should be image size agnostic and task driven Focal Self-attention for Local-Global Interactions in Vision Transformers

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.454271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.454271Z digest=sha256:42e7ea7809512d9015e3706b72449b6595799e7033046cc0b674104ac8cc0e42

Observation 8ffb39f8-4169-4983-87c7-b5a4406cbcc1 · outbound

This paper cites mixup: Beyond Empirical Risk Minimization.

Vision encoders should be image size agnostic and task driven mixup: Beyond Empirical Risk Minimization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.457813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.457813Z digest=sha256:6e0fde17215939b76d3a53bc2a2757c31b46649fd32d0e22bb031bd4864a1522

Observation 587bb495-b9dd-44c7-bd01-1131514e119e · outbound

This paper cites an unresolved cited work.

Vision encoders should be image size agnostic and task driven Unresolved cited work

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T17:27:27.378175Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:27:27.378175Z digest=sha256:fdccd8d8f2271342432e7d22c4d5b182afa18196a3f5801c9d36fffd7b2f1f2e

Pith citing papers

Observation 5adc7cf5-6218-46ef-94a0-ab5649bd6a69 · inbound

Self-supervised pretraining for an iterative image size agnostic vision transformer cites this paper.

Self-supervised pretraining for an iterative image size agnostic vision transformer Vision encoders should be image size agnostic and task driven

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-10T00:19:46.831894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T00:18:06.989944Z digest=sha256:fee4bc3a92ec974fc8fe1435cd6e24ea30b0bfb15cc7d78b60303648ac1e7fc8