Pith. sign in

Paper Citation Record · LEDGER

Images are Worth Variable Length of Representations

As of 22 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 2 inbound Pith citation observations for arXiv:2506.03643.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.03643 v2

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T11:05:22.416317Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-28T07:09:25.049534Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T07:06:44.335867Z

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved40
  • parse uncertain1
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5fa297b5-09a8-47db-b861-d400370d9dc8 · outbound

This paper cites GPT-4 Technical Report.

Images are Worth Variable Length of Representations GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.235943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.235943Z digest=sha256:f290e40b6a9c5b3d9135b0a2f60696b1e3933d35a1fb249766a6226cc9b802c1

Observation 45912b42-4cdf-44e2-9917-854b7273e822 · outbound

This paper cites Gqa: Training generalized multi-query transformer models from multi-head checkpoints, 2023.

Images are Worth Variable Length of Representations Gqa: Training generalized multi-query transformer models from multi-head checkpoints, 2023

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.239610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.239610Z digest=sha256:ee020c9f2430569da639885e5e713db00c533589bc87e2e4c35db0451d7f0189

Observation d03f0524-d703-4c29-84d8-f74e3ed3d7c6 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736, 2022.

Images are Worth Variable Length of Representations Flamingo: a visual language model for few-shot learning.Advances in neural information processing systems, 35:23716–23736, 2022

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.242682Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.242682Z digest=sha256:c645065f56d7f574755fc3912d6cdba9b574c81506b9ad8323669c6aca815de2

Observation a2f8e363-b706-42b5-8f02-f5da44c53879 · outbound

This paper cites Revisiting active perception.Autonomous Robots, 42:177–196, 2018.

Images are Worth Variable Length of Representations Revisiting active perception.Autonomous Robots, 42:177–196, 2018

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.934151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:05:22.245814Z digest=sha256:0d9f94bc5214e22600d5f20270f3902946898175df2ee1d2460cd8ecd466701d

Observation fd3a7952-9f8a-4c5d-a704-96bcde14e1db · outbound

This paper cites Blur image detection using laplacian operator and open-cv.

Images are Worth Variable Length of Representations Blur image detection using laplacian operator and open-cv

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.924943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:05:22.248818Z digest=sha256:4c1be461f1674359d38ec97226901e2db858fc04ba0ef8e3689bf4faf612e8fc

Observation f132ad6e-5e22-4f43-a69a-77c5872919e6 · outbound

This paper cites Statistical inference for probabilistic functions of finite state markov chains.The annals of mathematical statistics, 37(6):1554–1563, 1966.

Images are Worth Variable Length of Representations Statistical inference for probabilistic functions of finite state markov chains.The annals of mathematical statistics, 37(6):1554–1563, 1966

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.914977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:05:22.251721Z digest=sha256:7edb8c44185971095b476d891ca24aaecc0f65ec16ba6a95e7c1029c9fc50a4a

Observation 891de3fe-d2f9-477f-8142-191130b5ae5e · outbound

This paper cites Pythia: A suite for analyzing large language models across training and scaling.

Images are Worth Variable Length of Representations Pythia: A suite for analyzing large language models across training and scaling

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.904132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:05:22.255187Z digest=sha256:2850c478a2acfe2a017f4bfec76303154ad93bd42a8266f755124c8ef9ed1de4

Observation 01c49392-716c-4588-a8b7-6ca5ececbaaf · outbound

This paper cites Token merging: Your vit but faster.

Images are Worth Variable Length of Representations Token merging: Your vit but faster

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.257886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.257886Z digest=sha256:d480d5811d9a29471e29b74e80dfb5db20dcdaf6f13fa2be05682947202427e3

Observation ad4bbdee-535d-4a74-9173-da359f52b388 · outbound

This paper cites Food-101 – mining discriminative components with random forests.

Images are Worth Variable Length of Representations Food-101 – mining discriminative components with random forests

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.260665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.260665Z digest=sha256:caa56386d504060fd520ef17bb53de271f1518969cac45efb4f0bbbcffc2a7ed

Observation 2172ddcb-58ec-4afc-ac8b-c87d2e38167b · outbound

This paper cites Emerging properties in self-supervised vision transformers.

Images are Worth Variable Length of Representations Emerging properties in self-supervised vision transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.263449Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.263449Z digest=sha256:0c10653868ec6c4149d2b7515b545e01025633594d729cd34c2daedef1e2a849

Observation e44a12e4-71b8-4b62-b426-1cd75b084c70 · outbound

This paper cites Efficient large multi-modal models via visual context compression.

Images are Worth Variable Length of Representations Efficient large multi-modal models via visual context compression

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.873209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:05:22.266244Z digest=sha256:ffc862597a108554e195d864e8573cb44803ef46ad477baf8a5aca0eef060a62

Observation 01fa2fea-701b-42d8-a567-8bd7311bcea3 · outbound

This paper cites Review of image classification algorithms based on convolutional neural networks.Remote Sensing, 13(22):4712, 2021.

Images are Worth Variable Length of Representations Review of image classification algorithms based on convolutional neural networks.Remote Sensing, 13(22):4712, 2021

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.862922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:05:22.269137Z digest=sha256:1b90fb38f652c4c7775f454de011f64708c2c4eb24ad06798feead102e02893d

Observation ca84c8ff-6e6b-486e-b2f8-78b8f3a4b496 · outbound

This paper cites An empirical study of smoothing techniques for language modeling.

Images are Worth Variable Length of Representations An empirical study of smoothing techniques for language modeling

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.852630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:05:22.271750Z digest=sha256:aec497704626d5dc524da1fdceefd9840753e402ef0cf8707188a73f44d36084

Observation 8d482209-8cb5-4eab-a1d2-0e307aee191e · outbound

This paper cites Cimpoi, S.

Images are Worth Variable Length of Representations Cimpoi, S

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.274496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.274496Z digest=sha256:beff44fc4f93775ccdc50f0f0791ac3d471cdcaf3426cf234853d705e12d2304

Observation 6e7ac09b-2df3-4c41-b5d0-2c3d7ccca8a1 · outbound

This paper cites An analysis of single layer networks in unsupervised feature learning aistats.

Images are Worth Variable Length of Representations An analysis of single layer networks in unsupervised feature learning aistats

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.836894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:05:22.277525Z digest=sha256:31c501c661392fdd0b8c1a0bbfad1239c7c62f11787a308bbf4487aa0d1c7df0

Observation c9be1d94-02f7-47eb-b6b9-6470a20c06c5 · outbound

This paper cites Scaling up dataset distillation to imagenet-1k with constant memory.

Images are Worth Variable Length of Representations Scaling up dataset distillation to imagenet-1k with constant memory

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.280475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.280475Z digest=sha256:f89dcb1434a433fd4971f872bc31f3c7dee3f438c9a380218c16aeda4d081f31

Observation 9e6443a4-945f-40b8-b154-11e647a0239c · outbound

This paper cites Top-down control of eye movements: Yarbus revisited.Visual Cognition, 17(6-7):790–811, 2009.

Images are Worth Variable Length of Representations Top-down control of eye movements: Yarbus revisited.Visual Cognition, 17(6-7):790–811, 2009

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.820288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:05:22.283436Z digest=sha256:b78aaa2cca81146539fb071f06a8d007f2474d01b18bd72a5994e4adfe5e4ea9

Observation c76268b8-cb52-4d1e-a89b-0b55abb2c568 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Images are Worth Variable Length of Representations Imagenet: A large-scale hierarchical image database

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.809368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:05:22.286379Z digest=sha256:e27a2be8535e5b44595da5cf7e38e59033cb96458e28eca1790b2eeef00ef83a

Observation ff348b87-899f-4bd0-9660-cbc8c5a06b90 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Images are Worth Variable Length of Representations An image is worth 16x16 words: Transformers for image recognition at scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.289289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.289289Z digest=sha256:7a8ea38e8ebe8147666ab1e16ac2e12ecc6948018276530bd6990ffc606cd51b

Observation 92471137-7e22-4321-8b0b-a8cba96c9d0f · outbound

This paper cites Adaptive length image tok- enization via recurrent allocation.

Images are Worth Variable Length of Representations Adaptive length image tok- enization via recurrent allocation

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.791905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:05:22.292319Z digest=sha256:0863b51ab01ce626ae3a0a8e1ca18ce163589d0b50c6a332748e5b9ff1426210

Observation 6b25f19e-82c0-4e3e-90c4-99087b80d5c5 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Images are Worth Variable Length of Representations Taming transformers for high-resolution image synthesis

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.295367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.295367Z digest=sha256:ae90fc921078cce8da6093a37cfa7873247692737e6331501843b69f01bbe598

Observation 95c113f2-cfef-4850-8774-b2cb5ebb92f9 · outbound

This paper cites Multimodal autoregressive pre-training of large vision encoders, 2024.

Images are Worth Variable Length of Representations Multimodal autoregressive pre-training of large vision encoders, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.775879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:05:22.298398Z digest=sha256:f5bfc106cf763b6dc63b36d6ea2c5d4d6733969d66634965d7e3c9c31edf2da4

Observation d7e007a9-dfd4-4daf-8305-60905aef446f · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering, 2017.

Images are Worth Variable Length of Representations Making the v in vqa matter: Elevating the role of image understanding in visual question answering, 2017

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.301186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.301186Z digest=sha256:23d77e9552343df753a17f991433ee4ebd76ce9a20dca912157d06e52775cf20

Observation c6bd0c5f-092f-44b4-b79f-b4dcf9d75158 · outbound

This paper cites The Llama 3 Herd of Models.

Images are Worth Variable Length of Representations The Llama 3 Herd of Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.304085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.304085Z digest=sha256:48afb4d4c7911dd17ff6380973ccee3769f40ceab86cadf81f71bbe4f3ee114d

Observation fe918958-fb65-46b5-ba89-513bb74ace0d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Images are Worth Variable Length of Representations DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.307061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.307061Z digest=sha256:2c29bda266a460288a31a0e95b1a40f6074e6d748c94a64d4fac6809d4681a3e

Observation 43299d14-b086-4b06-b686-53aa23a0cc53 · outbound

This paper cites A review of semantic segmentation using deep neural networks.International journal of multimedia information retrieval, 7:87–93, 2018.

Images are Worth Variable Length of Representations A review of semantic segmentation using deep neural networks.International journal of multimedia information retrieval, 7:87–93, 2018

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.758858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:05:22.310089Z digest=sha256:7d619b3a5c41c0f1e1731be83680b044e3bdf63fa8f00ccefba06973f66f0188

Observation 019de7a3-6c95-41e8-9e30-3ea473633ec8 · outbound

This paper cites A brief survey on semantic segmentation with deep learning.

Images are Worth Variable Length of Representations A brief survey on semantic segmentation with deep learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.312920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.312920Z digest=sha256:d88455488d1d776b0bf86b1259880d10ba4b8af06dc55470b90fff14c0904e3e

Observation 48f972e5-dcc0-4816-9149-9f89f62bbbe1 · outbound

This paper cites Hierarchical cross-modal agent for robotics vision-and-language navigation.

Images are Worth Variable Length of Representations Hierarchical cross-modal agent for robotics vision-and-language navigation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.741335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:05:22.316327Z digest=sha256:3bcafa6688c9c6fd1c5bdf00006c22bc8e8660280fae0b0984535d0df519b795

Observation 6bece6e8-d26d-4f67-b9c6-82763fa8dad1 · outbound

This paper cites Perceiver IO: A General Architecture for Structured Inputs & Outputs.

Images are Worth Variable Length of Representations Perceiver IO: A General Architecture for Structured Inputs & Outputs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.319209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.319209Z digest=sha256:2685a50eaea4d9f61501cb278154eaa873d0d4cbac2ca015ab1270fc1eef0d65

Observation f3001747-ebb7-4a85-b3d1-db8d77ccf7f0 · outbound

This paper cites Perceiver: General perception with iterative attention.

Images are Worth Variable Length of Representations Perceiver: General perception with iterative attention

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.322257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.322257Z digest=sha256:9b17a8b86b6c5c141a7b8e065ea78900369d0f39bf12b895c53ea54ba05cb8de

Observation 9b001137-dc8a-4ab0-a6e0-3bccaf5394aa · outbound

This paper cites Auto-encoding variational bayes, 2013.

Images are Worth Variable Length of Representations Auto-encoding variational bayes, 2013

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.324996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.324996Z digest=sha256:03d73b3333740d0aa79604cb25028f33e278b0ec42e874fda94814d1128bc986

Observation 36a19078-aea1-4d25-ba63-b83344cfb3e3 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017.

Images are Worth Variable Length of Representations Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.327905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.327905Z digest=sha256:5707efb4b0647b9ebbcf331e994c00fe5bf20867f5c54ac95dc538d092ae5067

Observation 41ee6020-27b0-4d1c-8767-75c13e86566b · outbound

This paper cites Learning multiple layers of features from tiny images.

Images are Worth Variable Length of Representations Learning multiple layers of features from tiny images

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.330945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.330945Z digest=sha256:027c5229d608c23b038aa26cb1664b558f016db8cd5b85435cdc0e7c50dbde94

Observation 77620b32-7059-4b66-a838-71e014c302c0 · outbound

This paper cites an unresolved cited work.

Images are Worth Variable Length of Representations Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.334047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.334047Z digest=sha256:e12afcc6dd6162c13b40216e46f7391e4e4ba416a66b822c9416ee7085bd089b

Observation d9380372-0304-4957-9e22-d96bf19a583d · outbound

This paper cites The roles of vision and eye movements in the control of activities of daily living.Perception, 28(11):1311–1328, 1999.

Images are Worth Variable Length of Representations The roles of vision and eye movements in the control of activities of daily living.Perception, 28(11):1311–1328, 1999

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.698087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:05:22.336862Z digest=sha256:0a8500edf24b62bf25bf07b137573d80cc86100e8389ceca6192769a7a5ac27c

Observation 42c60635-8e1e-4e20-972e-3a8ba85d4ffe · outbound

This paper cites Microsoft coco: Common objects in context.

Images are Worth Variable Length of Representations Microsoft coco: Common objects in context

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.340117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.340117Z digest=sha256:2b09777a5499a460de52848d41f2a7881da6b70c07642cca31abeedbceb94e37

Observation f52a53eb-89ab-4977-b03b-1fc26cb8d4c5 · outbound

This paper cites Visual instruction tuning, 2023.

Images are Worth Variable Length of Representations Visual instruction tuning, 2023

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.681102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:05:22.342828Z digest=sha256:848199b973e9a769305208b983bd13c56f32faad2f8afededd3ef19ecbf8a390

Observation 96561ff4-5ad4-48de-8efe-babd7df27e49 · outbound

This paper cites A survey of image classification methods and techniques for improving classification performance.International journal of Remote sensing, 28(5):823–870, 2007.

Images are Worth Variable Length of Representations A survey of image classification methods and techniques for improving classification performance.International journal of Remote sensing, 28(5):823–870, 2007

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.670826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:05:22.345651Z digest=sha256:60496805ec24ecb1066db46a2acd089219f8c261d634929f1340538b8f0e0e87

Observation 4e3f4811-8715-4b05-a7b0-9bbde605c2c1 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022.

Images are Worth Variable Length of Representations Learn to explain: Multimodal reasoning via thought chains for science question answering, 2022

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.348474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.348474Z digest=sha256:bfff94ee8273b2bb204642fe897d0944aba29c51856db11d95226fb53d188200

Observation 5692391d-aa2f-4c6c-a73f-199b11dd5a1a · outbound

This paper cites Fine-grained visual classification of aircraft, 2013.

Images are Worth Variable Length of Representations Fine-grained visual classification of aircraft, 2013

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.351532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.351532Z digest=sha256:65316387d96c868d02a00fabc6e60ab63ebb901e971b8b8f7a3be3027ba3c3b0

Observation 00df95de-f42a-4d2f-9812-6f9301f61c5f · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge, 2019.

Images are Worth Variable Length of Representations Ok-vqa: A visual question answering benchmark requiring external knowledge, 2019

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.354914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.354914Z digest=sha256:3f3393cb3aab029a6a60a787826961a60550293a00fe196bfd8173c4fca9d911

Observation 7da9feef-d2de-41b9-a215-da20f4d4b611 · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning, 2022.

Images are Worth Variable Length of Representations Chartqa: A benchmark for question answering about charts with visual and logical reasoning, 2022

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.357727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.357727Z digest=sha256:267d20e66d6d98d1c28a47036e67c1efc29b5aa519857c3940ed00a79648ac71

Observation 7b8df15e-85fb-43e6-bb6a-f7c9738b71e3 · outbound

This paper cites V Jawahar.

Images are Worth Variable Length of Representations V Jawahar

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.360520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.360520Z digest=sha256:cc9e2d221616fecb80eb999f0fd12cd73db90f62dc142902f6f57beeeccd20ee

Observation d92654c2-b6b8-4e24-9f1a-081cd8ab8de1 · outbound

This paper cites an unresolved cited work.

Images are Worth Variable Length of Representations Unresolved cited work

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.363555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.363555Z digest=sha256:4705298c70203cd9588b64d1e57502880c9516a6c7336df467c8a3492afe1a99

Observation aaeaaffa-f164-4f98-8d82-5f7d2d7c1063 · outbound

This paper cites Stl-10, nov 2024.

Images are Worth Variable Length of Representations Stl-10, nov 2024

Reference 45

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T11:05:22.619520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:05:22.366118Z digest=sha256:83e0c70d001e3216b0898d7c012c7f7d7005cad10179589053ca1ba8f65d6234

Observation 2ae2af82-135c-4123-8a5c-2f48793d3f72 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Images are Worth Variable Length of Representations Learning transferable visual models from natural language supervision

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.368954Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.368954Z digest=sha256:30fadc2d67ef5519103a38ae2b0efbdaf203c22a7a08f074e3f777a74201c6a9

Observation 82b92334-15b0-45d2-94c7-f836d7634716 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Images are Worth Variable Length of Representations Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.371771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.371771Z digest=sha256:eed8708d834352df6ed519426963ae03b1b957e6e71bf62e3fdc01307583dd95

Observation 10ca220c-212c-444f-8b32-e6b1ef1576ec · outbound

This paper cites Dynamicvit: Efficient vision transformers with dynamic token sparsification.Advances in neural information processing systems, 34:13937–13949, 2021.

Images are Worth Variable Length of Representations Dynamicvit: Efficient vision transformers with dynamic token sparsification.Advances in neural information processing systems, 34:13937–13949, 2021

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.374870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.374870Z digest=sha256:cc1781e81cc74df2e520f4925940ff5e9ba9ae92a364f738b29a84170a5b9e7a

Observation 3be09f3f-052c-41df-aa96-d8b96c5b0430 · outbound

This paper cites Generating diverse high-fidelity images with vq-vae-2.

Images are Worth Variable Length of Representations Generating diverse high-fidelity images with vq-vae-2

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.377480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.377480Z digest=sha256:b1d91652680fdc2308170fb6f886bac9e0615e57207812313f5fed43f7519229

Observation 88a6ba60-4198-4c95-bf35-0f7e44931193 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Images are Worth Variable Length of Representations High-resolution image synthesis with latent diffusion models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.381721Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.381721Z digest=sha256:eb4952e9be7a343042bcd0f1c327467722812596c897e90e52fb6296207c4aae

Observation fe8720a3-f6eb-49ed-81c7-91b2fcbd33f4 · outbound

This paper cites Towards vqa models that can read, 2019.

Images are Worth Variable Length of Representations Towards vqa models that can read, 2019

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.384499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.384499Z digest=sha256:136503e6d4d9ef7abf81b8a8d3c5481d5d1d4e2310e4a879a8b5b62798be2855

Observation 05203ea4-08d8-4c28-9b26-bfc328771026 · outbound

This paper cites Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning.Artificial intelligence, 112(1-2):181–211, 1999.

Images are Worth Variable Length of Representations Between mdps and semi-mdps: A framework for temporal abstraction in reinforcement learning.Artificial intelligence, 112(1-2):181–211, 1999

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.387502Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.387502Z digest=sha256:8060ef665863d664540bc644443363eeb4e9e5c8d338bbfff26996f42ae13589

Observation 8866789d-7efd-46b3-91d9-5bbeb1ed0fcb · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

Images are Worth Variable Length of Representations Gemini: A Family of Highly Capable Multimodal Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.390224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.390224Z digest=sha256:76666fa50780dc8302d6e9eead5ea8a8111a4ed5784ec7d68cd0db2ddcc79b42

Observation 6e42155a-daaf-4746-9f3f-11801312b986 · outbound

This paper cites Neural discrete representation learning.Advances in neural information processing systems, 30, 2017.

Images are Worth Variable Length of Representations Neural discrete representation learning.Advances in neural information processing systems, 30, 2017

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.393356Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.393356Z digest=sha256:d325424d2b0467131d59da7b526a3c413f7b2c97f050ebccfb66931a36be7034

Observation d13295d0-1535-4bae-b745-3635609c0cdd · outbound

This paper cites Attention is all you need.

Images are Worth Variable Length of Representations Attention is all you need

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.396119Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.396119Z digest=sha256:3cb706b7698a850f0b3106d179dc838969e7cc0cde4481bd8b4acbd7c599444a

Observation c6c02abb-3fb2-480d-8926-60765e2584e4 · outbound

This paper cites Supervised hashing for image retrieval via image representation learning.

Images are Worth Variable Length of Representations Supervised hashing for image retrieval via image representation learning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.558474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:05:22.398997Z digest=sha256:f68080401d0d1cffb0d8604837abc716df93323b7b70670185c5955183efe0ca

Observation 584a92dd-b964-41ec-8b3b-20017a04943d · outbound

This paper cites Ehinger, Aude Oliva, and Antonio Torralba.

Images are Worth Variable Length of Representations Ehinger, Aude Oliva, and Antonio Torralba

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.548329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:05:22.401621Z digest=sha256:816525f68f17af07eea05809238e8641d3ae99b399a059e562708c741a33ae99

Observation bc20ad02-831e-4c8b-b824-d8e030d24534 · outbound

This paper cites A-vit: Adaptive tokens for efficient vision transformer.

Images are Worth Variable Length of Representations A-vit: Adaptive tokens for efficient vision transformer

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.404732Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.404732Z digest=sha256:7a22d848dd727764308c4c52794650ad8f823e23ef20245c12f89bad62c514f6

Observation b7f27687-3010-43f3-bdbe-4f0cc2ef0e23 · outbound

This paper cites Scaling Autoregressive Models for Content-Rich Text-to-Image Generation.

Images are Worth Variable Length of Representations Scaling Autoregressive Models for Content-Rich Text-to-Image Generation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.407403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.407403Z digest=sha256:697ba7b61b710d09284cfc36a744c8ce83adf8d1a4120e792c0be3c6f83547ed

Observation 2f1ff9be-878a-4111-aedb-0634abbdcec8 · outbound

This paper cites An image is worth 32 tokens for reconstruction and generation.Advances in Neural Information Processing Systems, 37:128940–128966, 2024.

Images are Worth Variable Length of Representations An image is worth 32 tokens for reconstruction and generation.Advances in Neural Information Processing Systems, 37:128940–128966, 2024

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T11:05:22.410713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:05:22.410713Z digest=sha256:528b027443d793aedc4e4c7fc65c4cfb811263f14e6c77aaec87925b2fc343f0

Observation 88729760-bbb3-4dce-b971-1b0cccb86c0b · outbound

This paper cites Object detection with deep learning: A review.IEEE transactions on neural networks and learning systems, 30(11):3212–3232, 2019.

Images are Worth Variable Length of Representations Object detection with deep learning: A review.IEEE transactions on neural networks and learning systems, 30(11):3212–3232, 2019

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T11:05:22.525814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:05:22.413512Z digest=sha256:067a9200566c05b30ed1b1b312b94110b786558bcd3751fd24b4b319f3de8bf4

Observation 73c305e8-6a6d-4ba0-b741-f6e23dfe44ca · outbound

This paper cites STOP” on a sign as “SHOP.

Images are Worth Variable Length of Representations STOP” on a sign as “SHOP

Reference 62

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T11:05:22.514936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-07T11:05:22.416317Z digest=sha256:1dccd8861fe6116a85896218ef3826afb174048439fa9563d6418143b3e24e4a

Pith citing papers

Observation acf9eb43-b6b6-45a5-8c17-8e97430f7680 · inbound

DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking cites this paper.

DC-DiT: Adaptive Compute and Elastic Inference for Visual Generation via Dynamic Chunking Images are Worth Variable Length of Representations

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T15:10:05.814629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-15T15:10:00.146694Z digest=sha256:e8cb569c10e69c59688ec4388715556e1eaa458c37d128f65a1ea29dc4f6b46a

Observation 2b64c07e-428e-4fbb-9902-0b96192c6b33 · inbound

ChannelTok: Efficient Flexible-Length Vision Tokenization cites this paper.

ChannelTok: Efficient Flexible-Length Vision Tokenization Images are Worth Variable Length of Representations

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-02T07:06:44.338154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-28T07:09:25.049534Z digest=sha256:0d55a7034e6efa634be7141597be65bb425e217784970b298da41323c9c3918d