Pith. sign in

Paper Citation Record · LEDGER

RADIO1D: Elastic Representations for Condensed Vision Modeling

As of 17 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2607.03624.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.03624 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-12T01:07:20.766474Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact2
  • verified fuzzy0
  • unresolved56
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 66caf685-be41-48be-a201-9ebaf35dc24a · outbound

This paper cites Learning transferable visual models from natural language supervision.

RADIO1D: Elastic Representations for Condensed Vision Modeling Learning transferable visual models from natural language supervision

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:463c129e7a09699259aa0e4873648a02c41dffa5b4e20d06eb200fe48de9d87d

Observation 649677a9-321a-42cc-99eb-bb259d835013 · outbound

This paper cites BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language models.

RADIO1D: Elastic Representations for Condensed Vision Modeling BLIP-2: Bootstrapping language- image pre-training with frozen image encoders and large language models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:391b9e0ceee1d97164c42ac8721eab1e50eed14fa54558872d6c8f5066d324e6

Observation 83d05092-8465-40ae-a363-34acb78c5246 · outbound

This paper cites Language is not all you need: Aligning perception with language models.

RADIO1D: Elastic Representations for Condensed Vision Modeling Language is not all you need: Aligning perception with language models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:cbe4b98268eef6f7b995da3abffc901701ef77f5e7f40bd4b32633107b3f3fa6

Observation 2729808d-b74d-4414-852e-f4de603e2957 · outbound

This paper cites Visual instruction tuning.

RADIO1D: Elastic Representations for Condensed Vision Modeling Visual instruction tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:399327a43a1022e56264d8c43a0f1b0f17445f18697aef42da8b0cde33684446

Observation cc0c6dda-8b77-4550-afd5-0a09274a61e1 · outbound

This paper cites Sigmoid loss for language image pre-training.

RADIO1D: Elastic Representations for Condensed Vision Modeling Sigmoid loss for language image pre-training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:e8670b752beb5f2da2a0d2ee5cdcbc0bab84b702f37b163d995a04332c383df3

Observation b4e8b263-c779-45a5-8665-88c5eeb5aa94 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

RADIO1D: Elastic Representations for Condensed Vision Modeling SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:6b9f73c1108f108c2faad36cc301bf466a270e1cf49327305367f75d36e7eb65

Observation ff0e86db-0c8c-4716-aee0-47e46c9bb751 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

RADIO1D: Elastic Representations for Condensed Vision Modeling PaliGemma: A versatile 3B VLM for transfer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:d37c5b11adcbd31f59b56da00dc9291a1871ce645df5fd1d3bb6a87b19f7a10c

Observation 3425b0fc-02ef-4160-84f0-d0889b1b3e97 · outbound

This paper cites What matters when building vision-language models?.

RADIO1D: Elastic Representations for Condensed Vision Modeling What matters when building vision-language models?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:57bcfdcf9e335cbeced442b45eba7f79c1351df9ca49385cb85811f528a646a3

Observation 6f740f26-d7de-4d0d-aeb2-0c7793c251cd · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

RADIO1D: Elastic Representations for Condensed Vision Modeling Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:c16360bc2e8fad8197270d89ce594f9ef67cc434ee2a244114f1325d99fc5ffb

Observation 89826ebe-5819-4ffd-a570-9dbf1521d10b · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

RADIO1D: Elastic Representations for Condensed Vision Modeling NVILA: Efficient Frontier Visual Language Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:14d7f72a589f56b5a293b46a4a5f9d7950bce2971cd8664ba95fd1aed9266c04

Observation 981597fb-7f4b-4ae1-89d9-0c02a279d687 · outbound

This paper cites Qwen3-VL Technical Report.

RADIO1D: Elastic Representations for Condensed Vision Modeling Qwen3-VL Technical Report

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:f4e00dcc814336f650243017f49b5293d3e12b6594364e4751e0e2eadae966c3

Observation 7330b42e-8738-4433-8747-a026a6ef14cc · outbound

This paper cites Radiov2.5: Improved baselines for agglomerative vision foun- dation models.

RADIO1D: Elastic Representations for Condensed Vision Modeling Radiov2.5: Improved baselines for agglomerative vision foun- dation models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:de996b36882049857ff3c1ec0a49b8a82d191bdfe22cef8e2a28a9a6b8bd2a6d

Observation b66eb5b0-fed4-4dbe-8022-4187bf71d53e · outbound

This paper cites an unresolved cited work.

RADIO1D: Elastic Representations for Condensed Vision Modeling Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:7ab48df711f9f2e6814e56730ee0fabe85a03a5fc867869a259e03565ef7349c

Observation 753be084-d9a9-4903-8007-a0d6e3da9b50 · outbound

This paper cites DINOv3.

RADIO1D: Elastic Representations for Condensed Vision Modeling DINOv3

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:5ed2552993efbdf4667d9f05ed5c605be616e4c40bfa7f187b5f3d242baa71da

Observation c59013d7-ec72-4d46-9bfc-5f41d561a554 · outbound

This paper cites SAM 3: Segment Anything with Concepts.

RADIO1D: Elastic Representations for Condensed Vision Modeling SAM 3: Segment Anything with Concepts

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:8bab9da5f160b42447defab772a9fb5b8de79cc822188f47e66a8cb6a3ccd6a2

Observation 7486a3ee-c219-41a7-b09e-48ad9bd30107 · outbound

This paper cites Eagle: Exploring the design space for multimodal LLMs with mixture of encoders.

RADIO1D: Elastic Representations for Condensed Vision Modeling Eagle: Exploring the design space for multimodal LLMs with mixture of encoders

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:837143c47e410575c164ab4d38b8da2950afe6ed3fe0ea71c17e837713648ad0

Observation 6f75ea9e-ab5a-4abf-9b66-1dda9a5b9129 · outbound

This paper cites VILA-u: a unified foundation model integrating visual understanding and generation.

RADIO1D: Elastic Representations for Condensed Vision Modeling VILA-u: a unified foundation model integrating visual understanding and generation

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:5ac67d7dba2917d78d7843b2bd2e1d5e6e9a5a09b59bd6b85b999db5c938f9e7

Observation d816267e-b8c2-4385-b7de-6804e3fa1cc6 · outbound

This paper cites Qwen-VL: A versatile vision-language model for under- standing, localization, text reading, and beyond,.

RADIO1D: Elastic Representations for Condensed Vision Modeling Qwen-VL: A versatile vision-language model for under- standing, localization, text reading, and beyond,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:8516580db611e600639a77d18b3908c32da8bd4fa97796b11ab31d1eed5c1566

Observation b529dc4a-ac26-4eb9-8b0a-1dc38d6b038e · outbound

This paper cites an unresolved cited work.

RADIO1D: Elastic Representations for Condensed Vision Modeling Unresolved cited work

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:877510718ae3bfb4b0a7fde49eaa674d7b14326cc8f087b4e90d08864a947852

Observation cdb03c63-ac96-41cd-b557-54e2890f3a6c · outbound

This paper cites LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer.

RADIO1D: Elastic Representations for Condensed Vision Modeling LLaVA-UHD v2: an MLLM Integrating High-Resolution Semantic Pyramid via Hierarchical Window Transformer

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:a57cd037933ffd7ea4844b6a36ff11e082eeba204573c0ad20649d8585341a58

Observation eec94014-4c85-4388-864c-9077dff19806 · outbound

This paper cites Scene parsing through ADE20K dataset.

RADIO1D: Elastic Representations for Condensed Vision Modeling Scene parsing through ADE20K dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:3a0d9903c7cc71ceac02a59339687fc1df7a1b9b90b2d624760a1a90ee328ce9

Observation 0b670147-5789-4d89-9d6b-49701cf1cbbd · outbound

This paper cites Schwing, Alexander Kirillov, and Rohit Girdhar.

RADIO1D: Elastic Representations for Condensed Vision Modeling Schwing, Alexander Kirillov, and Rohit Girdhar

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:9e65bd7fd62a2621a0cbd628a9cfa3c36e8687760cb0611bd82a84af9373e1cf

Observation 3f39cf81-c855-4dfe-81c8-23bd62c7077d · outbound

This paper cites an unresolved cited work.

RADIO1D: Elastic Representations for Condensed Vision Modeling Unresolved cited work

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:819301353fba3aea7e6796000b4316cf3df6a276b70d4aac1a36a52254f7610e

Observation e3efc387-ff03-40e9-b7d9-0c425878f774 · outbound

This paper cites DINOv3-driven se- mantic segmentation for landslide mapping in mountainous regions.Sensors, 26(2), 2026.

RADIO1D: Elastic Representations for Condensed Vision Modeling DINOv3-driven se- mantic segmentation for landslide mapping in mountainous regions.Sensors, 26(2), 2026

Reference 24

Resolution
verified exact
doi, observed 2026-07-12T01:08:22.975093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:b4a6ac6d97bb55552bfb7177c65a1bc1d69d4b3648b3b08a5f2ff509561c8bbb

Observation 7b208968-cf6b-45b3-80ee-1b1a2699b973 · outbound

This paper cites Similarity of neural network representations revisited.

RADIO1D: Elastic Representations for Condensed Vision Modeling Similarity of neural network representations revisited

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:cffe1277224001e17559cc2f9bafcac672fad244921b30d065e2d7accae3a9f7

Observation 3fcef4d6-2d43-484b-907e-78114ff4c0f3 · outbound

This paper cites an unresolved cited work.

RADIO1D: Elastic Representations for Condensed Vision Modeling Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:f626ddbb1a69750147d9ff8bc9e9f645e1f0df186d1f91151f6d23890e3830c0

Observation c4fa02a4-4d50-4f62-b8fc-00247176752f · outbound

This paper cites Lawrence Zit- nick, and Piotr Dollár.

RADIO1D: Elastic Representations for Condensed Vision Modeling Lawrence Zit- nick, and Piotr Dollár

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:eeabfb324c77000f11af82388761bad3aa355e3b6afe386b20d0412be878561b

Observation ce04983f-4c12-4b34-9514-9d3c258db8ee · outbound

This paper cites C-radiov4 (tech report), 2026.

RADIO1D: Elastic Representations for Condensed Vision Modeling C-radiov4 (tech report), 2026

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:6ab9f6534debf229d7ba5759a9410eef696342c9a5f9b523321e37daf8ebea79

Observation 5e653ed3-47fd-4bf1-b489-5dd587627d0e · outbound

This paper cites Pereira, and William Bialek.

RADIO1D: Elastic Representations for Condensed Vision Modeling Pereira, and William Bialek

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:0fa908efa42b6331598eb86dcbe7d3df02fbcea91e9294d85b071a469ff64374

Observation 0b70eefb-7a76-4dc7-9032-7eb2101de28b · outbound

This paper cites Modeling by shortest data description.Automatica, 14(5):465–471, 1978.

RADIO1D: Elastic Representations for Condensed Vision Modeling Modeling by shortest data description.Automatica, 14(5):465–471, 1978

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:3dd060fcc688ccf928b53b8fb3edbac85649678f1609268c52f7f5c4fc5ef584

Observation 4f332976-3e04-4e2e-ac85-88a8e7db861a · outbound

This paper cites AM-RADIO: Agglomerative Vision Foundation Model Reduce All Domains Into One.

RADIO1D: Elastic Representations for Condensed Vision Modeling AM-RADIO: Agglomerative Vision Foundation Model Reduce All Domains Into One

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:6bcc0fd760bf34ffd08020050ec378fe297866d6aa0b49d86e1f3c5bbfbb27ca

Observation 2e28156b-c771-4b87-8db6-2829b584e89f · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

RADIO1D: Elastic Representations for Condensed Vision Modeling An image is worth 16x16 words: Transformers for image recognition at scale

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:27773230d175ba2dc31823ba03fe5d558beed0fa0b4a7f68365ce8dd49dfa127

Observation 69362b1e-f22f-4eb7-b66c-e219cf945c0f · outbound

This paper cites Flextok: Resam- pling images into 1d token sequences of flexible length.

RADIO1D: Elastic Representations for Condensed Vision Modeling Flextok: Resam- pling images into 1d token sequences of flexible length

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:550441d120f7864008b1c99b38eebe2d3f6e8f99cacf5219b41259dd0d871474

Observation d0a3f5da-d4fe-4108-a26a-803d177fc429 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

RADIO1D: Elastic Representations for Condensed Vision Modeling Swin transformer: Hierarchical vision transformer using shifted windows

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:35e1fd683c990296ae9ef76fab36b85f52a41527ea7c5ff39ea7d31b4a65e26e

Observation 64f9987d-9b2e-495e-97c9-e375d3f97e63 · outbound

This paper cites DataComp: In search of the next generation of multimodal datasets.

RADIO1D: Elastic Representations for Condensed Vision Modeling DataComp: In search of the next generation of multimodal datasets

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:3148895497635d6e6762da76731e784e501c30f730911bd209b706f22277b244

Observation f4d1825f-7d6f-49c2-a2e0-48f4e84aeeed · outbound

This paper cites Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design.

RADIO1D: Elastic Representations for Condensed Vision Modeling Getting ViT in Shape: Scaling Laws for Compute-Optimal Model Design

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:baaee86329ce9288e3f0a96a147e07592043ad0ff9adc6c8a49adb2b5cc8e48b

Observation 29f51107-8494-4970-9be4-cb42825b3884 · outbound

This paper cites Cover and Joy A.

RADIO1D: Elastic Representations for Condensed Vision Modeling Cover and Joy A

Reference 37

Resolution
verified exact
doi, observed 2026-07-12T01:08:22.982569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:864b28793ed80b590a20b7c498b8ef425a0af5e00b56600644e5cf29489c1466

Observation 804dbc10-94b1-439c-a10b-d3979b0f1d76 · outbound

This paper cites NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model.

RADIO1D: Elastic Representations for Condensed Vision Modeling NVIDIA Nemotron Nano 2: An Accurate and Efficient Hybrid Mamba-Transformer Reasoning Model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:90342c8f31f809c3ba0e51827842f0e679f86b9a8c2ab21f8586362ec967498c

Observation dbefd43c-09a1-43fa-af84-f742e1ce83e9 · outbound

This paper cites Towards vqa models that can read.

RADIO1D: Elastic Representations for Condensed Vision Modeling Towards vqa models that can read

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:0816a9237735f88eefca378a0c0d8ad182f1af336660cf2146a115dc85f0c2a6

Observation d0a1ea24-90f5-47ea-a7ef-7e51dccc1a17 · outbound

This paper cites DocVQA: A Dataset for VQA on Document Images.

RADIO1D: Elastic Representations for Condensed Vision Modeling DocVQA: A Dataset for VQA on Document Images

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:2abb8c6b34c533fdd3a764ed576456800fd66507f8c9a86f902498ef6306bc4f

Observation 55a4d441-fd07-464d-8ab9-848e44d22386 · outbound

This paper cites an unresolved cited work.

RADIO1D: Elastic Representations for Condensed Vision Modeling Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:a83055b62c716bf5490ec73276ba9c8694f1bf7d93594cf1c7b7e12494440430

Observation ceabeb51-c265-4956-8071-29fcef4cb21f · outbound

This paper cites OCRBench: On the hidden mys- tery of OCR in large multimodal models.Sci- ence China Information Sciences, 67(12):220102, dec 2024.

RADIO1D: Elastic Representations for Condensed Vision Modeling OCRBench: On the hidden mys- tery of OCR in large multimodal models.Sci- ence China Information Sciences, 67(12):220102, dec 2024

Reference 42

Resolution
malformed identifier
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:b088c75700feb7090ef66cb8bb8b8c568fe2f27995fdd67381e8e57e5bdf7e3e

Observation 0b898045-7a72-44b1-a8b8-ee2d60cf0b6a · outbound

This paper cites OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning.

RADIO1D: Elastic Representations for Condensed Vision Modeling OCRBench v2: An Improved Benchmark for Evaluating Large Multimodal Models on Visual Text Localization and Reasoning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:ad74494235e91047de042b2ec1cdb224d521c11b6f090a6f17b8b5b7dd88cd4c

Observation 5b8448dd-d5af-458c-b582-ca231587385e · outbound

This paper cites A diagram is worth a dozen images.

RADIO1D: Elastic Representations for Condensed Vision Modeling A diagram is worth a dozen images

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:0e7084ef14742bb790c6b1bc051c76488cca72a17963c7cdcad28e5cc5b6e49a

Observation 461920ee-dd50-4318-ab0d-74fcaaaea9c2 · outbound

This paper cites ChartQA: A benchmark for question answering about charts with visual and logical reasoning.

RADIO1D: Elastic Representations for Condensed Vision Modeling ChartQA: A benchmark for question answering about charts with visual and logical reasoning

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:2788a590bb8873e89947cbb182b647f097a15a6cdde8b3cbb94abbf06bc031f9

Observation f69973c4-41e2-4e3e-8c2f-91e7df19fd66 · outbound

This paper cites findings-acl.177.

RADIO1D: Elastic Representations for Condensed Vision Modeling findings-acl.177

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:ca251ae231b1cc0a542969e3940239e18063e74ede07235681b6a1010e192175

Observation f364fbe3-0860-454d-82c8-afdf54fae839 · outbound

This paper cites Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi.

RADIO1D: Elastic Representations for Condensed Vision Modeling Mmmu: A massive multi-discipline multi- modal understanding and reasoning benchmark for expert agi

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:12c0491c0179b4e4f55fedd1cf5b6ad90df19deb90508b5e766ca06bb7275797

Observation 4d5edd92-9330-46c6-83b4-312464a78755 · outbound

This paper cites Seed-bench: Benchmarking multimodal large language models.

RADIO1D: Elastic Representations for Condensed Vision Modeling Seed-bench: Benchmarking multimodal large language models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:890f1c9472a03a6ce7d2341738f4efbe15da6bb72a357195870239179d2b9a3d

Observation 9bee5ee8-9c47-4b2c-83dc-8bdf3f6797f9 · outbound

This paper cites Longvideobench: A benchmark for long- context interleaved video-language understand- ing, 2024.

RADIO1D: Elastic Representations for Condensed Vision Modeling Longvideobench: A benchmark for long- context interleaved video-language understand- ing, 2024

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:2dfbbbcc9d27e2587741946fa459fe11808d98c0cab6c60c1668297fed27648c

Observation 8e10099c-2ee2-4bde-9bd8-18fad7fbfe66 · outbound

This paper cites Token merging: Your ViT but faster.

RADIO1D: Elastic Representations for Condensed Vision Modeling Token merging: Your ViT but faster

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:b4122122052692e8939fcb20f8de02dbe67ccfe37eaf50419a4cba11976e9260

Observation af6ec779-02cc-4cd9-9823-eff642eca4cb · outbound

This paper cites Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs.

RADIO1D: Elastic Representations for Condensed Vision Modeling Beyond Attention or Similarity: Maximizing Conditional Diversity for Token Pruning in MLLMs

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:96c3a04980bd877b03decd2640e7a61a0236cc028da1c2c64bc30553df01ea68

Observation f6521bc4-61be-4709-bdfd-a16a87087341 · outbound

This paper cites Generalized intersection over union.

RADIO1D: Elastic Representations for Condensed Vision Modeling Generalized intersection over union

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:387a70996a75df50c82226eda3b331c4194de96a3d668f15f162b395c681e27c

Observation 75a16e6d-d832-405c-ba20-5780c54e871b · outbound

This paper cites an unresolved cited work.

RADIO1D: Elastic Representations for Condensed Vision Modeling Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:17459bd1d9993156e617bf6a5d016cf814521450d68d7f444b22a3c4bed26ac5

Observation 54a38a24-d414-4cbe-95ab-73f04a3552b7 · outbound

This paper cites nuScenes: A multimodal dataset for autonomous driving.

RADIO1D: Elastic Representations for Condensed Vision Modeling nuScenes: A multimodal dataset for autonomous driving

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:38032830b7d9e938f34fb9ff453f4933901dc8e3e304274742bc10662adad829

Observation 9bd40bc9-03a8-4e3d-a9f8-a92971d1603a · outbound

This paper cites Vision transform- ers need registers.

RADIO1D: Elastic Representations for Condensed Vision Modeling Vision transform- ers need registers

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:a69fc36973cc01f5a19ea2243be814f5a486a4250c874391454487ce484346eb

Observation aeb46822-da92-4a36-b465-b232079d691e · outbound

This paper cites an unresolved cited work.

RADIO1D: Elastic Representations for Condensed Vision Modeling Unresolved cited work

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:1239c5603c14757be6ce4601f44824a611de5ec0e547c6c3f2c69f631adf8ce7

Observation 55116568-fc65-4b8a-a22b-30387427ee3b · outbound

This paper cites An image is worth 32 tokens for reconstruction and generation.

RADIO1D: Elastic Representations for Condensed Vision Modeling An image is worth 32 tokens for reconstruction and generation

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:6f5261f0f21343a21712cc1f492e8aa49ed2d9d5ee79f443d6b43d31b2b0ce66

Observation af783732-6037-44ee-ba36-c46a55a80944 · outbound

This paper cites Net2Net: Accelerating Learning via Knowledge Transfer.

RADIO1D: Elastic Representations for Condensed Vision Modeling Net2Net: Accelerating Learning via Knowledge Transfer

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:e5ac4674e5f9b8dd4a9f17561d3f20401dcf08d9a5b23459fb52efbfe3ab776b

Observation 75e8dad6-1e16-439d-8faa-08ce07560d8d · outbound

This paper cites an unresolved cited work.

RADIO1D: Elastic Representations for Condensed Vision Modeling Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:4dc3538ffa5890e3e9ccbabe2df734356dc9dc740dfcc70b5f0305a20b7f556b

Observation a5a0d2a4-68a1-4e6e-a731-74058dca946f · outbound

This paper cites InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks.

RADIO1D: Elastic Representations for Condensed Vision Modeling InternVL: Scaling up Vision Foundation Models and Aligning for Generic Visual-Linguistic Tasks

Reference 60

Resolution
malformed identifier
no resolver link, observed 2026-07-12T01:07:20.766474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T01:07:20.766474Z digest=sha256:0027bba860656207fb2ec17107bd104b4f1e143cba07534f8c046c149f8d70b7

Pith citing papers

No inbound Pith citation observations are available.