Pith. sign in

Paper Citation Record · LEDGER

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models

As of 15 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 0 inbound Pith citation observations for arXiv:2607.23373.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.23373 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-31T00:01:39.867668Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

70 of 70 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved69
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 094bf45c-ea4c-4235-a8c0-cd09ff089442 · outbound

This paper cites LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:32.624103Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:32.624103Z digest=sha256:b22d1ee84a30636b7f7e5adf66b9e3c9472d500f1a3b26891d40dbb3378619cd

Observation f1d37e2b-ebad-45e0-a7aa-d5dc00248eb2 · outbound

This paper cites In: Proceedings of the AAAI Conference on Artificial In- telligence.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the AAAI Conference on Artificial In- telligence

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:32.725286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:32.725286Z digest=sha256:981e61a985bb863737dfe9030c25e64bf2ef40762efad98dab5c4202aa9558ed

Observation 965803ed-75c0-4000-891e-cb6e51bad4ff · outbound

This paper cites Qwen3-VL Technical Report.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Qwen3-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:32.858371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:32.858371Z digest=sha256:0cd3e5d5231d367082d7d731435820c3b9fa2f6e37284a2f0f804027a3fd6fee

Observation ff630cb5-78ae-4394-97c4-266bfe7f780e · outbound

This paper cites Qwen2.5-VL Technical Report.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Qwen2.5-VL Technical Report

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:32.922639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:32.922639Z digest=sha256:6e86ea18aabc9c380afe76bb5409022d7e3a741400f441f679e047ced0131e4c

Observation 07347f26-0010-46dc-a137-c82a5de590c0 · outbound

This paper cites Longformer: The Long-Document Transformer.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Longformer: The Long-Document Transformer

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:33.003935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:33.003935Z digest=sha256:4e9156d5292fbc45d4db666244887897a839f860aa23e97374a89bc2b87e26c6

Observation 5be5d2c5-d832-4de9-8fa8-ff1cbf161ede · outbound

This paper cites Perception Encoder: The best visual embeddings are not at the output of the network.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Perception Encoder: The best visual embeddings are not at the output of the network

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:33.119140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:33.119140Z digest=sha256:adb079a67a621b0965ee54eea2bf2143a1301d8bf69ac92867a75f7c604dc258

Observation d24f56e4-f760-42ae-8421-ae9c615a4d8c · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:33.171167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:33.171167Z digest=sha256:e245c51a9eaf101e20412248048ddcd7b194c89eadf5b80d34bb21ba46810142

Observation 5354d74b-5494-41bc-a951-10a73d61707e · outbound

This paper cites In: The Thirty-ninth Annual Conference on Neural Information Processing Systems (2025).

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: The Thirty-ninth Annual Conference on Neural Information Processing Systems (2025)

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:33.245005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:33.245005Z digest=sha256:3f44b3b9ba15850fab8c13ae97bb6b7f42386a2078f98aa6897b7ae81fd40acd

Observation d659ef2f-90cb-45e5-99d1-923c05e3f84b · outbound

This paper cites Matryoshka Multimodal Models.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Matryoshka Multimodal Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:33.328665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:33.328665Z digest=sha256:90cba0b4aef60f86e4d84f3fe253fe04324bf1bf9b6ad01d5aea5088c722e183

Observation b80b76b5-a54f-4290-88b2-59a35f89dfad · outbound

This paper cites In: ICLR (2025).

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: ICLR (2025)

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:33.418212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:33.418212Z digest=sha256:b28773912c7f8ff8a85109e5ae53eba39255c2deb2f5bcae4255de1b6d69d9c3

Observation cc0bd62c-9fc6-4c5a-ad42-aaed435e4fd3 · outbound

This paper cites an unresolved cited work.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:33.516741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:33.516741Z digest=sha256:a71e40138915a8633ffdb22a8bd6b7f017dec3e96b6eb45e0251feaae4c596bd

Observation f31595c2-338e-4243-870f-03797b5e21e3 · outbound

This paper cites Generating Long Sequences with Sparse Transformers.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Generating Long Sequences with Sparse Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:33.623542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:33.623542Z digest=sha256:b8cbf6ed127009c80804dfd8477975c0125cbe6cc912bb4c5f57b50a8e686a0f

Observation ccfa5a24-c162-4fba-9837-a13979c35d7a · outbound

This paper cites MobileVLM V2: Faster and Stronger Baseline for Vision Language Model.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models MobileVLM V2: Faster and Stronger Baseline for Vision Language Model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:33.718457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:33.718457Z digest=sha256:0bdd2b015de37566c5ee9a9fb8e0bcccaab1a742d805bd0fea43a26592389168

Observation 966b53dd-de11-41d0-ba08-6fe1fc8d6a5d · outbound

This paper cites In: Proceedings of the IEEE/CVF International Confer- ence on Computer Vision.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF International Confer- ence on Computer Vision

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:33.803370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:33.803370Z digest=sha256:68618f8c90591ba7355d8f18f056241a545d1b7c916ec1b6b14d4e870dea8675

Observation f8c973fc-007f-411c-ac44-3ca37dce4eef · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Conference.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the Computer Vision and Pattern Recognition Conference

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:33.864979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:33.864979Z digest=sha256:1f57fefe989246cce8957bf3bfb5fe4fbb6f1ad96bbdf00b9e41c7d7e494106e

Observation 898f5330-0203-4f9f-8d82-d5b332a18204 · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:33.911479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:33.911479Z digest=sha256:0bd8f5cb521f79be8887c17274b7b40c3e0d1b4de103b847cf2e6d3d3e3a1f04

Observation 304e0e1a-e400-486c-bb2e-85c4149e3d25 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:34.050706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:34.050706Z digest=sha256:7142ff82e9c52c354a3da99ea61e1ba12dddb301afd15097eb2f43c663fda6df

Observation 8920c283-0f8c-49ef-a77c-6d7068b26b8e · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:34.113031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:34.113031Z digest=sha256:2221f900be078434ce9faa1e08f22730a2a744b14288a67e54b8200fd5a0bb54

Observation c3fb02bf-5c23-40ce-91cf-149d5d25eb02 · outbound

This paper cites In: Proceedings of the IEEE/CVF international conference on computer vision.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF international conference on computer vision

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:34.211569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:34.211569Z digest=sha256:209679d4cccab068b5a0d9f973f741d54cc57b8e2719d35d600eaede21979aca

Observation c9262c34-fb72-4534-aed5-7ebca176928d · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:34.326764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:34.326764Z digest=sha256:6275f05b4c4c1addadbe8e634c6d9b48c648d0efe1fd55c6451f61acdd9b12d8

Observation 902cb00f-32af-4c09-9427-9dad7aedafb2 · outbound

This paper cites Matryoshka Query Transformer for Large Vision-Language Models.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Matryoshka Query Transformer for Large Vision-Language Models

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:34.400885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:34.400885Z digest=sha256:34b5656c2a7b9314767efeb31a3b35abbb36caf36e7af02e21852850bd0e9ab4

Observation d6084cd6-9915-4075-8734-7891f403498d · outbound

This paper cites In: Proceedings of the IEEE/CVF confer- ence on computer vision and pattern recognition.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF confer- ence on computer vision and pattern recognition

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:34.480312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:34.480312Z digest=sha256:20e8e1436a74f7f6537fe9a7d0470eaf4c12152331e54e6896ad5518ecfbb75d

Observation 690679a4-1fea-4ca2-9699-b901c36c8feb · outbound

This paper cites In: European conference on computer vision.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: European conference on computer vision

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:34.628057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:34.628057Z digest=sha256:b107a728e0bc6a3fdf8ea1f2aaeca0e1761db96af69d20ade03fb8c09e0559f5

Observation e9675575-2c75-4aa4-a505-0b5133cc54cd · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models LLaVA-OneVision: Easy Visual Task Transfer

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:34.711225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:34.711225Z digest=sha256:87a040c6ad342232a4411732e95a0f23c95264305c7efdd068cbe54c4ed81b6d

Observation e62b6a09-3d5c-40f4-9451-00bf76278783 · outbound

This paper cites In: Proceedings of the 2023 conference on empirical methods in natural language processing.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the 2023 conference on empirical methods in natural language processing

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:34.779557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:34.779557Z digest=sha256:8badb1b826c7f1f5d239f6d85381ec2a93fac85961e49eddf764650708cb9596

Observation 9313867b-121f-4fb1-8843-5667837e5615 · outbound

This paper cites an unresolved cited work.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:34.854510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:34.854510Z digest=sha256:bae8236021c7ea5087b2c77d312a58821d9658d9aa9b262394e815970b2831e8

Observation bf7eea37-43dc-4d33-9d9d-ab3ab76ff443 · outbound

This paper cites an unresolved cited work.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:34.941159Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:34.941159Z digest=sha256:e096486c87f44a2131174fed8e40ef17a49d341550398d8146b9470f27ab3523

Observation 1a5f5f1c-d61b-4d93-ba34-c416777594c5 · outbound

This paper cites Science China Information Sciences67(12), 220102 (2024).

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Science China Information Sciences67(12), 220102 (2024)

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:35.043229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:35.043229Z digest=sha256:e4979a355c8e59d5e970a996a5853d87ef5759053f51de68ef58958a521d8051

Observation 6b5350e6-8f23-4bca-9f1f-b7bb25cca713 · outbound

This paper cites an unresolved cited work.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:35.188413Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:35.188413Z digest=sha256:65930066861b528c52036a29ee1873cbef9200e8149364d5ce1b1606696083bf

Observation dac0d11d-68b8-4d53-a6df-1d698f3ad44e · outbound

This paper cites Advances in neural information processing systems35, 2507– 2521 (2022).

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Advances in neural information processing systems35, 2507– 2521 (2022)

Reference 30

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:35.260051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:35.260051Z digest=sha256:3303c8d7222d04c4269fc42e9092f875ddb835a87faad178c288c97550df30b8

Observation 518b8411-0511-451d-9c2f-3f7e386c7d83 · outbound

This paper cites In: Proceedings of the European conference on computer vision (ECCV).

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the European conference on computer vision (ECCV)

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:35.361614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:35.361614Z digest=sha256:e5d470bd90f68a60ab68adf443a34f3a9fe83223a016e1eb97b598eeaada54cd

Observation 90f639cf-4016-478a-88b1-479bd5ad07ca · outbound

This paper cites SmolVLM: Redefining small and efficient multimodal models.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models SmolVLM: Redefining small and efficient multimodal models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:35.403324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:35.403324Z digest=sha256:c806ac48fc7bd27d5691b3d27840f50ed7d293089ce07eb207141c168e83c204

Observation 0e022fb1-b882-4283-aa1d-3587102bf34d · outbound

This paper cites In: Findings of the association for computational linguistics: ACL 2022.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Findings of the association for computational linguistics: ACL 2022

Reference 33

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:35.480435Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:35.480435Z digest=sha256:66ad11ef985cf98b7678d4a7ef8020f9970d3f46cb436d4792220424a9f3055d

Observation 1ef4d817-01bb-415c-b67f-34fc60606b83 · outbound

This paper cites In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:35.564805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:35.564805Z digest=sha256:1d35f2efecf951749650487fbd567ffe2c4023db30f2da04088d30fb2a57aa13

Observation b0d7c3ed-df52-414c-b6dc-69d9332eef8e · outbound

This paper cites In: Proceedings of the IEEE/CVF winter conference on applications of computer vision.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF winter conference on applications of computer vision

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:35.625122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:35.625122Z digest=sha256:d0b20de35c15d3d4a9e3ab0b5db87167872c016c6ea5023eb818e3f006d54dec

Observation efdc97a7-c351-4f85-8dea-eaeb47bfb5e6 · outbound

This paper cites Separable Self-attention for Mobile Vision Transformers.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Separable Self-attention for Mobile Vision Transformers

Reference 36

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:35.791774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:35.791774Z digest=sha256:1d57ce815d649896a349f2a0ba06f80a4ffc06d25bde9b3164ec00092018818b

Observation fec8168e-c8aa-4f72-88ff-aadeb31d509b · outbound

This paper cites In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:35.933720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:35.933720Z digest=sha256:14e5638aa592bfccb9f1aaf27197f50f9d02c3c52a8fd5537ac1e1b3810ab7f4

Observation 392e581b-4f1c-4446-b7fb-56f766bf4b95 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models DINOv2: Learning Robust Visual Features without Supervision

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:36.053011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:36.053011Z digest=sha256:ef8cf5c1d45522f8c9193601ed1f20c796a7a275dad9ceb16709ecee2856e9cf

Observation 66e7714e-842d-400b-a8c8-389fec46d6ea · outbound

This paper cites In: European conference on computer vision.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: European conference on computer vision

Reference 39

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:36.212439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:36.212439Z digest=sha256:594c730e9d89126185a23b569c9122eb7e83b7cab88bd4053ca4c2405eea6c70

Observation a9979fba-d67b-4110-9b91-9c616ac07627 · outbound

This paper cites Advances in neural information processing sys- tems32(2019).

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Advances in neural information processing sys- tems32(2019)

Reference 40

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:36.327194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:36.327194Z digest=sha256:425ebabd75762c2240fdd27959c0a61e665267fdef24e80ac7b28ec9272bd0e4

Observation 4f957f63-a273-4a61-8e57-1132151a4b4d · outbound

This paper cites In: International conference on machine learning.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: International conference on machine learning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:36.383614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:36.383614Z digest=sha256:bad252da90ebd5b7fc5f84efb101e1ae7977bf82bdae392e414ca4aacca03588

Observation ac691506-1418-418f-8642-8f1a39d66ff0 · outbound

This paper cites In: Proceed- ings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceed- ings of the 26th ACM SIGKDD international conference on knowledge discovery & data mining

Reference 42

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:36.578277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:36.578277Z digest=sha256:83caf35735b97673a47b226833dd5ca093bd608cd98ef5e536b6c4b6e50af956

Observation e08d5a83-4b32-47d6-99b5-eb49a7e53b41 · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:36.708581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:36.708581Z digest=sha256:2258d666a22a16e74c22864a341d930915c05e29a73b7f74f5950254a5e7ec8c

Observation 8a8ec369-3840-4719-8f29-940fd50b73b5 · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 44

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:36.832203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:36.832203Z digest=sha256:070dd012825ece23d46f12e55d72e99fc666702e881bb46f4611f812045cc352

Observation b761f5b3-a653-4cf2-a64b-2b9000002e1c · outbound

This paper cites In: European conference on computer vision.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: European conference on computer vision

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:36.985781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:36.985781Z digest=sha256:a199b5c445c32bb0a21c4df8d41056ad9c0113d5398593313b9846948012a02e

Observation 02974f28-d28a-42dc-84c6-9b21e8e4d729 · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:37.103468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:37.103468Z digest=sha256:37a03540aa48f17b28687ddc34ffa212f6ee0d162a8a1061dc72118cc83a80d9

Observation 0deded97-91c5-4093-8024-557daf45ee63 · outbound

This paper cites In: International conference on machine learning.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: International conference on machine learning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:37.255686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:37.255686Z digest=sha256:9d76b4569f516e1a4a9c48487432f80c5163cb8fdad9d42989d758f70a1f1233

Observation 13e07a65-2d5a-469d-bf35-4e99ee468c0e · outbound

This paper cites TULIP: Towards Unified Language-Image Pretraining.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models TULIP: Towards Unified Language-Image Pretraining

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:37.402513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:37.402513Z digest=sha256:8de1caf566a4de8186b8730d75736069c364d5b495adb910f2f243052465a138

Observation 26ef6db0-6940-46be-b7d9-cdf0d43f86c6 · outbound

This paper cites Patches Are All You Need?.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Patches Are All You Need?

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:37.527536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:37.527536Z digest=sha256:b3b51f299e458f06180bf30d898f4d23a31624776501184ae17c78ff838c47d0

Observation 3d24a58a-acf8-4e8d-b015-bf246446a3bd · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 50

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:37.632360Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:37.632360Z digest=sha256:795a4907529e7e3a0b50d25afeac4a228210fae031c2c0b14964ee167f134d4e

Observation 7cefa37e-9d15-48fe-97e3-bc8a10b72dc0 · outbound

This paper cites Advances in Neural Information Pro- cessing Systems36, 46830–46855 (2023).

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Advances in Neural Information Pro- cessing Systems36, 46830–46855 (2023)

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:37.739735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:37.739735Z digest=sha256:d52ba034b0713df25100e426104ea000c05163c3c03db1f21aa8b5b204893c14

Observation 43709f27-d5ee-443b-98f2-53beaa0abf43 · outbound

This paper cites In: Proceedings of the Computer Vision and Pattern Recognition Conference.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the Computer Vision and Pattern Recognition Conference

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:37.864270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:37.864270Z digest=sha256:92d11246c08320898fb597e34e1da61286daa7201013a7b95d6a6789e53579b1

Observation 67ecff82-707c-409e-8f91-d081c975443c · outbound

This paper cites In: Proceedings of the IEEE/CVF international conference on computer vision.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF international conference on computer vision

Reference 53

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:37.978165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:37.978165Z digest=sha256:b0067e251e1d453eafd1556269a3180c79064e7c5962829b9358a5fb1e0a7333

Observation 4ff79c6b-ed9f-4880-a567-30c7ae0c945c · outbound

This paper cites Advances in neural information pro- cessing systems30(2017).

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Advances in neural information pro- cessing systems30(2017)

Reference 54

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:38.050580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:38.050580Z digest=sha256:568b8f3b4bfb466d45e5879fad6c1100b31af02748e2d1ed919b83b45941e2fe

Observation f2cf568c-90f1-4973-80d3-e2ad1412f34a · outbound

This paper cites GIT: A Generative Image-to-text Transformer for Vision and Language.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models GIT: A Generative Image-to-text Transformer for Vision and Language

Reference 55

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:38.158513Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:38.158513Z digest=sha256:0f64157b6e4765645787faacccae611aa12d2db2f8ff4dddf905d8762fb4cc0e

Observation 215009d9-eed3-4b8f-a244-fb2b3fb76d13 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 56

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:38.278782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:38.278782Z digest=sha256:27dfb1137fa3886297db98821ceb73a8bca3eecef37b8bb51272b96f2df2e25d

Observation 9e989e71-289b-4c80-acec-be3ad75f102b · outbound

This paper cites In: Proceedings of the IEEE/CVF international conference on computer vision.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF international conference on computer vision

Reference 57

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:38.405105Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:38.405105Z digest=sha256:62325d37474ae607065213460e8c2c538ba1fbbb8a82550e9e6400e5322d4757

Observation 73ff9c81-59b2-4b4f-a334-417a205cd2b2 · outbound

This paper cites CVPR (2025).

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models CVPR (2025)

Reference 58

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:38.562942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:38.562942Z digest=sha256:c1cd8eecccd671fb594283f272fcd83106e0a3d2830bfc129894d85cfe52feb2

Observation 2d6969b0-9ef0-468c-9bac-7297c78604e5 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 59

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:38.737392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:38.737392Z digest=sha256:31fd343a7081c19c3b08e34e1dc7a8af13fee19ca0c99041bddd06d04c604385

Observation 879cc3ac-12cd-49fd-a0d2-a935d2e7f9d5 · outbound

This paper cites In: The Thirteenth International Conference on Learning Representa- tions.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: The Thirteenth International Conference on Learning Representa- tions

Reference 60

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:38.805471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:38.805471Z digest=sha256:78b8e0d7a86995e6172dcb6adf802eaa9ba41f77540ce4f2a73433c2c6f16513

Observation db39c310-8faf-4e46-bb8d-fd63e5af1d47 · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 61

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:38.927785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:38.927785Z digest=sha256:0af982d2d04808f003d836a9f4299c14682d1d3c88cc43b97a498a244fac0260

Observation 982fa184-a23d-49d3-9749-8c5c1ed3f9d4 · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 62

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:39.027880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:39.027880Z digest=sha256:0918a1d547ace3e6e992248eae13a6b634abc36dfb4c7ba06b1e11831b7c24d7

Observation 4a7a6ce6-7135-4398-809d-52fc868c2d20 · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pat- tern recognition.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF conference on computer vision and pat- tern recognition

Reference 63

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:39.145738Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:39.145738Z digest=sha256:77a59226aea5b67c022cc276d2d0669747d35f183f20e2cbf243438e8c5d7e15

Observation ca06b46d-6b45-4976-b46d-6285fe82513a · outbound

This paper cites In: Proceedings of the IEEE/CVF international conference on computer vision.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF international conference on computer vision

Reference 64

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:39.208088Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:39.208088Z digest=sha256:498aeaaa826507985554729bb833ba7491b2f8f2d3a486bf890519d62f8fb605

Observation 5a712220-6839-4948-a204-d866fc4a2017 · outbound

This paper cites In: Proceedings of the IEEE/CVF International Conference on Computer Vision.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE/CVF International Conference on Computer Vision

Reference 65

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:39.326347Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:39.326347Z digest=sha256:f5fb5d3beeac84f208a367e438a209ec8356285aee5aa7ac9194e5a9221a9c80

Observation 62f12636-3703-4b1f-aa19-017e96d3ab07 · outbound

This paper cites arXiv e-prints pp.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models arXiv e-prints pp

Reference 66

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:39.434396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:39.434396Z digest=sha256:8e27ff9858dafb262e42b17a134af09ad61bcc0090c4ac8c4fe97d5b88cf8b3a

Observation 399b288d-08f2-4733-8220-7de4df0a0dba · outbound

This paper cites In: Proceedings of the IEEE conference on computer vision and pattern recognition.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models In: Proceedings of the IEEE conference on computer vision and pattern recognition

Reference 67

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:39.532976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:39.532976Z digest=sha256:36029a64ca2545ed076ada051207f3d6a40e9c478e5de334216f1468bdb15cf1

Observation 3814d9e8-3e61-434a-a671-c35ad689f654 · outbound

This paper cites SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models SparseVLM: Visual Token Sparsification for Efficient Vision-Language Model Inference

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:39.654463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:39.654463Z digest=sha256:e27ddfbc29307aa7916428a484baa474d982957c179f24dbc787952f7fb9f5ff

Observation f1d4c1ff-9071-4986-8938-1af28d583a6b · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 69

Resolution
unresolved
no resolver link, observed 2026-07-31T00:01:39.738511Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:39.738511Z digest=sha256:ee46866fac67b1e4fffcf152cea4f60ff9671e2bbdedecad9533f1eb46d2e11a

Observation 8fe9eb04-1671-44ed-966d-39493cdc64f6 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

UltraViT: Latency-Optimized On-device Vision Encoder for Large Vision-Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 70

Resolution
malformed identifier
no resolver link, observed 2026-07-31T00:01:39.867668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-31T00:01:39.867668Z digest=sha256:682aa7e56cdf6e566674f2d1e87f121ed092a73c5766dcff2b8c8c44af190058

Pith citing papers

No inbound Pith citation observations are available.