Pith. sign in

Paper Citation Record · LEDGER

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs

As of 8 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2507.00754.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00754 v2

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:17:14.935499Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cf57972b-3253-4d67-8cc2-8dbcf447dd24 · outbound

This paper cites Language models are few-shot learners.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Language models are few-shot learners

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:24.184422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:17:06.174747Z digest=sha256:e57da9a13b7f9b81648a7df19a49787ec9dd9cf3879ccc2fb8df5eee2c0a916a

Observation 0d832a76-bf8a-4879-b8af-89e85431c105 · outbound

This paper cites In these works, they provided litmus tests for measuring how focused the attention patterns of particular models are and how they relate to model robustness.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs In these works, they provided litmus tests for measuring how focused the attention patterns of particular models are and how they relate to model robustness

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:22.984748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:17:11.783929Z digest=sha256:3f231633e9ee8b2889a860e088cfd7c34f5b4bf5ca62caedf6ceb5eea3ef6ac2

Observation 80bb878e-7eab-488d-bf54-fac98f9d577b · outbound

This paper cites the standard deviation of the accuracy values divided by the number of different seeds.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs the standard deviation of the accuracy values divided by the number of different seeds

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:20.966369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:17:12.633122Z digest=sha256:19e186ecb7f19fbb4c4df29bb85f8d68f6bb1e87c562e16bc21594170e1d354f

Observation 7df17b61-bd53-4908-bfe4-0825ff4c8666 · outbound

This paper cites Frozen Transformers in Language Models Are Effective Visual Encoder Layers.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Frozen Transformers in Language Models Are Effective Visual Encoder Layers

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:17:16.472633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:17:06.952393Z digest=sha256:cc315f85f913a7074bf203562269ee450639725c2ef6c51f44f4932d632f0efb

Observation 8067ec4b-4bd5-4eab-8942-b273bfb5247d · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:07.140211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:07.140211Z digest=sha256:9752ea60877adf0c664498166c366a6a22b78697d6a8bdd66c35ef64d253b64c

Observation f0432910-0c1c-432f-b169-d2edd481baa3 · outbound

This paper cites Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:08.054862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:08.054862Z digest=sha256:6e97a6b4a1decbb09b48338390186c39f4d9a12817df34df72898882bd53cca0

Observation a6a28afd-2ec9-47d8-a024-7db7890eb3fc · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs DINOv2: Learning Robust Visual Features without Supervision

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:08.132343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:08.132343Z digest=sha256:a08e5c83cf320afce4028842a0e884d77e1cd21c33140e7776cbd595ece809c4

Observation f5907d37-48d9-4cc7-9e57-bc8d05756ac8 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Imagenet: A large-scale hierarchical image database

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:08.286198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:08.286198Z digest=sha256:6a55ec5cb742e2830007536a26d035a2bebe3411030ff52d67d4f9a70fc46d67

Observation b3724c4e-8c25-47a4-8b1f-0025c935b7d9 · outbound

This paper cites Microsoft coco: Common objects in context.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Microsoft coco: Common objects in context

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:08.730823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:08.730823Z digest=sha256:62645cf38f429e45d4a1c6d823635e1a74c9866c45ff9404d0a6c4c05c219ee9

Observation dbb7affd-7b08-4c85-9bcd-6eb3e47e0215 · outbound

This paper cites Attention Entropy is a Key Factor: An Analysis of Parallel Context Encoding with Full-attention-based Pre-trained Language Models.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Attention Entropy is a Key Factor: An Analysis of Parallel Context Encoding with Full-attention-based Pre-trained Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:08.954914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:08.954914Z digest=sha256:63fb529768c4e093e9f9128e4a83d5de384c83216176e69f098db9216f8b9451

Observation 54bfea88-6ae9-486d-80d8-fe4d093f67bc · outbound

This paper cites Torch.manual_seed(3407) is all you need: On the influence of random seeds in deep learning architectures for computer vision.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Torch.manual_seed(3407) is all you need: On the influence of random seeds in deep learning architectures for computer vision

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.069919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.069919Z digest=sha256:cf428534ba5467ae4dcc547d16e417ff9d6440103db42c6389abeb6ccf760e6d

Observation 6a846dd8-e028-4ee5-ad8d-aa0c10fafebe · outbound

This paper cites Understanding Why Neural Networks Generalize Well Through GSNR of Parameters.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Understanding Why Neural Networks Generalize Well Through GSNR of Parameters

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.234749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.234749Z digest=sha256:4f76d4ae218e29eb82b0da48837e6d5a8b1744c16be6a86766fb5bd09959f78c

Observation a964eac8-6510-4018-b4b1-d0208ec5c971 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Distilling the Knowledge in a Neural Network

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.364748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.364748Z digest=sha256:57df05d88e5efa57f566a488665c97067b0dd2236119548d94909d06e929f864

Observation c1b79421-ce28-4c47-b15b-bc68e211d958 · outbound

This paper cites SOLO: A Single Transformer for Scalable Vision-Language Modeling.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs SOLO: A Single Transformer for Scalable Vision-Language Modeling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.614748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.614748Z digest=sha256:fa7b0525d87ca869813f5dbfb861eed19df17a6bf95c5b9ae050772f343f749f

Observation 7eb34fdb-15b0-4387-afd5-973902d3b467 · outbound

This paper cites Layer Normalization.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Layer Normalization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.764750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.764750Z digest=sha256:7e6a4f6bc641b26408795346e80d56e1ff9aaa120afe0afab12fcf6138b3061a

Observation 36fc9a21-bae4-4c9b-82ca-5261a805f6d7 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Adam: A Method for Stochastic Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:10.275076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:10.275076Z digest=sha256:0bc09a5fa787059af59abbce184b3d4d49a05825ffc4ba118ec8bd22870b1139

Observation b9ac64ec-a47d-4723-b68d-338970fc7d57 · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:10.544749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:10.544749Z digest=sha256:3126b9b017f101fa528aa6807db46e018b1d7da19396b2011cb9ae3ac6e5dbf9

Observation 330ec6c4-24c3-4029-9c70-faf41c147242 · outbound

This paper cites Deep networks with stochastic depth.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Deep networks with stochastic depth

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:23.564748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:17:10.764825Z digest=sha256:95f1ec9313cef24601cdcd3c48772bbc29981682623688e058044ee568527a4f

Observation dde40791-2763-47ba-85c7-68ab4f603424 · outbound

This paper cites mixup: Beyond Empirical Risk Minimization.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs mixup: Beyond Empirical Risk Minimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:10.884728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:10.884728Z digest=sha256:cdd993726d7dd0cdd8f74260c6a9c384fb8c2577c49082f83f5dbc19961a9351

Observation e882166e-8e34-4822-a2a3-3639df5af93f · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:11.024571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:11.024571Z digest=sha256:3fca6d76997ea6988e2bd1fdac1135e0917a691c9ce46bef06305632b065d695

Observation 3551f281-1f14-42b8-9919-51b480181047 · outbound

This paper cites MMDetection: Open MMLab Detection Toolbox and Benchmark.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs MMDetection: Open MMLab Detection Toolbox and Benchmark

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:11.224870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:11.224870Z digest=sha256:4f1d77329fef159ecb3e6d329cce7d1229e2c0647b1614eeaf9b9921ade7f15d

Observation ea12eca6-b266-45dc-be83-1cae9595b108 · outbound

This paper cites Quantifying the Carbon Emissions of Machine Learning.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Quantifying the Carbon Emissions of Machine Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:11.392668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:11.392668Z digest=sha256:23ae949351bc540b7c39f9971335c231e1816b5961b02bd74054015db6b07cdd

Observation fa83d959-90c6-4d5a-91d8-f7ecd2af6b26 · outbound

This paper cites Attention Entropies.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Attention Entropies

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:23.194826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:17:11.624886Z digest=sha256:0fa2a10c6571fa8e33d62959ee96bbb1094e51accf97db0b137838ec77d88da7

Observation fb0eeb87-84ec-4df2-8f9f-f3d8bb842029 · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:22.707460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:17:11.899321Z digest=sha256:c18c8015be89bc105a5e61ed571301f5fa9463d9ccc70b30322bce9983c4496c

Observation b46e9033-6c7b-4434-83ac-9d401faf3c38 · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:22.273985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:17:12.054837Z digest=sha256:e7c22c28e4b5bde60424df10f93d68bc068c12a3b4e512523890183814c178e2

Observation 23b300f3-d459-43fd-81fd-970983943bcc · outbound

This paper cites Finally, we highlight the high quality of the attention maps of LUViT in the Imagenet-Segmentation dataset [Gao et al., 2022] in Section B.3.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Finally, we highlight the high quality of the attention maps of LUViT in the Imagenet-Segmentation dataset [Gao et al., 2022] in Section B.3

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:21.994766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:17:12.195882Z digest=sha256:5bf2057f242cac8fdf1cdc715f05edef564267663bb0cf5de8f28058cac8c62e

Observation 308f5b8b-7143-4025-8fdd-d4e5232c97ad · outbound

This paper cites The results of both our reproduction of Pang et al.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs The results of both our reproduction of Pang et al

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:21.721932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:17:12.304770Z digest=sha256:148b5ec549d5007b850ba7b09db0b4ff435e3593df2008c9542dec9e4b398922

Observation 2c539026-45a4-408b-8d8f-6f6784ee44f9 · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:21.364775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:17:12.514747Z digest=sha256:ad4d35c8168acd14b63fc7eecd3fe0810a1bea5e6bf6e5fc4dc12cab81a03a00

Observation 5dc585a5-23fb-4816-8628-515f380fa421 · outbound

This paper cites Each reported value is an average of three training runs with three seeds, (0, 1, 2), and the subscript ± denotes the standard error for each setting.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Each reported value is an average of three training runs with three seeds, (0, 1, 2), and the subscript ± denotes the standard error for each setting

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:20.534750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:17:12.805484Z digest=sha256:299a2df99b9b85335dd04cd69c620e4752c39a0ac9c93116c1abcd9cf80c0103

Observation e26f0bfe-e2bf-4b04-861c-9b75690a5919 · outbound

This paper cites A cell in the downsampled mask is assigned a value of 1 if it overlaps with the original high-resolution mask.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs A cell in the downsampled mask is assigned a value of 1 if it overlaps with the original high-resolution mask

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:20.254747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:17:12.974750Z digest=sha256:f116d75bf0832c390045ccef80a8cc57921c9e81713fc3944552372f2eb8dc46

Observation 839f31c0-7bc6-4924-8a4c-354758aaeef4 · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:19.894763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:17:13.145214Z digest=sha256:1f99cd1ba75ff9ebfcf3ab5610c8fedd3de50e2d5b0b582718697d2192ab8500

Observation c1b92f23-9be1-49a0-89cd-15577e5ca49c · outbound

This paper cites While LUViT differs from Pang et al.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs While LUViT differs from Pang et al

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:19.505420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:17:13.371752Z digest=sha256:e7f18abbadf92572d700b1f2b135edc4ace5150468cde8d743536e5e6108e09e

Observation da3ffe0c-fa7e-4a2a-9b93-62a4ef5b3ec6 · outbound

This paper cites The authors quantified this alignment through demonstrating improved gradient-signal-to-noise ratio (GSNR) under the presence of the LLM block.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs The authors quantified this alignment through demonstrating improved gradient-signal-to-noise ratio (GSNR) under the presence of the LLM block

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:19.186193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:17:13.524735Z digest=sha256:46c48e4edcefa01538b162f660d51156a702b0c62f167c2c09a0ebf5973ebfe1

Observation 5de44f84-dac2-4840-ac07-61247b0d0f55 · outbound

This paper cites Following up from this observation and taking inspirations from Tiwari and Shenoy [2023], Bai et al.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Following up from this observation and taking inspirations from Tiwari and Shenoy [2023], Bai et al

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:18.974743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:17:13.682924Z digest=sha256:b23d0de453c5d759b3fc95611a4ae640c0f3deabb816c94eb09c105a1a54e040

Observation b93c1078-e07f-4bf8-b17e-4ba73ffc0095 · outbound

This paper cites This auxiliary training objective distills the representations of the frozen-LLM-appended ViT to a vanilla ViT through a similarity loss in-between [Hinton et al., 2015].

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs This auxiliary training objective distills the representations of the frozen-LLM-appended ViT to a vanilla ViT through a similarity loss in-between [Hinton et al., 2015]

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:18.727982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:17:13.865294Z digest=sha256:17109b67f088df4dd739b034ea613d7e06c1b8142b19c72c6c2b1f94e97650e4

Observation d454685a-4cad-4bbe-8758-f8265300bd6a · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:17.959249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:17:14.464832Z digest=sha256:ba14982a64bd042f65470a1877e2b3933080a3f37c560ed6169d99a753321f9c

Observation 9fdb8a9a-67fb-40f9-a4c6-fe9a7102f9ac · outbound

This paper cites Notably, we utilize average pooling setting instead of relying on the [CLS] token for performing classification.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Notably, we utilize average pooling setting instead of relying on the [CLS] token for performing classification

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:17.384759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:17:14.713999Z digest=sha256:1e713bc64727b4b984f1cbeedcd2b85eb5d2885f5460716ffd46504f2dcba7ce

Observation 698b20cc-0325-4398-a81c-e2335a8c7cb1 · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:17.103304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:17:14.794778Z digest=sha256:383c86b143b971e0065ace3186dd7fe552bc970128282105a885ced761a080f0

Observation fc64965e-b720-4e62-8ec6-4f0240a71397 · outbound

This paper cites renditions.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs renditions

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:16.834523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:17:14.935499Z digest=sha256:c4b73246187f7e2d003aaecafb56bc03450cabdee3ef9d9e69fde72368ed2c8b

Observation e6e964db-c461-44ad-aad6-b0501c535c97 · outbound

This paper cites Finally, the additional capacity baselines in Section 4 all have additional linear projection layers at the head, analogously with where they are placed in LUViT.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Finally, the additional capacity baselines in Section 4 all have additional linear projection layers at the head, analogously with where they are placed in LUViT

Reference 512

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:17.747453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:17:14.573708Z digest=sha256:54bc80782ff6ac109959d30a9d7467bd185de7d89f3e151e4cd6a563c4e988dc

Observation 7d8a0460-13f5-4f0f-ae51-22d3e18d2b07 · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 768

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:18.423887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:17:14.072992Z digest=sha256:486670117c0dd5f28f4faf484fe5b4bae09cd00f2f220b5c258db24923e38c1b

Observation 75a915c7-eb95-4384-a911-b85a519eb630 · outbound

This paper cites Benchmarking Neural Network Robustness to Common Corruptions and Perturbations.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Benchmarking Neural Network Robustness to Common Corruptions and Perturbations

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:08.425779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:08.425779Z digest=sha256:b9a7019cdc3dc99a2968efec9c876a3d13803b0dae4a6c4644434d76edd78cd7

Observation 027f2cbc-e45a-4f23-a60b-184b396ade95 · outbound

This paper cites BEiT: BERT Pre-Training of Image Transformers.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs BEiT: BERT Pre-Training of Image Transformers

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:07.315858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:07.315858Z digest=sha256:18bde7f9d39e058d670572065c4db65e4a5c4c577905d9fac117e04ba553587a

Observation 47bfb469-ff1f-4a19-829b-6e2b16311c60 · outbound

This paper cites Decoupled Weight Decay Regularization.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Decoupled Weight Decay Regularization

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:10.424828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:10.424828Z digest=sha256:5733cb82fc88c46abfe1fec5407efb75e02f452eb5ccb0fa34375302c0a6a669

Observation d54a2bed-4601-40be-852d-3d6a8fecd470 · outbound

This paper cites EVEv2: Improved Baselines for Encoder-Free Vision-Language Models.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.482217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.482217Z digest=sha256:a0098c35c2bdbc41ead5d7504b84078655824336038a616121871624083cb395

Observation ffa1e9e9-f158-4635-8296-6e50ca07ab6e · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.984740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.984740Z digest=sha256:6acc0388d124587834bc74e639345b968ca0293a5b06c525694e80940e122feb

Observation 5d8588d1-12b7-4214-92d8-2a3c65addc25 · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:10.104743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:10.104743Z digest=sha256:15a780f107c7f6e6ae93010f42996b75e398884c7b80090e5f0c8fe330aba81d

Observation 3bb7e0bc-074e-4c44-b5fc-3728678743ab · outbound

This paper cites Noise or Signal: The Role of Image Backgrounds in Object Recognition.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Noise or Signal: The Role of Image Backgrounds in Object Recognition

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:08.554972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:08.554972Z digest=sha256:473e9bbf0157c9f101e74af0cebebe9e421ff6f193cca77d009c9643559266ad

Observation e2a04bf1-b2b9-4a7c-9aef-03e91098f537 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs LLaMA: Open and Efficient Foundation Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:06.272206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:06.272206Z digest=sha256:98f7caaf9eb7f835c67d45631a3b6a91e726a09ff629dd8e29fcd16543a6e666

Observation ebad85fb-f08e-4e2c-83cd-bf4aabd2f59e · outbound

This paper cites iBOT: Image BERT Pre-Training with Online Tokenizer.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs iBOT: Image BERT Pre-Training with Online Tokenizer

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:07.479266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:07.479266Z digest=sha256:2205c2cf379ac89774aa4e17ef80ceed2a1fa6caec9fbc6c8651d53c133c289f

Observation 8976a14b-d71a-4b71-a24b-64905332c8cf · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:06.519528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:06.519528Z digest=sha256:e88732f0f6dc9462f4966eb723f5c8024632d3192e75884caa338e05dca1f868

Observation 9c1ae321-34d2-4f7c-8dfb-f4ce852f1aee · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:07.660329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:07.660329Z digest=sha256:f2b1f9a59426e351d40bb1cfcc5926840d2bc9a5f2c9b2e57116146c6bba30ef

Observation 2f706a71-111d-4a0b-8e01-496b98b0c5c7 · outbound

This paper cites Vision as LoRA.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Vision as LoRA

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:07.858228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:07.858228Z digest=sha256:311db05352fa8e3c04e7c645f468a93e5d1813b0bc29b2dba775759c65883816

Observation ccd954c7-0920-4a4d-97e2-fd6224cf7758 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:06.711803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:06.711803Z digest=sha256:ef940d6b8a19217c814bdc99e3396586d3421f7624b433a81b58c0e367b60698

Observation 6d3a361d-a3ff-45e4-8a4f-53d1e34cc2da · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 4096

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:18.206814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T21:17:14.265854Z digest=sha256:6245944f79e3657ff09ee550ee16df60bc0e0de23f2b705dbd5da5387ae8405d

Pith citing papers

No inbound Pith citation observations are available.