Pith. sign in

Paper Citation Record · LEDGER

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs

As of 10 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2507.00754.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00754 v2

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:17:14.935499Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy16
  • unresolved38
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation cf57972b-3253-4d67-8cc2-8dbcf447dd24 · outbound

This paper cites Language models are few-shot learners.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Language models are few-shot learners

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:24.184422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:17:06.174747Z digest=sha256:1fac4116396334b827abd6369151b8c036639b975547e001539886f19920ac38

Observation 0d832a76-bf8a-4879-b8af-89e85431c105 · outbound

This paper cites In these works, they provided litmus tests for measuring how focused the attention patterns of particular models are and how they relate to model robustness.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs In these works, they provided litmus tests for measuring how focused the attention patterns of particular models are and how they relate to model robustness

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:22.984748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:17:11.783929Z digest=sha256:3b563d7fa2f8837ecc1693195e159ef65514dc6ae9202d2c51f191b703dd6a21

Observation 80bb878e-7eab-488d-bf54-fac98f9d577b · outbound

This paper cites the standard deviation of the accuracy values divided by the number of different seeds.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs the standard deviation of the accuracy values divided by the number of different seeds

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:20.966369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:17:12.633122Z digest=sha256:47e7c922b421284f52625ee797a0066c6e73cbc2d4fb62acd9a90af2a4aa3b9f

Observation 7df17b61-bd53-4908-bfe4-0825ff4c8666 · outbound

This paper cites Frozen Transformers in Language Models Are Effective Visual Encoder Layers.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Frozen Transformers in Language Models Are Effective Visual Encoder Layers

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T21:17:16.472633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:17:06.952393Z digest=sha256:78c86812dc2fc960f397e5e479c6779c301e129d8d405f303c14628aab9bb967

Observation 8067ec4b-4bd5-4eab-8942-b273bfb5247d · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:07.140211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:07.140211Z digest=sha256:23ad9600776e7fd905f91f540c6797aa1507510631fb35f00745c2078c726af9

Observation f0432910-0c1c-432f-b169-d2edd481baa3 · outbound

This paper cites Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Mono-InternVL: Pushing the Boundaries of Monolithic Multimodal Large Language Models with Endogenous Visual Pre-training

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:08.054862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:08.054862Z digest=sha256:fb3dfccdcc5dbec268a49095bf3658a194e76a60829dd05d68b3e9ff63d6c406

Observation a6a28afd-2ec9-47d8-a024-7db7890eb3fc · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs DINOv2: Learning Robust Visual Features without Supervision

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:08.132343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:08.132343Z digest=sha256:df97c764d126ec4fde4bdba484578fc75cf0c7d6400d051b3d4730ba2b3f7945

Observation f5907d37-48d9-4cc7-9e57-bc8d05756ac8 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Imagenet: A large-scale hierarchical image database

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:08.286198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:08.286198Z digest=sha256:9e07e46f76308700ddff6a08b4ae1dcea66bc8278f1551c53e84db6df643051d

Observation b3724c4e-8c25-47a4-8b1f-0025c935b7d9 · outbound

This paper cites Microsoft coco: Common objects in context.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Microsoft coco: Common objects in context

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:08.730823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:08.730823Z digest=sha256:e369fdf699dc87a9cfda875051aaf3402c490815aa090fb6f4422e070befeea6

Observation dbb7affd-7b08-4c85-9bcd-6eb3e47e0215 · outbound

This paper cites Attention Entropy is a Key Factor: An Analysis of Parallel Context Encoding with Full-attention-based Pre-trained Language Models.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Attention Entropy is a Key Factor: An Analysis of Parallel Context Encoding with Full-attention-based Pre-trained Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:08.954914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:08.954914Z digest=sha256:7fae8c20698835aa73ebe19a645bbb590c08cfa0691cb43a35dfe46f869ff583

Observation 54bfea88-6ae9-486d-80d8-fe4d093f67bc · outbound

This paper cites Torch.manual_seed(3407) is all you need: On the influence of random seeds in deep learning architectures for computer vision.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Torch.manual_seed(3407) is all you need: On the influence of random seeds in deep learning architectures for computer vision

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.069919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.069919Z digest=sha256:2639d0f137621b66b1b6e60285e1371f0041df4165ee75dd5e09385e5f6d79f1

Observation 6a846dd8-e028-4ee5-ad8d-aa0c10fafebe · outbound

This paper cites Understanding Why Neural Networks Generalize Well Through GSNR of Parameters.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Understanding Why Neural Networks Generalize Well Through GSNR of Parameters

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.234749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.234749Z digest=sha256:ef5283fcc958c94068ad049e3fb0a0d31fe56556e4ad5fbe94a7a505bc9f925a

Observation a964eac8-6510-4018-b4b1-d0208ec5c971 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Distilling the Knowledge in a Neural Network

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.364748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.364748Z digest=sha256:ed40c7c26220b4467ce094d17d5baca3ab8f0af99f8b1006a9a2fc0471e4d2b7

Observation c1b79421-ce28-4c47-b15b-bc68e211d958 · outbound

This paper cites SOLO: A Single Transformer for Scalable Vision-Language Modeling.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs SOLO: A Single Transformer for Scalable Vision-Language Modeling

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.614748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.614748Z digest=sha256:3fc3d16cfd61b933021165d2097e9620cc2dcc350e2e2f972478c59a5796dabb

Observation 7eb34fdb-15b0-4387-afd5-973902d3b467 · outbound

This paper cites Layer Normalization.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Layer Normalization

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.764750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.764750Z digest=sha256:8e5d5296fc80ddc16af7a62bb3ab42b8850e1d86646c8bdb039b133108c33534

Observation 36fc9a21-bae4-4c9b-82ca-5261a805f6d7 · outbound

This paper cites Adam: A Method for Stochastic Optimization.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Adam: A Method for Stochastic Optimization

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:10.275076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:10.275076Z digest=sha256:329cd3294e568ad32c73323bf5974494de738f1d915b0630a819206ef0f12fdd

Observation b9ac64ec-a47d-4723-b68d-338970fc7d57 · outbound

This paper cites Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Accurate, Large Minibatch SGD: Training ImageNet in 1 Hour

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:10.544749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:10.544749Z digest=sha256:b60be3505c65d611f8ded8e03f04c0daba0f421587c0e8f22df97da49b75371c

Observation 330ec6c4-24c3-4029-9c70-faf41c147242 · outbound

This paper cites Deep networks with stochastic depth.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Deep networks with stochastic depth

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:23.564748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:17:10.764825Z digest=sha256:cb5fb65c7c724099b4d89439e7e77efa78105296ee47ccd7c7a9ebf2ac04bb4f

Observation dde40791-2763-47ba-85c7-68ab4f603424 · outbound

This paper cites mixup: Beyond Empirical Risk Minimization.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs mixup: Beyond Empirical Risk Minimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:10.884728Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:10.884728Z digest=sha256:c5d1f488122a948bf5a00c5f282badfbca94abe285c3e276aee9e6bb516c6825

Observation e882166e-8e34-4822-a2a3-3639df5af93f · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:11.024571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:11.024571Z digest=sha256:6117fb84a585b9208d63941b3a491c202bf024779bb6e6f2adb73834c8db25e5

Observation 3551f281-1f14-42b8-9919-51b480181047 · outbound

This paper cites MMDetection: Open MMLab Detection Toolbox and Benchmark.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs MMDetection: Open MMLab Detection Toolbox and Benchmark

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:11.224870Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:11.224870Z digest=sha256:043da966b13b2038879a38e6f48f5f334a2b66cdb9c62e5085df80676bcd0bcb

Observation ea12eca6-b266-45dc-be83-1cae9595b108 · outbound

This paper cites Quantifying the Carbon Emissions of Machine Learning.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Quantifying the Carbon Emissions of Machine Learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:11.392668Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:11.392668Z digest=sha256:d44ec25e693575dc0e2df315c3b66e7c8f35fa150ed88df0bc2122ccb5e1f574

Observation fa83d959-90c6-4d5a-91d8-f7ecd2af6b26 · outbound

This paper cites Attention Entropies.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Attention Entropies

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:23.194826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:17:11.624886Z digest=sha256:fd2b96f77fb6a218d3f8dd050393f18c36ad275d6d8eab6d2fabc358d776d3da

Observation fb0eeb87-84ec-4df2-8f9f-f3d8bb842029 · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:22.707460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:17:11.899321Z digest=sha256:e57d645a0124264402eb9bb9c2080d808a20ebe33660befc94346e5140b69888

Observation b46e9033-6c7b-4434-83ac-9d401faf3c38 · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:22.273985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:17:12.054837Z digest=sha256:e5d072aa7a7435a00d1775a63a875ad65ef6656c1bf59e42f1dfb22723c01253

Observation 23b300f3-d459-43fd-81fd-970983943bcc · outbound

This paper cites Finally, we highlight the high quality of the attention maps of LUViT in the Imagenet-Segmentation dataset [Gao et al., 2022] in Section B.3.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Finally, we highlight the high quality of the attention maps of LUViT in the Imagenet-Segmentation dataset [Gao et al., 2022] in Section B.3

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:21.994766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:17:12.195882Z digest=sha256:89df122be4684d8ae7b668819ebb9d108b46ecbd36369928b4cbb58b13594aa2

Observation 308f5b8b-7143-4025-8fdd-d4e5232c97ad · outbound

This paper cites The results of both our reproduction of Pang et al.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs The results of both our reproduction of Pang et al

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:21.721932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:17:12.304770Z digest=sha256:97225a041a3bfef908a87a7422680c0bd67d03e6a701baa2393d13492c7a0cad

Observation 2c539026-45a4-408b-8d8f-6f6784ee44f9 · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:21.364775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:17:12.514747Z digest=sha256:d31cd48ecca459a8836aadeb8332b9ee7ec9851035c677e0cbbe4e571e9d7431

Observation 5dc585a5-23fb-4816-8628-515f380fa421 · outbound

This paper cites Each reported value is an average of three training runs with three seeds, (0, 1, 2), and the subscript ± denotes the standard error for each setting.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Each reported value is an average of three training runs with three seeds, (0, 1, 2), and the subscript ± denotes the standard error for each setting

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:20.534750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:17:12.805484Z digest=sha256:3ce3b369f76e56b68c787baece78cc1e1410ba9e5689a484809873259f446c27

Observation e26f0bfe-e2bf-4b04-861c-9b75690a5919 · outbound

This paper cites A cell in the downsampled mask is assigned a value of 1 if it overlaps with the original high-resolution mask.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs A cell in the downsampled mask is assigned a value of 1 if it overlaps with the original high-resolution mask

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:20.254747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:17:12.974750Z digest=sha256:42c917531f33a62630d4bc260b683f1cc987b1cb7cb68dc3c417d6efd59f6b57

Observation 839f31c0-7bc6-4924-8a4c-354758aaeef4 · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:19.894763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:17:13.145214Z digest=sha256:01261806507abca2befd7c5ed24661401d25a3a1acd6880d708caa83815a94e1

Observation c1b92f23-9be1-49a0-89cd-15577e5ca49c · outbound

This paper cites While LUViT differs from Pang et al.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs While LUViT differs from Pang et al

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:19.505420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:17:13.371752Z digest=sha256:73d28e929f691c85d0941c2071fcd8a5fc3c98548c3431dd98ab0aa28370477e

Observation da3ffe0c-fa7e-4a2a-9b93-62a4ef5b3ec6 · outbound

This paper cites The authors quantified this alignment through demonstrating improved gradient-signal-to-noise ratio (GSNR) under the presence of the LLM block.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs The authors quantified this alignment through demonstrating improved gradient-signal-to-noise ratio (GSNR) under the presence of the LLM block

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:19.186193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:17:13.524735Z digest=sha256:a50b8d06dd5a8834bcc894bf7aa3eed73ab17290693ebad1ac5032242a6dac20

Observation 5de44f84-dac2-4840-ac07-61247b0d0f55 · outbound

This paper cites Following up from this observation and taking inspirations from Tiwari and Shenoy [2023], Bai et al.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Following up from this observation and taking inspirations from Tiwari and Shenoy [2023], Bai et al

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:18.974743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:17:13.682924Z digest=sha256:c416c3b761eaae36c0f31164f45956ae069a6e839b79226cf5026ce98cad0374

Observation b93c1078-e07f-4bf8-b17e-4ba73ffc0095 · outbound

This paper cites This auxiliary training objective distills the representations of the frozen-LLM-appended ViT to a vanilla ViT through a similarity loss in-between [Hinton et al., 2015].

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs This auxiliary training objective distills the representations of the frozen-LLM-appended ViT to a vanilla ViT through a similarity loss in-between [Hinton et al., 2015]

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:18.727982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:17:13.865294Z digest=sha256:f1d4afb655bd47bd124765e42f44e21fa6c99bca999e15c2bffbfd14e15082e0

Observation d454685a-4cad-4bbe-8758-f8265300bd6a · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 51

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:17.959249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:17:14.464832Z digest=sha256:2183b203119dab9cd77d387cc582b7e16f5d61df3a0cd716e9618250c6c9ea84

Observation 9fdb8a9a-67fb-40f9-a4c6-fe9a7102f9ac · outbound

This paper cites Notably, we utilize average pooling setting instead of relying on the [CLS] token for performing classification.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Notably, we utilize average pooling setting instead of relying on the [CLS] token for performing classification

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:17.384759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:17:14.713999Z digest=sha256:ee4d70262892e7e10571ce560554c1d1fcb1a554fbeab3f63281b8c923feefdf

Observation 698b20cc-0325-4398-a81c-e2335a8c7cb1 · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:17.103304Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:17:14.794778Z digest=sha256:5365000d8a0e09bb504575dfeb37ac84a182c2108fc0e385f9b4daea53ec09ff

Observation fc64965e-b720-4e62-8ec6-4f0240a71397 · outbound

This paper cites renditions.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs renditions

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:16.834523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:17:14.935499Z digest=sha256:0ca4ddb8b77007b06cf057530af5a6aa6ae8364cee41358ee0c3f79374282525

Observation e6e964db-c461-44ad-aad6-b0501c535c97 · outbound

This paper cites Finally, the additional capacity baselines in Section 4 all have additional linear projection layers at the head, analogously with where they are placed in LUViT.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Finally, the additional capacity baselines in Section 4 all have additional linear projection layers at the head, analogously with where they are placed in LUViT

Reference 512

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:17:17.747453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:17:14.573708Z digest=sha256:845a5eb895891f5803fd5c58dec733424b2f3548601574d918aa29c9466f42b6

Observation 7d8a0460-13f5-4f0f-ae51-22d3e18d2b07 · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 768

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:18.423887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:17:14.072992Z digest=sha256:bd21f12f699e811cc676898a23802ca6c31b7551d87d88ec1a99b8535fd3c846

Observation 75a915c7-eb95-4384-a911-b85a519eb630 · outbound

This paper cites Benchmarking Neural Network Robustness to Common Corruptions and Perturbations.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Benchmarking Neural Network Robustness to Common Corruptions and Perturbations

Reference 2009

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:08.425779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:08.425779Z digest=sha256:61231b7dab629842ecff69af6921236b71dc9664bce2d102f8714b6a482b64ea

Observation 027f2cbc-e45a-4f23-a60b-184b396ade95 · outbound

This paper cites BEiT: BERT Pre-Training of Image Transformers.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs BEiT: BERT Pre-Training of Image Transformers

Reference 2010

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:07.315858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:07.315858Z digest=sha256:4a30b3c6d0ec992b4605dbecb889de17f81da8226a3d096ddf8c5d8ac74f201d

Observation 47bfb469-ff1f-4a19-829b-6e2b16311c60 · outbound

This paper cites Decoupled Weight Decay Regularization.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Decoupled Weight Decay Regularization

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:10.424828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:10.424828Z digest=sha256:72aad023dfff5b0fd21bff1b6816839bfd16e1161410e83b3d74dcfe3a2aa0a8

Observation d54a2bed-4601-40be-852d-3d6a8fecd470 · outbound

This paper cites EVEv2: Improved Baselines for Encoder-Free Vision-Language Models.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs EVEv2: Improved Baselines for Encoder-Free Vision-Language Models

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.482217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.482217Z digest=sha256:7d9e1bc51fe57dd230eceb5aeaf1ff9599ee57e41cd1ee12cdf66e28cb24a985

Observation ffa1e9e9-f158-4635-8296-6e50ca07ab6e · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:09.984740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:09.984740Z digest=sha256:9427c1582a07137ef00c774cb889b27b08dfcb85409869a9fe6eff17dc221942

Observation 5d8588d1-12b7-4214-92d8-2a3c65addc25 · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:10.104743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:10.104743Z digest=sha256:89a8da9fa97bfc905a6d6f64781ece87dc7ae984ef3b8a16883a8d2dece063c8

Observation 3bb7e0bc-074e-4c44-b5fc-3728678743ab · outbound

This paper cites Noise or Signal: The Role of Image Backgrounds in Object Recognition.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Noise or Signal: The Role of Image Backgrounds in Object Recognition

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:08.554972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:08.554972Z digest=sha256:2d6782ba0d9e1ce56ca450ee1a7bbc14477f2da296afe20a2d1b7a7a66074bd9

Observation e2a04bf1-b2b9-4a7c-9aef-03e91098f537 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs LLaMA: Open and Efficient Foundation Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:06.272206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:06.272206Z digest=sha256:349608adb7a0292ca9fbd383b6d99aa3588cb293a14b1cd6966b3d856fa84336

Observation ebad85fb-f08e-4e2c-83cd-bf4aabd2f59e · outbound

This paper cites iBOT: Image BERT Pre-Training with Online Tokenizer.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs iBOT: Image BERT Pre-Training with Online Tokenizer

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:07.479266Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:07.479266Z digest=sha256:24e840610fe6e9b62562034106b67b3f50a1f3f37550f12fc64267f5792e1015

Observation 8976a14b-d71a-4b71-a24b-64905332c8cf · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:06.519528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:06.519528Z digest=sha256:dd251bd6fcf6b9e5bcb929de8dd1c2b3ae8d947c9523c4c0cf5be3d8f67deebd

Observation 9c1ae321-34d2-4f7c-8dfb-f4ce852f1aee · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:07.660329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:07.660329Z digest=sha256:083e604ac2fb97bd530b54b5bc0a8f877b0e7e36707be7875a00d6aabffecc89

Observation 2f706a71-111d-4a0b-8e01-496b98b0c5c7 · outbound

This paper cites Vision as LoRA.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Vision as LoRA

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:07.858228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:07.858228Z digest=sha256:c240d9ece0a954314571cdc4d748a6dcc636358635ab7249e53e26e50644ecb5

Observation ccd954c7-0920-4a4d-97e2-fd6224cf7758 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T21:17:06.711803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:17:06.711803Z digest=sha256:3b15728e2624e689af806daebbc6859c55a6ad17ae8b58ac77b978abc5b35b85

Observation 6d3a361d-a3ff-45e4-8a4f-53d1e34cc2da · outbound

This paper cites an unresolved cited work.

Language-Unlocked ViT (LUViT): Empowering Self-Supervised Vision Transformers with LLMs Unresolved cited work

Reference 4096

Resolution
unresolved
raw_fallback, observed 2026-08-06T21:17:18.206814Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T21:17:14.265854Z digest=sha256:c2142b722f78a1a0db672386e73722b7f4f654ef381af404600858a21297a33f

Pith citing papers

No inbound Pith citation observations are available.