Pith. sign in

Paper Citation Record · LEDGER

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception

As of 18 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 1 inbound Pith citation observation for arXiv:2606.19584.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.19584 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-26T20:54:48.203293Z

measured 38 of 38 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-04T14:48:51.126553Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

37 of 37 outbound references displayed

  • verified exact23
  • verified fuzzy0
  • unresolved8
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b409a6b4-2b1c-469c-a19c-41e9306614ba · outbound

This paper cites Exploring Visual Prompts for Adapting Large-Scale Models.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception Exploring Visual Prompts for Adapting Large-Scale Models

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:19.219783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:cacf24d1a920e5bde096dccf456a74b3ccd7d93e3f6af6f52123fa3fa553a5cf

Observation 4d801b32-0b86-45eb-a176-91299175cc09 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception PaliGemma: A versatile 3B VLM for transfer

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:19.235400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:b615584ae3e8f925234a2ea5a765f808bb924b9aff5038dee2561b36a38d4106

Observation f65f8726-a036-4254-9f67-75bccbe48d5f · outbound

This paper cites All You May Need for VQA are Image Captions.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception All You May Need for VQA are Image Captions

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:49:19.182325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:4b391f28899901ae47d00354c3606d726fab9de2e5d3357b185421be1ed75010

Observation 04765852-f0e7-47bf-a8a0-0e3d3f4d71f5 · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:19.151225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:e4d73cca934e61cf9e3813674c3427d9bc96ffed75ecb3f22d7508ccac358e9b

Observation af8d776a-b5dc-4e54-93ff-56c205391feb · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:19.186065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:ba86d1e5d7c57f704be452a71e53bb8e06aeef09fc926b77e6b52da0a98676cc

Observation fe8d1ff0-b339-427f-a452-b870f0b87148 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.Advances in neural information processing systems, 36:49250–49267,.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception Instructblip: Towards general-purpose vision-language models with instruction tuning.Advances in neural information processing systems, 36:49250–49267,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-06-26T20:54:48.203293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:a707b1a34a5f9c1a05b2865b86641e0256779df5366feda38c8a9201c8a4d9e8

Observation d0a11f6c-9b33-4edd-a4bf-3dbaede5828b · outbound

This paper cites Selective Visual Representations Improve Convergence and Generalization for Embodied AI.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception Selective Visual Representations Improve Convergence and Generalization for Embodied AI

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:19.230707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:3f4c2878726e9c432f956f7dcb6e920232d0f7fb1e44e4bc8d5d9cc45800b52b

Observation 74603111-ced3-4d18-9828-7fd57917637a · outbound

This paper cites Data Filtering Networks.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception Data Filtering Networks

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:19.225762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:156ee334e8fdc08014ad776193730c7a845297898dd864185174d713b14c270c

Observation d0b878aa-956d-45f4-8003-365ac9d3cd15 · outbound

This paper cites The Llama 3 Herd of Models.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception The Llama 3 Herd of Models

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T00:49:19.239610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:e24b8703e22b4a1d9c78d2e3c022ba326f18310b7e5e26d9615635c1ed6b2c8f

Observation 500065c9-714e-448b-823f-4fa100ce157f · outbound

This paper cites E5-V: Universal Embeddings with Multimodal Large Language Models.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception E5-V: Universal Embeddings with Multimodal Large Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:19.214012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:dc8e17406ff6d0d336dd7cf37d23faf2941ae83af77572a275fa1ecbbba0c072

Observation 0f3e7447-3d28-4bdd-bc44-5b2744a5d4ff · outbound

This paper cites Referitgame: Referring to objects in photographs of natural scenes.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception Referitgame: Referring to objects in photographs of natural scenes

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-26T20:54:48.203293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:24482de45503468e2e65e5ebc86b42d42e2666b9e07091039dd8ea60fdc0cedd

Observation 2f15f18c-c861-4fad-bd23-32145fe9126d · outbound

This paper cites Modeling Caption Diversity in Contrastive Vision-Language Pretraining.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception Modeling Caption Diversity in Contrastive Vision-Language Pretraining

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:19.203184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:7ec1a58bb36b4fcfdf8f8df6c33bd57ae605dd6e46ec8379e217621ac8e047a6

Observation eb94859a-361c-4f75-bf42-d1c570f5915a · outbound

This paper cites Vx2text: End-to-end learning of video-based text generation from multimodal inputs.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception Vx2text: End-to-end learning of video-based text generation from multimodal inputs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-06-26T20:54:48.203293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:366f1c5df3aec986b0f4d84138133945c942c73292c93ad460a51a0bebe5ad5a

Observation db8e253e-c488-4afd-b7a7-b04411b59a66 · outbound

This paper cites Training-free deep concept injection enables language models for video question answering.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception Training-free deep concept injection enables language models for video question answering

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-26T20:54:48.203293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:f7ad05003b544ff44b9100389c010389f983d93b878dd5be3dc137f325af45ed

Observation 24da8478-d867-4092-ad40-ec9f08a7b5fd · outbound

This paper cites Understanding Zero-Shot Adversarial Robustness for Large-Scale Models.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception Understanding Zero-Shot Adversarial Robustness for Large-Scale Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:19.208818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:48beaeeeceb01bdd8f5edeaa150e478a49c094551acbb66cadd1c1fc15095b1f

Observation 075770d9-5c3b-40e0-99be-b6180696ee69 · outbound

This paper cites Visual Classification via Description from Large Language Models.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception Visual Classification via Description from Large Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:19.244441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:ea17d9819f89f6b3867c2d0ffc0b9a5e04fbc5e4d12f00d30369e6c3d3d7dffc

Observation 39dc97ab-cd15-4c86-b94d-147a7176662e · outbound

This paper cites Task Bias in Vision-Language Models.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception Task Bias in Vision-Language Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:19.264204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:2ac61ed3942a3e93f412badd02ff0ec04d9c9dff758b4cc43e284bb452007de4

Observation 0ed28888-7932-4c8c-af28-f5be5c797b1d · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception DINOv2: Learning Robust Visual Features without Supervision

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:19.186873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:67a76cfe69504573e08af23bce83f2f2d505100b3a91ac5125fc3f25a7593e90

Observation b09248ad-332a-46dc-94bf-8d04a110f576 · outbound

This paper cites Pre-training image-language transformers for open-vocabulary tasks.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception Pre-training image-language transformers for open-vocabulary tasks

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:49:19.212641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:2c9bfd747042f4403480681e9e30628582189c88183e817f0740df5b64babcde

Observation 67f47f13-970e-4356-a8aa-cecdc6daa478 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:19.177192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:843ae3f0a69a1034066f5c71b8127187683d8a27bb1e75d8ec1b43191e63543c

Observation b75a33ca-9a94-4554-8f86-a03749aacd6a · outbound

This paper cites X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception X-Former: Unifying Contrastive and Reconstruction Learning for MLLMs

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:19.166875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:c21ebc90011f9b8ae04e83c64779f941ff847968329c08c9d486647cba66fd8e

Observation 650ab919-5ed6-4535-aa2b-f874432f07df · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:19.230531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:e222b6543f8b98d154a5f853f5d76015339ce6ccd8ba01faf156d73e6265c865

Observation 6969ea20-7901-4607-abca-b9018cfccf80 · outbound

This paper cites Gemma 3 Technical Report.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception Gemma 3 Technical Report

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:19.190898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:98b44cd5095124b4feb4a077eea9fa66b9283c40f382f89f1583e6d7d587a781

Observation abfccb1c-3e97-4a79-b498-cd004b12dba1 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:19.172207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:4beb92195dc97b307675f5a68e9bdc29ad8a59a2ee05cf9c8bb16e170e0c7081

Observation bcf2fc55-31ca-4fa8-abe4-c2168bb3593c · outbound

This paper cites Scaling Pre-training to One Hundred Billion Data for Vision Language Models.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception Scaling Pre-training to One Hundred Billion Data for Vision Language Models

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:19.195832Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:efa802d1f598ee5f69af1ad1145c9734c1b93336f0d732e2e7f63a77776ed5b4

Observation 0b3c985a-f1f3-4b88-b7af-7c74314198ab · outbound

This paper cites Demystifying CLIP Data.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception Demystifying CLIP Data

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:19.181436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:73aefcf25871d292d0eff65822a475975ddd58262b2dfc4729b2723ff6c1b7ac

Observation cf4f49d8-1f70-489b-8683-ce1c00aa6e29 · outbound

This paper cites Large Language Models as Optimizers.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception Large Language Models as Optimizers

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:19.235077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:7b6055cf2f24b7d35442e72a2666c7a9a743caddb6745a572da21bbc70c2f935

Observation 950179fd-155c-499d-88e5-8da671c702dd · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:19.259387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:6cc793c53fc6502c1d0a0074054489e1696aecaf3579c7c250905721971ffcc5

Observation 34451644-0e85-47e9-980a-7eefb87abe42 · outbound

This paper cites Scaling vision transformers.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception Scaling vision transformers

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-26T20:54:48.203293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:e56570ea9b116407514fa9fb71b91c1492e330fd56f414df0ee8108446c78c0e

Observation 43eea2a8-91f0-4808-b12d-7c5e1088cdb6 · outbound

This paper cites MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception MagicLens: Self-Supervised Image Retrieval with Open-Ended Instructions

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:19.250279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:a358cce703ec2f3a18544442b59f3d94e0102b9cf5602c2b6e31876d53df1213

Observation c764cc1d-0431-4631-ac7a-56970e5dd1b3 · outbound

This paper cites VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception VISTA: Visualized Text Embedding For Universal Multi-Modal Retrieval

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:19.255293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:332f844d508753e76cd3e55d7482b030820cbeab36dced790e754b04a7a51bae

Observation c24c57e1-e3a7-4ef1-b4f8-eda8f79322fb · outbound

This paper cites However, its practical application and further development are subject to certain limitations, which also open avenues for future research.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception However, its practical application and further development are subject to certain limitations, which also open avenues for future research

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-26T20:54:48.203293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:6e2865398745be904e172657142be6ed8094d6fd0af6c766400a755f93962adf

Observation 3026f273-31af-4762-aecd-544733f4dfb6 · outbound

This paper cites (2021), SigLip Zhai et al.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception (2021), SigLip Zhai et al

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-26T20:54:48.203293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:9ed60eb4bacb20b453d5d19dde26e293367903f1ae3ea4f0863a55789789c3a1

Observation b31ba3f7-e307-4343-9c7d-5eac1b8b6016 · outbound

This paper cites Interestingly, by leveraging Gemini to evolve and generate different text prompts Yang et al.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception Interestingly, by leveraging Gemini to evolve and generate different text prompts Yang et al

Reference 34

Resolution
malformed identifier
no resolver link, observed 2026-06-26T20:54:48.203293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:c44a0764c1222060e08b3931ee984ef360fddf9c3926fedd4e82bbf56b82204b

Observation 4afd2aa5-0bde-41c6-a55b-97b803725cd8 · outbound

This paper cites an unresolved cited work.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception Unresolved cited work

Reference 35

Resolution
malformed identifier
no resolver link, observed 2026-06-26T20:54:48.203293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:468616bea971d4491d262ebe037e380dbba34a4f69b14d27fcf7a2e6c7991901

Observation f6d734c5-ba85-4df6-843c-595b9a9c37be · outbound

This paper cites The resulting distribution, visualized in Figure 13, reveals significant variations in instruction fre- quency.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception The resulting distribution, visualized in Figure 13, reveals significant variations in instruction fre- quency

Reference 36

Resolution
malformed identifier
no resolver link, observed 2026-06-26T20:54:48.203293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:502cf1ecec35ece0e51b7aa3680f812543c4031efdd512f663efd5b3ba016059

Observation 1f6a59b5-a165-44a2-99d0-43b7855c8550 · outbound

This paper cites an unresolved cited work.

Language-Instructed Vision Embeddings for Controllable and Generalizable Perception Unresolved cited work

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-26T20:54:48.203293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:54:48.203293Z digest=sha256:ccfe1ebb179e7b1f7a8bf3de435487132134638965c53a772a805470fdbe12ef

Pith citing papers

Observation bfdd4eff-2c2d-400e-bbbe-6f7e1f2e63ef · inbound

HAFI-VLM: A Frequency Perspective for Diagnosing and Enhancing Visual Perception in Vision-Language Models cites this paper.

HAFI-VLM: A Frequency Perspective for Diagnosing and Enhancing Visual Perception in Vision-Language Models Language-Instructed Vision Embeddings for Controllable and Generalizable Perception

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-04T14:48:51.126553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T14:48:51.126553Z digest=sha256:295970d386598c5b76cc9ce4067013e79b2a8a7bf6e9520b6822fdcd0f0e7f98