Pith. sign in

Paper Citation Record · LEDGER

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders

As of 11 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2607.06445.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.06445 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-08T05:49:25.572256Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact13
  • verified fuzzy4
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch12

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation aad35dad-f979-4e0c-89a9-72819d3edc08 · outbound

This paper cites Qwen2.5-VL Technical Report.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Qwen2.5-VL Technical Report

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-08T05:54:33.573554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:f40513f6b6e7b1da613f4d569483ca3e095783d5cd583d2ad1308952b969de39

Observation 0e49c2a6-39cb-408a-ab15-2dfd325304d4 · outbound

This paper cites Dar, G., Geva, M., Gupta, A., and Berant, J.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Dar, G., Geva, M., Gupta, A., and Berant, J

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-07-08T05:54:33.547907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:4987c062e89f38e7cea05368b13d36febae1d9fb09df276c654cc49093dc28da

Observation 51cb6c9f-4942-40ed-a626-e84896284fc5 · outbound

This paper cites Analyzing Transformers in Embedding Space.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Analyzing Transformers in Embedding Space

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-08T05:54:33.545444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:7616dd2bf528c942b3d8260f031f636fdcdb3fa82285899706edc435ecaf7673

Observation 2d884595-620a-4dae-a9f4-eab353ce17d5 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-08T05:54:33.551545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:7758520f3af2c36fe05bb6a0893375d8a3c72a0b32cfe28ade4208b01d56012b

Observation 6a6d1ddc-11f7-4971-b729-9a09a6eb69e9 · outbound

This paper cites Transformer Feed-Forward Layers Are Key-Value Memories.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Transformer Feed-Forward Layers Are Key-Value Memories

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T05:54:33.542671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:6ec9fb9c5e65e86a27846db5aba180045566b0532c6088a46d9e13307a9bbedc

Observation c1d40b2a-e127-4743-b13c-8b960016e953 · outbound

This paper cites Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Transformer Feed-Forward Layers Build Predictions by Promoting Concepts in the Vocabulary Space

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T05:54:33.557100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:6c6473e0ba80bb2b8444a0bc335b0487c0f0e1e7add606d00d684220fe8c92ea

Observation 153eacbd-34eb-4737-9076-67a803745d13 · outbound

This paper cites Generating an image from 1,000 words: Enhancing text-to-image with structured captions.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Generating an image from 1,000 words: Enhancing text-to-image with structured captions

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-08T05:54:33.569864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:fd0e156633cf855db11731aafdbd6b18370e0c68bc204dbb18ebee3f5b614e2a

Observation 57ddb86a-554b-47ca-8298-635aba06c398 · outbound

This paper cites LTX-2: Efficient Joint Audio-Visual Foundation Model.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders LTX-2: Efficient Joint Audio-Visual Foundation Model

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T05:54:33.539837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:adde4b9353397a44dea99b6815c614903092c63bab4ef84b20b43e207c80a51f

Observation 86cf63ad-536e-4fb5-86c9-285f1322b384 · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-07-08T05:54:33.568137Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:4d420e5e0901ee3ee528179823c54a95befa99bd7d79aab2c4cadad83f71d049

Observation a7f0eec0-b394-48cd-8606-a5ede115aa8a · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders LoRA: Low-Rank Adaptation of Large Language Models

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-07-08T05:54:33.566698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:2ad19c5fded3d81e7bc32ecde076eabe026252c0fc22edba8df259c024356e79

Observation 67d74a6e-4e64-4164-a191-9bcd82c72c8b · outbound

This paper cites Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T05:54:33.562879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:8f75375c1975c36c4412de7f1f5fa3a09fe16cffdc26914b03aa0207ab33ec0f

Observation f099b476-2c82-4664-bae0-8de9a7610a42 · outbound

This paper cites What's in the Image? A Deep-Dive into the Vision of Vision Language Models.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders What's in the Image? A Deep-Dive into the Vision of Vision Language Models

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-07-08T05:54:33.577235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:055ddf89e47bf2a42c30bb80d241f987ff5f5effa94ff6476916aead4a8acc71

Observation 1a2f12fb-9cd9-43d0-af2f-1dc30bcc99ab · outbound

This paper cites FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders FLUX.1 Kontext: Flow Matching for In-Context Image Generation and Editing in Latent Space

Reference 13

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T05:54:33.547841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:2c5dd5e093d20efaacdfe049d4e4e6b1acbdf11e89abe8f1f8f47f609b72dfce

Observation dd27a126-1e67-4c4f-acd4-0443c975e361 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T05:54:33.523101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:2c77c346ecfd8fbee860c60e3198beb7cc1e64a3ac445039862605973fdf495d

Observation 5a879df2-b80d-49ff-827a-d00aea8bb235 · outbound

This paper cites A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders A Survey of State of the Art Large Vision Language Models: Alignment, Benchmark, Evaluations and Challenges

Reference 15

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T05:54:33.572461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:58730d9c13675ed45b5f560f7d11396883ff0a578e3039900af6e250960006c3

Observation 36b15485-8046-4abd-9af9-931f173d5ae2 · outbound

This paper cites Evaluating Text-to-Visual Generation with Image-to-Text Generation.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Evaluating Text-to-Visual Generation with Image-to-Text Generation

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T05:54:33.536126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:84a30d9ef83148b24c6414fb12b4679ccf4acb6706a63e0a8a4f16ddaefe0be9

Observation e9ae30cd-31e1-49b4-b55f-ea999ddbf0a9 · outbound

This paper cites Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Playground v3: Improving Text-to-Image Alignment with Deep-Fusion Large Language Models

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T05:54:33.575294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:78354c79d1c01e40dc1cd43ab933c7084988ae7c35d71b0dfa1c5a9dc483c1e7

Observation f58b044a-45d3-4f44-b54d-0e9b11fbc372 · outbound

This paper cites Seeing but not believing: Probing the disconnect between visual attention and answer correctness in vlms.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Seeing but not believing: Probing the disconnect between visual attention and answer correctness in vlms

Reference 18

Resolution
metadata mismatch
arxiv_id, observed 2026-07-08T05:54:33.526094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:446a68121972167ef3ee7e5e2019c756b4784fd385e53fbca887e6baa9cad7f4

Observation 5125282e-3e2d-40f2-9cec-e3e95d556f27 · outbound

This paper cites Linearly Mapping from Image to Text Space.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Linearly Mapping from Image to Text Space

Reference 19

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T05:54:33.565640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:9e83986b0994f2400501717217f0b926af4418b93c7ab39c89456149d7d35819

Observation 1e7fcd8f-d830-4ffa-ad08-ca62f0b97579 · outbound

This paper cites Towards Interpreting Visual Information Processing in Vision-Language Models.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Towards Interpreting Visual Information Processing in Vision-Language Models

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-07-08T05:54:33.570640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:2ab0275ae924e15aaddf798bf358c63fb8508259319e512758e6811783b1cd95

Observation 481f6402-548c-417a-b779-a5103a20432d · outbound

This paper cites Same task, different circuits: Disentangling modality-specific mechanisms in vlms.arXiv preprint arXiv:2506.09047, 2025a.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Same task, different circuits: Disentangling modality-specific mechanisms in vlms.arXiv preprint arXiv:2506.09047, 2025a

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-07-08T05:54:33.554573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:a7cae64650387dc8e2d94806c659a033a862407cbc7ba4167e90805333822a97

Observation 80aa4da0-cffb-44f4-9d95-5807e94bcce7 · outbound

This paper cites Rezatofighi, H., Tsoi, N., Gwak, J., Sadeghian, A., Reid, I., and Savarese, S.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Rezatofighi, H., Tsoi, N., Gwak, J., Sadeghian, A., Reid, I., and Savarese, S

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:54:33.883501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:448315472ec06e6fd516262cdc41fc5b409e4fa9dc71e3b03d8c3c8f007ef5d9

Observation b8c777e6-c106-4ede-88e0-ee5a5a87e18e · outbound

This paper cites High-Resolution Image Synthesis with Latent Diffusion Models.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders High-Resolution Image Synthesis with Latent Diffusion Models

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-07-08T05:54:33.553301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:753c9fee56f34e127b610a8cfa8c46601ec6968b5364d349cace84d3d9f1574a

Observation 00ecc256-a3cd-4a70-97ea-4ffd0be5259c · outbound

This paper cites Multimodal Few-Shot Learning with Frozen Language Models.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Multimodal Few-Shot Learning with Frozen Language Models

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-07-08T05:54:33.580326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:b6638aee315f33454d470515227792a38a1a5527be577ca02909d8b644517cb5

Observation 7a15bd3c-2bc1-4997-b72a-a2b2311f9306 · outbound

This paper cites OmniEdit: Building Image Editing Generalist Models Through Specialist Supervision.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders OmniEdit: Building Image Editing Generalist Models Through Specialist Supervision

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-07-08T05:54:33.521874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:72c0706729c4ca7e09e68ad98a6807769869e6b23b204f81eeadfe72d78b814b

Observation 8311148d-6b60-4cda-b5cc-6498e1eaee1d · outbound

This paper cites The Unreasonable Effectiveness of Deep Features as a Perceptual Metric.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders The Unreasonable Effectiveness of Deep Features as a Perceptual Metric

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-07-08T05:54:33.519294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:97755689bdc0edb399e2a554c4b2cae38f645e1d1b1192e4c95a603af464353b

Observation 9ffb9559-89cb-44ed-96c6-c0135f99bc60 · outbound

This paper cites We choose FIBO due to its strong adherence to spatial layouts.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders We choose FIBO due to its strong adherence to spatial layouts

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:54:33.887144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:ca3dadd21732e60c3881b816d1892a5ac25501cedd564fdca2806906c5c9598a

Observation 98cc5f86-8a70-4af4-b166-dd5f26f380bf · outbound

This paper cites To stabilize the initial training phase, we apply a warmup period of 100 steps during which the model is optimized using only the L1 loss.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders To stabilize the initial training phase, we apply a warmup period of 100 steps during which the model is optimized using only the L1 loss

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:54:33.885324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:f9f6322fa3c2a2daa876a35b09611620fb0f9897a8a39f1940c070a817cdb59d

Observation 896f0a33-f442-4b4d-a9f4-1d83dfc38318 · outbound

This paper cites Recolour the bottom right clownfish to be black and white.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders Recolour the bottom right clownfish to be black and white

Reference 29

Resolution
malformed identifier
arxiv_id, observed 2026-07-08T05:54:33.563947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:32571c5b74e9f47912263400f5d6d2c7751a1533b7c3af396f6a83575e9bb709

Observation e62b7b48-034c-4947-bb26-0b5a997dc9ac · outbound

This paper cites All automated Vision Question Answering evaluations and prompt generations utilizing the Gemini 2.5 Pro API (Gemini Team, Google,.

Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders All automated Vision Question Answering evaluations and prompt generations utilizing the Gemini 2.5 Pro API (Gemini Team, Google,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T05:54:33.881647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-07-08T05:49:25.572256Z digest=sha256:58e9d16e7bd561984fb3e29e27284cee3eb4dd23127a0ce29cce862902b4b7f4

Pith citing papers

No inbound Pith citation observations are available.