Pith. sign in

Paper Citation Record · LEDGER

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models

As of 10 August 2026, this Paper Citation Record lists 19 of 19 outbound references and 0 inbound Pith citation observations for arXiv:2607.16305.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.16305 v1

Coverage vector

measured 19 of 19 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T06:33:14.960504Z

measured 19 of 19 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

19 of 19 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 784c5693-f5f3-466a-be2e-2504a5d758a8 · outbound

This paper cites STEM: Scaling Transformers with Embedding Modules.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models STEM: Scaling Transformers with Embedding Modules

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.733672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.733672Z digest=sha256:ec83cb399c93d7080053f4f4629a2142ca4d65a4c70ff0e9280e28a145844467

Observation 8c04fa4d-8e33-418b-b576-0591770ed8f3 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:13.908702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:13.908702Z digest=sha256:ce6b7102dfd9589053f0ee05476b7da767dd23df7fa19fcd9add8c266f4e2549

Observation d30f5a49-2cbe-4992-96d7-716a8ed7b9b0 · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.240999Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.240999Z digest=sha256:68b11fd8c94f6f280969e98873f394a0455590086fac033609168ccad6aa7cc4

Observation 051a2367-bb40-4d66-9543-81b4c9972659 · outbound

This paper cites Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models Dosovitskiy, A.; Beyer, L.; Kolesnikov, A.; Weissenborn, D.; Zhai, X.; Unterthiner, T.; Dehghani, M.; Minderer, M.; Heigold, G.; Gelly, S.; Uszkoreit, J.; and Houlsby, N

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.292946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.292946Z digest=sha256:f760a9226d97610f8435c651b23bb4c8cfd20065265d39fa85c1a9cc224c1ee1

Observation 57187ae8-b034-49f7-b94c-ce96b817a597 · outbound

This paper cites Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models Infinity-MM: Scaling Multimodal Performance with Large-Scale and High-Quality Instruction Data

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.485161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.485161Z digest=sha256:6fd98b31329f1a208f27f21faf008ecffd86967471265b17de93a6fdfc59233c

Observation 6ff32895-0b63-44fc-a9a8-248689c4b32f · outbound

This paper cites an unresolved cited work.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.625608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.625608Z digest=sha256:dd04d37ab3e40b60c7884a1b2abcaaf7c880da58301b3959dcdfcc71eb53ac3f

Observation 54b5273f-5f12-4f84-afcf-746401790819 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.677496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.677496Z digest=sha256:9a64d0d99ef93e75ba939af0a1c60cdcbe7888bcda0e1f155aa27096d5b0012d

Observation 4598b522-f40d-4cee-a40e-a82abf03f44a · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.864080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.864080Z digest=sha256:d0e9d38765fb6ab1f8f97dd979f5d8fb6f88ae8f441369a4b06daf38277870d6

Observation 9c39c865-9dfc-45e6-b333-9540fc28ef60 · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models CogVLM: Visual Expert for Pretrained Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.945938Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.945938Z digest=sha256:6f36980cb8dfe147451f6469abfad0b5315d8d6856e67912083fc5eb1d4f3a29

Observation 90a3810e-e305-42f1-ad6f-b6eea24e9090 · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.950666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.950666Z digest=sha256:4287e756f5dba46ca9d44814e18500e904f6c91f2abbb397f10b85effd726a54

Observation 5a5680dd-462b-49e5-a284-ee7adb09ce1f · outbound

This paper cites Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models Tarsier2: Advancing Large Vision-Language Models from Detailed Video Description to Comprehensive Video Understanding

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.955758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.955758Z digest=sha256:4aa48d7ad540f8ca6a2d48df6fe980137b28aef8222fef3f99fedffee8e40b29

Observation e8af6555-174f-49e8-9fb3-4932eeb61e6d · outbound

This paper cites HiMix: Reducing Computational Complexity in Large Vision-Language Models.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models HiMix: Reducing Computational Complexity in Large Vision-Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.960504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.960504Z digest=sha256:85ba63417ee0031b02a5b2f8bd2b96879c982f7925f1f9a869aa9a2eadfd69c9

Observation 42d5e4bf-7495-4e17-83bf-3cfd89cc6bfc · outbound

This paper cites GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.797043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.797043Z digest=sha256:686e0499d5d55ccb46151c18e2cea63b70dfbe16cda9575d450d2a8d90efb22d

Observation 3caceb63-a5fb-44fe-8d15-e6ac4b03e2c4 · outbound

This paper cites Scaling Laws for Neural Language Models.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models Scaling Laws for Neural Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.541340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.541340Z digest=sha256:ab9d912c70bcf609158d38ee5b84370dd60cacf32a9b83fdd15c714ad35d752a

Observation 968911c8-70be-4eee-8090-d09f98d0d120 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.400943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.400943Z digest=sha256:f8f504ffcb00dc75c1b6e443c47250078551eafa0bbafdf536338a8ca1b88d26

Observation 4ec2d70e-32a6-46ad-8d09-a3446753db47 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:13.537859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:13.537859Z digest=sha256:696ac61cbdf21720baa1b87eb74fdfd19701794cb14701752d8074933a5b0d52

Observation f39be760-bfdb-4680-ac30-fe4801ecc271 · outbound

This paper cites DeepSeek-V3 Technical Report.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models DeepSeek-V3 Technical Report

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.128871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.128871Z digest=sha256:a01c37f11293ea4b6ba68a3f52e38dffe1b68f2109982f7cb0b1005bafb6faeb

Observation 6f5fc2fc-0d84-4509-993b-c3462782ca18 · outbound

This paper cites Qwen2.5-VL Technical Report.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models Qwen2.5-VL Technical Report

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:13.628554Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:13.628554Z digest=sha256:a1cb03c3ffd46e5499ac7eb6caad88b5b856d044747e200dc9e144fee2091814

Observation 33dd7702-4e23-4626-8262-b2c0afafed97 · outbound

This paper cites Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models Conditional Memory via Scalable Lookup: A New Axis of Sparsity for Large Language Models

Reference 2026

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:13.744408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:13.744408Z digest=sha256:451ab4d66e6b3bec333e4c3ca34d862f75c67683bc33908221e2d941fe79d64d

Pith citing papers

No inbound Pith citation observations are available.