Pith. sign in

Paper Citation Record · LEDGER

Visual Lexicon: Rich Image Features in Language Space

As of 14 August 2026, this Paper Citation Record lists 89 of 89 outbound references and 1 inbound Pith citation observation for arXiv:2412.06774.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.06774 v1

Coverage vector

measured 89 of 89 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T19:22:55.635316Z

measured 90 of 90 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:54:09.796839Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T18:54:12.830432Z

Reference resolution

89 of 89 outbound references displayed

  • verified exact0
  • verified fuzzy47
  • unresolved42
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fa53fbcf-cfba-451d-b919-ac685e9b21d2 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Visual Lexicon: Rich Image Features in Language Space Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:54.902029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:54.902029Z digest=sha256:defaaf13539d057660ea403b7f5329fb2bce7240328a1a73d0b64ffc72bf744b

Observation df5898d1-f414-46de-844d-3616b039a90c · outbound

This paper cites BEit: BERT pre-training of image transformers.

Visual Lexicon: Rich Image Features in Language Space BEit: BERT pre-training of image transformers

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:54.911285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:54.911285Z digest=sha256:3953f90dab1ec1c9a0d8e8fe125aff0e3e840711cab48f7c03141b7c42a6c779

Observation 8110e71e-2571-433c-812b-960dd4e6a763 · outbound

This paper cites Label-efficient se- mantic segmentation with diffusion models.

Visual Lexicon: Rich Image Features in Language Space Label-efficient se- mantic segmentation with diffusion models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:54.918280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:54.918280Z digest=sha256:a0e51d9b76635dae7d01bf7e232d1881115be482e790fd16e65d32c05f5a0afe

Observation 11a12ab5-d0aa-4f21-b0e2-42b25e537971 · outbound

This paper cites Generalized denoising auto-encoders as generative models.

Visual Lexicon: Rich Image Features in Language Space Generalized denoising auto-encoders as generative models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:54.925181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:54.925181Z digest=sha256:3910e5c648b5c5fe19d1a2f36855b67edf1da36bb949875f78bb0d03c88c0a9a

Observation 2bf1a14f-f047-4925-b72f-44ddb677ce87 · outbound

This paper cites Improving image generation with better captions.

Visual Lexicon: Rich Image Features in Language Space Improving image generation with better captions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:54.931763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:54.931763Z digest=sha256:d57aa1957091a6b38176d3f9e573abb9a738697552aea1493aeb8b9e233f62c6

Observation 187cf068-12cc-4afb-89fd-f1b572370631 · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Visual Lexicon: Rich Image Features in Language Space PaliGemma: A versatile 3B VLM for transfer

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:54.939829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:54.939829Z digest=sha256:dfbe0898f0b12a3c318b83a6a3db3aff23a99b5213d049dd6f04373d58f7b9fc

Observation a7233933-ab07-44fb-91cf-bdef1edd0a01 · outbound

This paper cites Language Models are Few-Shot Learners.

Visual Lexicon: Rich Image Features in Language Space Language Models are Few-Shot Learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:54.949489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:54.949489Z digest=sha256:7186cb9a34ca24d89fcc0dab08a4a574aadd2e508008aa55ec5619960f36f95d

Observation 34f3191c-b220-4d09-b037-40fae470963c · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

Visual Lexicon: Rich Image Features in Language Space Emerg- ing properties in self-supervised vision transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:54.957417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:54.957417Z digest=sha256:8b2afa9a85d3ec9d9d46b8ab104211008f1d45eba338703210facfc69f1ed976

Observation 1cf3a614-1b2c-4dd0-b362-f39cbf9c9fa8 · outbound

This paper cites Generative pre- training from pixels.

Visual Lexicon: Rich Image Features in Language Space Generative pre- training from pixels

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:54.966657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:54.966657Z digest=sha256:f2212bfc6fffb074ee494d8624502d0462cc65d386d5fbde6b9aae294fbee2ed

Observation ed4090ed-fc6a-48d9-acaf-358c37ae14e3 · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Visual Lexicon: Rich Image Features in Language Space Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:54.975741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:54.975741Z digest=sha256:e22573bd1521450dbf614d1dcd0b12c367c6839c6b2ec2636546fbab14151a1e

Observation 4471d54f-f548-4395-83ad-6a7ecc68b03c · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

Visual Lexicon: Rich Image Features in Language Space PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:54.983804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:54.983804Z digest=sha256:83b95166839345f2652e139fb899a780d7aea9c44d3e01df9bf043993daaf279

Observation 5ed146b8-028f-4577-a6e2-05d74d7da021 · outbound

This paper cites Deconstructing Denoising Diffusion Models for Self-Supervised Learning.

Visual Lexicon: Rich Image Features in Language Space Deconstructing Denoising Diffusion Models for Self-Supervised Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:54.993646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:54.993646Z digest=sha256:e32033cb0033b0bfdc05cb4f86b3ef04d4c60f269fa3d275b740aaf63bb73b46

Observation 7541b5ae-c73c-49bb-853f-026e0c757138 · outbound

This paper cites Reproducible scal- ing laws for contrastive language-image learning.

Visual Lexicon: Rich Image Features in Language Space Reproducible scal- ing laws for contrastive language-image learning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.003755Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.003755Z digest=sha256:834f78b7315bd288405317d02fd5f593f66d477a3bb32fcc767a850c6e19f198

Observation 5f2108ee-e53b-4aaa-9503-3a674420d1c9 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Visual Lexicon: Rich Image Features in Language Space Imagenet: A large-scale hierarchical image database

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.010841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.010841Z digest=sha256:4a61bdec0c5b461076229e4759ce5d31b6de839e35a380b4d2952e18eeddf255

Observation 03c98d72-164b-46e9-8f91-76a18d6c12ef · outbound

This paper cites Large scale adversarial representation learning.

Visual Lexicon: Rich Image Features in Language Space Large scale adversarial representation learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.017006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.017006Z digest=sha256:8cb41a959cc6e5941ad90195d3265db6fc7610bafe73b53f37da7ca6113ae753

Observation dfbf0931-9762-45ec-a0db-b6e1c13dcffa · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Visual Lexicon: Rich Image Features in Language Space An image is worth 16x16 words: Transformers for image recognition at scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.023580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.023580Z digest=sha256:875e0391893545a3449784e4d4f4c4f827cf2ea1acfda62a57cf724590a73909

Observation 6ccb5774-6814-46c8-833f-2be2bdda1bf9 · outbound

This paper cites A new algorithm for data compression.

Visual Lexicon: Rich Image Features in Language Space A new algorithm for data compression

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:57.410098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.029405Z digest=sha256:24ae2dd8b6ede70ac00a953f8d551d512fb1e94c3ab8ba47d2bd1ad01c42e243

Observation af91a1d3-e1e5-4ffb-918d-634b705575b9 · outbound

This paper cites An image is worth one word: Personalizing text-to-image gen- eration using textual inversion.

Visual Lexicon: Rich Image Features in Language Space An image is worth one word: Personalizing text-to-image gen- eration using textual inversion

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:57.383161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.034815Z digest=sha256:32563cda04f231ac113c35015392a8a60cc4d33e6bb23eccc46a21ff6e34ed8d

Observation 1d414d07-28c6-4f8d-89ae-da66703c8d72 · outbound

This paper cites Instructcv: Instruction- tuned text-to-image diffusion models as vision generalists.

Visual Lexicon: Rich Image Features in Language Space Instructcv: Instruction- tuned text-to-image diffusion models as vision generalists

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:57.352919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.046351Z digest=sha256:87287372036327cdac6234f43b947ab5c84fc38a8c14e6bc4cde91230ee1473a

Observation 9ebe5898-7155-478e-9f7c-984e83ac06bf · outbound

This paper cites Imagebind: One embedding space to bind them all.

Visual Lexicon: Rich Image Features in Language Space Imagebind: One embedding space to bind them all

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.056655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.056655Z digest=sha256:bacb004f7e6637820fa778013d9655de12981a0ba6f328a450d9c732a4fca3de

Observation 8fcdbec4-24b4-4a5d-9cd8-4672cdd60922 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

Visual Lexicon: Rich Image Features in Language Space Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:57.313708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.065525Z digest=sha256:980e7b1f596b72e08e2059a2acf045f788f66d90b00ea28aa479ca34320cda46

Observation 3d9bdb7a-b3b2-4aba-8c93-62428a98577e · outbound

This paper cites Vizwiz grand challenge: Answering visual questions from blind people.

Visual Lexicon: Rich Image Features in Language Space Vizwiz grand challenge: Answering visual questions from blind people

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.072741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.072741Z digest=sha256:5cff2b1d0ef29f74cacf1ea87a3eafceb1f14a7d4cb8ab901973ba7b101ee410

Observation dfda47b2-9d57-4352-bc29-30acf26db66c · outbound

This paper cites Masked autoencoders are scalable vision learners.

Visual Lexicon: Rich Image Features in Language Space Masked autoencoders are scalable vision learners

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:57.281433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.078927Z digest=sha256:bcd34ae4d6c6d1e845875ad5a7f875674337be199675e70dba152fb4016b4626

Observation 7bc44f81-38ce-4805-bc07-4364fc897ec9 · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilib- rium.

Visual Lexicon: Rich Image Features in Language Space Gans trained by a two time-scale update rule converge to a local nash equilib- rium

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.089100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.089100Z digest=sha256:86759866e39b424584e23e4cdc860a2988c4d749c09a2679ffb755bd1b0f2b5f

Observation ec23e6a0-1c94-4aab-a2a6-dc8159441e07 · outbound

This paper cites Autoencoders, mini- mum description length and helmholtz free energy.Advances in neural information processing systems, 6, 1993.

Visual Lexicon: Rich Image Features in Language Space Autoencoders, mini- mum description length and helmholtz free energy.Advances in neural information processing systems, 6, 1993

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:57.250336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.098671Z digest=sha256:338aabe8b2431a2e33a2952a8ceebd63be85bda878b5df6c35853db532c62e8c

Observation f8c2c37b-6659-46f5-b568-2aeafbbc2a9a · outbound

This paper cites Classifier-free diffusion guidance.

Visual Lexicon: Rich Image Features in Language Space Classifier-free diffusion guidance

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.108679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.108679Z digest=sha256:6d0bf6bc3ee975dc796cfb74234373e991c283bc45df124b54ee5497289fa7be

Observation 38770467-6fb6-47fe-a205-3d69acc679aa · outbound

This paper cites Denoising dif- fusion probabilistic models.

Visual Lexicon: Rich Image Features in Language Space Denoising dif- fusion probabilistic models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.115996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.115996Z digest=sha256:25615467369288c5176144efe9ac1599e9b66faebb4750c41deeffa350ff9691

Observation bd3f989b-bf23-49db-b91e-65274e5f4643 · outbound

This paper cites SciCap: Generating Captions for Scientific Figures.

Visual Lexicon: Rich Image Features in Language Space SciCap: Generating Captions for Scientific Figures

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.125210Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.125210Z digest=sha256:ce051a4e6f0ed1031053b6c47811df982e347067b8f1440d04504d11bb87c63e

Observation 77acda40-70fd-45e1-8a62-9ffb034197ec · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Visual Lexicon: Rich Image Features in Language Space LoRA: Low-Rank Adaptation of Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.132713Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.132713Z digest=sha256:929e4f253e9298ef88d8ede480fd5c104030b7353d1e74110d0d4242e8b48992

Observation 820a3d38-192c-4903-8b28-2b68487e3ded · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Visual Lexicon: Rich Image Features in Language Space Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:57.205052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.140234Z digest=sha256:e31a662437f7f725ced9e1b255b7d2fd980b58f0c2b637bbf806b171758adf76

Observation 214ad2ed-739b-401b-a78e-4a7e7ceb5abb · outbound

This paper cites Soda: Bottleneck diffusion models for representation learning.

Visual Lexicon: Rich Image Features in Language Space Soda: Bottleneck diffusion models for representation learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:57.187455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.153036Z digest=sha256:f26aa19e5a30d736a0a1a827d1d458837ced06adfe025b126f67800dd60752e1

Observation d048583f-2d3a-4248-995e-33dbe42e710a · outbound

This paper cites Open- clip, 2021.

Visual Lexicon: Rich Image Features in Language Space Open- clip, 2021

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:57.168440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.163968Z digest=sha256:870c37cc148dfdd7e09efa21d8577952f02fb60725865ef4c6bf81b989a8cfb7

Observation cbb5b517-af0f-4458-aca6-fa74ef836242 · outbound

This paper cites Referitgame: Referring to objects in pho- tographs of natural scenes.

Visual Lexicon: Rich Image Features in Language Space Referitgame: Referring to objects in pho- tographs of natural scenes

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:57.145340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.171855Z digest=sha256:f032dd0a7abc754c51d4fa3cc22ede29a6161bda15d02b239285a361e3417d2f

Observation 48d02701-c186-4f6d-a14f-6b27347adc43 · outbound

This paper cites Auto-Encoding Variational Bayes.

Visual Lexicon: Rich Image Features in Language Space Auto-Encoding Variational Bayes

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.180039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.180039Z digest=sha256:816918eb0d1256ad92672bdb36d624fb6d682a2a6782a8fdd4035a03e16276c5

Observation 5722e5a2-58b1-469f-8f16-4016ff76d8ee · outbound

This paper cites Your diffusion model is secretly a zero-shot classifier.

Visual Lexicon: Rich Image Features in Language Space Your diffusion model is secretly a zero-shot classifier

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.189781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.189781Z digest=sha256:fa075f374b2e08148fe72546455e1f36c1e22ad800e0de5939aa187f24f37418

Observation 60760f2a-1876-4691-81ba-2bfdefc76522 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

Visual Lexicon: Rich Image Features in Language Space LLaVA-OneVision: Easy Visual Task Transfer

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.196482Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.196482Z digest=sha256:b3fab6710ac8831e2c9518c5b116f2048fc8fa8cf826049a91686fb34f16fb2b

Observation 0f506868-c561-40dd-ba55-39aa22e4f875 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation.

Visual Lexicon: Rich Image Features in Language Space Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:57.108917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.203918Z digest=sha256:e1ca22184a96a62015536bc395d78f71f0b6660e3e3e9bdbfa76c0ebcca4eb30

Observation 2881563e-06ee-46ce-a4d5-84293a2247eb · outbound

This paper cites ImageFolder: Autoregressive Image Generation with Folded Tokens.

Visual Lexicon: Rich Image Features in Language Space ImageFolder: Autoregressive Image Generation with Folded Tokens

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.217594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.217594Z digest=sha256:be811b2125f1b505da25dffa798be77b22e3bea473b2c3518e6bb90514d047f4

Observation 4234095f-c9a6-412e-920f-8b00f32cbcad · outbound

This paper cites Vila: On pre-training for visual language models.

Visual Lexicon: Rich Image Features in Language Space Vila: On pre-training for visual language models

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:57.090493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.226402Z digest=sha256:678fdb49ed0d1d5eec0664d722b4d341476f41bccc2da045082845329c7347c4

Observation 51f159c9-6975-4607-9083-ce26e41d3649 · outbound

This paper cites Microsoft coco: Common objects in context.

Visual Lexicon: Rich Image Features in Language Space Microsoft coco: Common objects in context

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:57.071696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.231979Z digest=sha256:ecc2d77664653a896f8c96d15403f35c9beaccdb4e667695df5a90071fe49f95

Observation 31ef5b21-af4d-4adb-af7c-7cf765f11b39 · outbound

This paper cites Improved baselines with visual instruction tuning.

Visual Lexicon: Rich Image Features in Language Space Improved baselines with visual instruction tuning

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.237076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.237076Z digest=sha256:8b5ede883e0ab4169e795e1013a95d294f6e5d1a8364103a5655ff7aa288294a

Observation 73a78549-8fda-415f-b5fd-31b0fe850442 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Visual Lexicon: Rich Image Features in Language Space Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:57.038651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.244513Z digest=sha256:3d8cacb52e9af468515a10a7cedb7763d49de1625612c40c31cd1c3b1cd585fb

Observation 16f3383e-3957-4fb6-87bd-f7260ef7f4b9 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

Visual Lexicon: Rich Image Features in Language Space Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.251306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.251306Z digest=sha256:dbb8ce309dfb27264d56f69040aba3d35242a7d1a3de19042ce19eead465d7b9

Observation e9a4194e-ccde-4b59-b7a6-20b105456862 · outbound

This paper cites Prompting hard or hardly prompting: Prompt inversion for text-to-image diffusion models.

Visual Lexicon: Rich Image Features in Language Space Prompting hard or hardly prompting: Prompt inversion for text-to-image diffusion models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:57.004154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.256987Z digest=sha256:e32b857a8250b424848792a53930d0ca85bea8b368e41341b896db23f823f6e3

Observation 2c298753-c6e2-42fc-aa86-fc2884543f2c · outbound

This paper cites Generation and comprehension of unambiguous object descriptions.

Visual Lexicon: Rich Image Features in Language Space Generation and comprehension of unambiguous object descriptions

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.985685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.265281Z digest=sha256:41ee67455397af0f3b24cacb8329debec5e644d56118a6540c909129d6018cc7

Observation 11192204-bfe0-41cd-99de-3e2e4013ef88 · outbound

This paper cites Ok-vqa: A visual question answering 10 benchmark requiring external knowledge.

Visual Lexicon: Rich Image Features in Language Space Ok-vqa: A visual question answering 10 benchmark requiring external knowledge

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.966163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.273236Z digest=sha256:87b3861397c4654edbe45080a0504ddbefadc16d6d4141ac979c560b8ce99a00

Observation 343df608-9cd4-4d54-8b63-dabefd91fb01 · outbound

This paper cites Finite scalar quantization: Vq-vae made simple.

Visual Lexicon: Rich Image Features in Language Space Finite scalar quantization: Vq-vae made simple

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.947678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.282252Z digest=sha256:78174921cc4025dda1c74e76e6d9f40b30c7729289f6b42ded5447cbd5b40138

Observation 7e9a96d0-4db1-4a8f-85d1-d0e9c6846b71 · outbound

This paper cites Null-text inversion for editing real im- ages using guided diffusion models.

Visual Lexicon: Rich Image Features in Language Space Null-text inversion for editing real im- ages using guided diffusion models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.288368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.288368Z digest=sha256:9ab4d104590cb1cf168ecc8e6b676c1ccfb2c6127acc2e0cb26a76507c3cdb6f

Observation 25210b1e-3cda-4333-a301-8ae368bff820 · outbound

This paper cites Improved denoising diffusion probabilistic models.

Visual Lexicon: Rich Image Features in Language Space Improved denoising diffusion probabilistic models

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.917686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.297128Z digest=sha256:a9e44ee9a1a2754a29676a522dff0a224b01c46c2902f5413c801053dc195f83

Observation c3683f6d-6730-4d9f-b0d9-bcad9a0e70ad · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Visual Lexicon: Rich Image Features in Language Space DINOv2: Learning Robust Visual Features without Supervision

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.305714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.305714Z digest=sha256:8027b445c949d976de738f94e3550bf7a085b6d0ea6a07e79603e2a5c9cc267e

Observation 780337ba-4573-47fb-973f-c07c08df00a6 · outbound

This paper cites Styleclip: Text-driven manipulation of stylegan imagery.

Visual Lexicon: Rich Image Features in Language Space Styleclip: Text-driven manipulation of stylegan imagery

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.314781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.314781Z digest=sha256:c92bcd5d6349b1aacd36d8fc9c9bdd28eb248beca9c60cfc23d8af729cf8a136

Observation 6972a36d-6ad5-4e07-979e-bd5d9eae5668 · outbound

This paper cites Sdxl: Improving latent diffusion models for high-resolution image synthesis.

Visual Lexicon: Rich Image Features in Language Space Sdxl: Improving latent diffusion models for high-resolution image synthesis

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.883818Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.329481Z digest=sha256:f7db59c98dc47ed4eec7e8d9ebe1b7e90134e2943fcdb44ee4e9a305d7a6209b

Observation 3526c95c-494c-45a0-9b7d-392445279843 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Visual Lexicon: Rich Image Features in Language Space Learn- ing transferable visual models from natural language super- vision

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.863188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.346679Z digest=sha256:2a8acae1cec60ebf7ef8015d975280139442a7e594fa4d87cc1d63874b826e6a

Observation 07ccb4c7-7fe7-4226-9664-470eb4c50e37 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer.

Visual Lexicon: Rich Image Features in Language Space Exploring the limits of transfer learning with a unified text-to-text transformer

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.839149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.352693Z digest=sha256:0b5b8d099cf9ca9eb0b5354dca14e45d27d7feba38bc9afe08927ede1188f915

Observation 97de4469-4aca-463b-ac3c-1f848dbc3502 · outbound

This paper cites Hierarchical Text-Conditional Image Generation with CLIP Latents.

Visual Lexicon: Rich Image Features in Language Space Hierarchical Text-Conditional Image Generation with CLIP Latents

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.360946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.360946Z digest=sha256:ce36130a8d6ba07af17a02cb518e2293d38bf1cb02a8a0507ad5dd69a377b720

Observation 08351e34-873d-4549-b142-3807ea059427 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

Visual Lexicon: Rich Image Features in Language Space High-resolution image syn- thesis with latent diffusion models

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.819093Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.367294Z digest=sha256:01f5e36db1b4b7ad74ad63d123861964a8d988250e95a18be6cc5b0cce44267e

Observation 1741204b-4a84-47fc-8dc5-cca29e09855f · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

Visual Lexicon: Rich Image Features in Language Space U-net: Convolutional networks for biomedical image segmentation

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.799309Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.375945Z digest=sha256:1f46f54a854ee1970e839b37bc9ecab99194e4244289c12e211db79f84280ccf

Observation e9180746-7448-4671-93cd-2e7235e8e822 · outbound

This paper cites Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation.

Visual Lexicon: Rich Image Features in Language Space Dreambooth: Fine tuning text-to-image diffusion models for subject-driven generation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.772700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.383169Z digest=sha256:bad8dc7092dde46aac333aa6c9fc267494e40dda6cc76edc69165a8250c5f6a5

Observation 726ffdd5-ad93-4285-875c-1b79fdb9ee05 · outbound

This paper cites Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models.

Visual Lexicon: Rich Image Features in Language Space Hyperdreambooth: Hypernetworks for fast personalization of text-to-image models

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.751649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.390493Z digest=sha256:1ce1069ab7a6330d7610e9ed759a62e9082358297996f94be03088d112681a12

Observation b21586a3-038e-4654-ad79-7db984e8432f · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding.

Visual Lexicon: Rich Image Features in Language Space Photorealistic text-to-image diffusion models with deep language understanding

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.730199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.398079Z digest=sha256:b35cc00a5977440540fc60448291d4f3909c87db7ab5575fe28054df48356441

Observation 9cb61f0b-e9c7-4677-80ee-4c886789ddef · outbound

This paper cites Progressive distillation for fast sampling of diffusion models.

Visual Lexicon: Rich Image Features in Language Space Progressive distillation for fast sampling of diffusion models

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.711979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.405021Z digest=sha256:248d45709d8f114190c1eaa2527c9365dee53e21b6659117b644131a46edad1f

Observation e1062836-6fdf-430a-b5fe-b1e2ceafd75c · outbound

This paper cites Neural Machine Translation of Rare Words with Subword Units.

Visual Lexicon: Rich Image Features in Language Space Neural Machine Translation of Rare Words with Subword Units

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.415902Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.415902Z digest=sha256:7d84614b6b1d2130ded546b6fbcd79108573f932c7fb21d19cc1f71d267524d0

Observation dbb4e182-7d1f-47d6-aba4-56cd33b86dd6 · outbound

This paper cites Adafactor: Adaptive learning rates with sublinear memory cost.

Visual Lexicon: Rich Image Features in Language Space Adafactor: Adaptive learning rates with sublinear memory cost

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.690894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.423417Z digest=sha256:89640017a59ead0d8a78cb0c36811246503babd1b429779892a151a43aeab603

Observation 98c570c5-8c58-4703-9090-7ceb5fd670a4 · outbound

This paper cites Textcaps: a dataset for image caption- ing with reading comprehension.

Visual Lexicon: Rich Image Features in Language Space Textcaps: a dataset for image caption- ing with reading comprehension

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.432155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.432155Z digest=sha256:ae0f2560aefab84a96f8f7f2a4d5e0a1570dc62f37a5866353291d1e5416bb09

Observation 6a0cd468-0675-418e-adb7-e63546c0b297 · outbound

This paper cites Towards vqa models that can read.

Visual Lexicon: Rich Image Features in Language Space Towards vqa models that can read

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.442633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.442633Z digest=sha256:f7623e41ffae91c1237b4bcec1bdf2db2ae170cda13c4302a69d00dc863b3016

Observation 40168739-51f3-494d-82b2-52916c5d5d37 · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

Visual Lexicon: Rich Image Features in Language Space Deep unsupervised learning using nonequilibrium thermodynamics

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.638259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.455532Z digest=sha256:4bc59414794b36bf208dbe8f3d67f71671efe5f8ec3b0a907c0a7f8765fa79b0

Observation 9d5d67f3-4ce2-4d40-8d65-5b72cb5ddac6 · outbound

This paper cites Generative modeling by esti- mating gradients of the data distribution.

Visual Lexicon: Rich Image Features in Language Space Generative modeling by esti- mating gradients of the data distribution

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.617318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.465049Z digest=sha256:63835507402a074dd73f484ac6fdd94cc36ea66dee005564d450df04a62a139a

Observation c0291128-0b6e-4d13-b777-f912c63b328c · outbound

This paper cites Rethinking the inception archi- tecture for computer vision.

Visual Lexicon: Rich Image Features in Language Space Rethinking the inception archi- tecture for computer vision

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.471866Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.471866Z digest=sha256:8ab01e5f25db009e8510fab2624887877347b96721f5b3286b506ab57da5e2fd

Observation 92366efc-2d27-4d38-a7f1-4ad0d2a9d062 · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Visual Lexicon: Rich Image Features in Language Space Gemma: Open Models Based on Gemini Research and Technology

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.481135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.481135Z digest=sha256:d556e5f56210fb7524c741c6dc16f96b316c52834211e5d2991d4be82ddc8215

Observation 72cfda0f-2a18-454e-90c0-ee95e2dbf7d5 · outbound

This paper cites Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset.

Visual Lexicon: Rich Image Features in Language Space Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.487818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.487818Z digest=sha256:c2b88526839842ab1c4b865e48ce86a881e9a3c9020e5dd194adfc1a3376febd

Observation b6ef4604-a898-4f20-9585-20a277c9e7ea · outbound

This paper cites Visual autoregressive modeling: Scalable image generation via next-scale prediction.

Visual Lexicon: Rich Image Features in Language Space Visual autoregressive modeling: Scalable image generation via next-scale prediction

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.579690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.493448Z digest=sha256:c6bb4ebca05a06087cb8eedfa01003412910354ab36dcb7dcec3740d9f2cf64e

Observation 3617e7a4-14ec-4c25-9ab0-df3055614c94 · outbound

This paper cites Con- trastive multiview coding.

Visual Lexicon: Rich Image Features in Language Space Con- trastive multiview coding

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.554902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.498748Z digest=sha256:eae8ef1bda26f61b05b66b2e8db9f48e43b9dbda2191d9029a4377eb2ec0407f

Observation 787f805e-c67a-4a2f-b217-0d5cea682b99 · outbound

This paper cites Extracting and composing robust features with denoising autoencoders.

Visual Lexicon: Rich Image Features in Language Space Extracting and composing robust features with denoising autoencoders

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.535837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.506084Z digest=sha256:4cfb91958148c2e867bd8ef187178508e35542938ce87c1edc3035680a712cb1

Observation bed01bb0-dc17-4af1-a093-924f36bfa87c · outbound

This paper cites Diffusion Feedback Helps CLIP See Better.

Visual Lexicon: Rich Image Features in Language Space Diffusion Feedback Helps CLIP See Better

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.515775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.515775Z digest=sha256:99d90aabe16161edced92df5e49e644e657e42e3fca8d8c8318b58c41ffc4cd1

Observation 7ca3aeca-eb7a-4e63-b017-670a2fa7dc70 · outbound

This paper cites Unsupervised feature learning by cross-level instance-group discrimina- tion.

Visual Lexicon: Rich Image Features in Language Space Unsupervised feature learning by cross-level instance-group discrimina- tion

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.508284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.525424Z digest=sha256:459a9bdd219cce2106d8b3b02b778a5bae9c35171fb13b1d353f2beb3b2cf75b

Observation cefa8917-1ddc-48fc-8ebb-d62875f10f97 · outbound

This paper cites De-diffusion makes text a strong cross- modal interface.

Visual Lexicon: Rich Image Features in Language Space De-diffusion makes text a strong cross- modal interface

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.477684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.535074Z digest=sha256:c8c1b8da33264168ea8eeb55d3431fbac7a5bd8ab550b435b458cac4cbfbc627

Observation d6b9ab6d-99e3-45f6-abac-85d7c6cd926e · outbound

This paper cites Unsupervised feature learning via non-parametric instance discrimination.

Visual Lexicon: Rich Image Features in Language Space Unsupervised feature learning via non-parametric instance discrimination

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.450250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.544008Z digest=sha256:70d2a346f1d73d895c385e009e8839bffce7f0338eda1119b8360697d7927e30

Observation ea83b458-1ff4-4fb3-aa0f-d4a9ebc34188 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

Visual Lexicon: Rich Image Features in Language Space Msr-vtt: A large video description dataset for bridging video and language

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.427896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.553657Z digest=sha256:23c1798910d62d0fd88ffb2a45419387aa1d788e5190448ff63842db1c642113

Observation c8af8a4d-ca0b-4740-8d1b-0dc86c4cc6d6 · outbound

This paper cites Open-vocabulary panop- tic segmentation with text-to-image diffusion models.

Visual Lexicon: Rich Image Features in Language Space Open-vocabulary panop- tic segmentation with text-to-image diffusion models

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.401306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.559898Z digest=sha256:6fcb5b4028d61c07a30e24999dfcc6bcdab43c7ed7dff99b51bca4dd4cd4c094

Observation 13b41e26-cd54-491f-a928-9d7d24d32d7a · outbound

This paper cites Diffusion model as repre- sentation learner.

Visual Lexicon: Rich Image Features in Language Space Diffusion model as repre- sentation learner

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.371408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.569386Z digest=sha256:232295e08038cf84f0c3cc2ab201bbbcbdb6deba460c0389e8f344046bea5ca3

Observation 8ef7244e-c05c-46dc-8b6c-94b0e5886f8e · outbound

This paper cites Vector-quantized image modeling with improved vqgan.

Visual Lexicon: Rich Image Features in Language Space Vector-quantized image modeling with improved vqgan

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.343354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.579225Z digest=sha256:6eb1308d4bd8a2f0aaaf5789f73cff19df2e8632905ee2d329cf6a3ccba0d25b

Observation cfc5324d-733c-4bbe-9044-d46b3037695a · outbound

This paper cites Coca: Contrastive captioners are image-text foundation models.

Visual Lexicon: Rich Image Features in Language Space Coca: Contrastive captioners are image-text foundation models

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.320682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.592159Z digest=sha256:ee1bf2efa2683e47ea18d81d890272c1fcf270eaa488dd6f008ae05cb6e68cbc

Observation 17770810-7065-4cf0-9f7f-97014c3b3370 · outbound

This paper cites Modeling context in referring expres- sions.

Visual Lexicon: Rich Image Features in Language Space Modeling context in referring expres- sions

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.298202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.598680Z digest=sha256:b407155942de4b0724d5bff3c0917ac861b3fc8a622ac1cfa498b3bbf76e9aea

Observation cf55616b-e5eb-4651-afb4-d0794ba586d7 · outbound

This paper cites Spae: Semantic pyramid autoencoder for multimodal generation with frozen llms.

Visual Lexicon: Rich Image Features in Language Space Spae: Semantic pyramid autoencoder for multimodal generation with frozen llms

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.273396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.605122Z digest=sha256:ce6c62b5ff1b03025c6c48205d4e1fc45961a3981df85f71c600cde3dcc3b2ac

Observation e45783fa-396d-4806-9a3c-bc9e234330df · outbound

This paper cites Language model beats diffusion–tokenizer is key to visual generation.

Visual Lexicon: Rich Image Features in Language Space Language model beats diffusion–tokenizer is key to visual generation

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.612012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.612012Z digest=sha256:3ea45da9ba17287c5ae6317f097a5c098a47f5a50363a1892a3628eebed4d3a7

Observation fbda5d99-2584-4952-9df4-2b260fc56b0d · outbound

This paper cites An image is worth 32 tokens for reconstruction and generation.

Visual Lexicon: Rich Image Features in Language Space An image is worth 32 tokens for reconstruction and generation

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.232376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.618321Z digest=sha256:527c90dbb2931c09e3ec0a2ce9e6662de06e457beec3012dc8ece399a16038f2

Observation 9cfe70ec-22c6-428e-8c94-9b9fd7270e0a · outbound

This paper cites Soundstream: An end- to-end neural audio codec.

Visual Lexicon: Rich Image Features in Language Space Soundstream: An end- to-end neural audio codec

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.207434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.624640Z digest=sha256:12e06f7a3b356ab112ff067933eef01d9c4a24f356163ff4f3f88098a2c79348

Observation 69467f88-3b85-4b8c-ba8d-eb8f8942aaf0 · outbound

This paper cites Sigmoid loss for language image pre-training.

Visual Lexicon: Rich Image Features in Language Space Sigmoid loss for language image pre-training

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T19:22:56.185620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T19:22:55.629661Z digest=sha256:2c1bf32245589db5d88c5cf0c66c243439334f1dc7a39ca157d488391fdf767e

Observation 2e2c4c98-8c3b-434a-857d-ddd25aef3edd · outbound

This paper cites Image and Video Tokenization with Binary Spherical Quantization.

Visual Lexicon: Rich Image Features in Language Space Image and Video Tokenization with Binary Spherical Quantization

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-11T19:22:55.635316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T19:22:55.635316Z digest=sha256:77fa9204fcf0c353a82349583caea79f2c118fa50939ceefd347db88b3675562

Pith citing papers

Observation e599697c-7bea-411d-b4f2-60dcfa7a302b · inbound

Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models cites this paper.

Vision-Language-Vision Auto-Encoder: Scalable Knowledge Distillation from Diffusion Models Visual Lexicon: Rich Image Features in Language Space

Reference 66

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:54:12.941828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T18:54:09.796839Z digest=sha256:f44784c0c24b08311bc2dafc8e85b596488c01b3c1bf176fa3a26c9f01adae18