Pith. sign in

Paper Citation Record · LEDGER

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor

As of 18 August 2026, this Paper Citation Record lists 70 of 70 outbound references and 1 inbound Pith citation observation for arXiv:2507.07106.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07106 v1

Coverage vector

measured 70 of 70 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:52:37.501612Z

measured 71 of 71 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T19:30:30.517564Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

70 of 70 outbound references displayed

  • verified exact2
  • verified fuzzy37
  • unresolved29
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e8b87fd0-240e-418d-b6cf-8ec87ed4eb41 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Eyes wide shut? exploring the visual shortcomings of multimodal llms,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.207480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:31.474283Z digest=sha256:8dd5444f094190994667c4ce9484cc3049405c78c3a364dfbcebd4e3c7ff86ae

Observation a12d77e5-20f2-4573-90e9-81ca44e1ab33 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Learning transferable visual models from natural language supervision,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.197431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:31.523702Z digest=sha256:6c483399efea1b3e1e1f02b370bed729dd1f2676682219e51d972b1175f81803

Observation b602ef68-7beb-436a-b7af-92b84f6a69a2 · outbound

This paper cites DetailCLIP: Detail-Oriented CLIP for Fine-Grained Tasks.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor DetailCLIP: Detail-Oriented CLIP for Fine-Grained Tasks

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:31.581690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:31.581690Z digest=sha256:c8dcd993b3cfdc986fd59662ad7f363e0744932babdd2705ddc0e96d1ddc944b

Observation b3734393-a69e-4769-a1e9-4e7e7c1d021c · outbound

This paper cites Is clip the main roadblock for fine-grained open-world perception?,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Is clip the main roadblock for fine-grained open-world perception?,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.187174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:31.646858Z digest=sha256:d5be037931a878df226d86836c6f542da8583fdcd149b785e2027f360e1e55d2

Observation 625bdc2b-6844-4ae1-ac8f-0e68ed309af3 · outbound

This paper cites Discffusion: Discriminative Diffusion Models as Few-shot Vision and Language Learners.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Discffusion: Discriminative Diffusion Models as Few-shot Vision and Language Learners

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:52:38.207894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:31.697090Z digest=sha256:c83b27009f41829961dab5c250e6508ffa140281027dc0b55e5d5f9325066f03

Observation bf309773-992a-4ce3-a6c3-4332c53bc974 · outbound

This paper cites BRAVE: Broadening the visual encoding of vision-language models.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor BRAVE: Broadening the visual encoding of vision-language models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:31.767820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:31.767820Z digest=sha256:ea4e10abd1463760db49089e0a6608b2686109baae424269e89286a0b8b5dc22

Observation 67cf8678-d817-4b9f-bb5c-00e6edf6170c · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:31.825326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:31.825326Z digest=sha256:bbab568955665b45de73ba29ab6ca07b1a1d3d2cfb96de92502a0d5912776b73

Observation ea4cd2ef-2034-4a1b-8b7a-f9b908f0b8ea · outbound

This paper cites From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:31.897432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:31.897432Z digest=sha256:15b5f54e760195d15e56964730a20f17dfe0d856e00066311c14b4f23e2a0a61

Observation a6fd8af5-b0b1-4324-8852-e6e4af2443fc · outbound

This paper cites Mini-gemini: Mining the potential of multi-modality vision language models,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Mini-gemini: Mining the potential of multi-modality vision language models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.176426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:31.974928Z digest=sha256:64fe98cb7831e161aec0e94749635eea344c2a8b59c2a66adbd22e06e2e6f70d

Observation e9b9b89d-4508-462d-80ed-2c3b2ab5c80e · outbound

This paper cites Prismer: A Vision-Language Model with Multi-Task Experts.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Prismer: A Vision-Language Model with Multi-Task Experts

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:32.060515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:32.060515Z digest=sha256:2430d9c0443e02da8a5bb018aae8c9a5936adcf38803ced8743dacf40498d088

Observation 7cb52f24-dbbb-4ab1-81ec-a3d5c5cbafc0 · outbound

This paper cites Vcoder: Versatile vision encoders for multimodal large language models,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Vcoder: Versatile vision encoders for multimodal large language models,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.166549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:32.175329Z digest=sha256:4d2adb1232ea8e6769e63e702f4d158d255cc09cf3864d6e967d2ed6878c1bc9

Observation b7103154-94f7-4435-bfa8-ac2c332fb38a · outbound

This paper cites Question aware vision transformer for multimodal reasoning,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Question aware vision transformer for multimodal reasoning,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.156651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:32.267593Z digest=sha256:62294358286a305a5b9ef41ce3c49d5a5723be99fd8cc927839e351e295e0545

Observation 7283de56-cd47-43d2-a110-a861513486d7 · outbound

This paper cites Api: Attention prompting on image for large vision-language models,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Api: Attention prompting on image for large vision-language models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.146498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:32.368424Z digest=sha256:10579a2dfd1446de1d9745cd30db59840c6811c2f59d1fdfcd75436d408f29bf

Observation cefb533f-5313-443f-b4f5-cca3a74cf1db · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Instructblip: Towards general-purpose vision-language models with instruction tuning,

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:32.475667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:32.475667Z digest=sha256:3084cf266a29a82855f48107a2019c62b501ef9cda2204bf93519c8e323ad80f

Observation 51279896-4373-4237-84bd-1fdf372892e6 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:32.491396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:32.491396Z digest=sha256:3aa5527474c7f05d1293053cad75a68c771d9895b741feb0f919a5d6ac86a583

Observation df85e6de-4f47-4c97-afdc-2213358707fb · outbound

This paper cites High-resolution image synthesis with latent diffusion models,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor High-resolution image synthesis with latent diffusion models,

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:32.533667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:32.533667Z digest=sha256:44b9899ebc1adde02e86ff4f795965c5461224b80b3d6e1168319656225c6b02

Observation bbcd5294-5e78-4b1c-bd97-3c5db57e9a4a · outbound

This paper cites Photorealistic text-to-image diffusion models with deep language understanding,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Photorealistic text-to-image diffusion models with deep language understanding,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.116420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:32.611237Z digest=sha256:bf23f00a7d8f07b809735f068a027108e20168bce324965fef8ba7b553b8f93b

Observation 296a3972-4983-498b-ad94-39a6c8bc9d1a · outbound

This paper cites Hierarchical text-conditional image generation with clip latents,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Hierarchical text-conditional image generation with clip latents,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:32.727639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:32.727639Z digest=sha256:831f7f2d87d15545c2eba3bc4dea183a9b8874aec72847b435fdb99d0e866986

Observation c0901327-ae10-4002-9c21-bba7c7ce2154 · outbound

This paper cites Sdxl: Improving latent diffusion models for high-resolution image synthesis,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Sdxl: Improving latent diffusion models for high-resolution image synthesis,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.097243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:32.845664Z digest=sha256:b661d818ac1120cf47fde195b20a7c43e026837c7ccbedfb3768177245c18214

Observation 55412a86-d16b-48bd-a6f9-528b92f8f1d8 · outbound

This paper cites Prompt-to-prompt image editing with cross attention control,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Prompt-to-prompt image editing with cross attention control,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.086788Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:32.958030Z digest=sha256:5cff27d2b27ff398d25c57203f23b8613194181bd19ec8b8f832166d282cd1b2

Observation b99865be-9966-49d5-ae52-cd4fc760b7dc · outbound

This paper cites What the daam: Interpreting stable diffusion using cross attention,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor What the daam: Interpreting stable diffusion using cross attention,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.077053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:33.072184Z digest=sha256:534ed8531964d57c05413aa1ece1786ff55975f6f9f5f411f86c9ab7cf063bf3

Observation 5715627b-f108-4911-a7ea-6493cff780c5 · outbound

This paper cites Towards understanding cross and self-attention in stable diffusion for text-guided image editing,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Towards understanding cross and self-attention in stable diffusion for text-guided image editing,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.067728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:33.171759Z digest=sha256:becaba2dbe5496b53baf0413d115aa79b9e41b1d2cb82a72558880efa24d041a

Observation f68a552d-440a-4c04-83f3-8c6fe427aa1e · outbound

This paper cites Plug-and-play diffusion features for text-driven image- to-image translation,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Plug-and-play diffusion features for text-driven image- to-image translation,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.058139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:33.269016Z digest=sha256:a8ef2b5287fa48150a49a13403990789372768091f3608f066bd478f1cb7e51e

Observation 3de6ec26-b7fb-41ef-9c92-b1afb1298272 · outbound

This paper cites Diffusion Model is Secretly a Training-free Open Vocabulary Semantic Segmenter.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Diffusion Model is Secretly a Training-free Open Vocabulary Semantic Segmenter

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:33.421637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:33.421637Z digest=sha256:6a4b6e070ea01669af3c1a43c18b9f064923af3ebee5b7563eb62b83aa810d2a

Observation 8e32935a-f2fa-4f14-879f-b4239b68da79 · outbound

This paper cites Repurposing diffusion-based image generators for monocular depth estimation,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Repurposing diffusion-based image generators for monocular depth estimation,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.048007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:33.537106Z digest=sha256:f6ec621bbb0fb47b74bdd6ffb6ec6f6ee3873411b31f4f12541803629557f807

Observation 4c4dce4f-505e-4580-bd26-cbd6e386ba1f · outbound

This paper cites Do text-free diffusion models learn discriminative visual representations?.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Do text-free diffusion models learn discriminative visual representations?

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:33.646987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:33.646987Z digest=sha256:f5b78ffe11d71b4fe1deb1e0392c49e735a850f2b8853b0b6e1eeff3a8cace60

Observation 44a895a7-2e78-42cf-bce6-5f52c4eb49a3 · outbound

This paper cites Deconstructing denoising diffusion models for self-supervised learning,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Deconstructing denoising diffusion models for self-supervised learning,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.038820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:33.726982Z digest=sha256:4ef6793c6ff2267965bf6865bfc2564d8ddc998ee43358a68b1e66fdfe64a9f0

Observation 98a009da-c65d-4f22-9c99-2857b6e2c9a8 · outbound

This paper cites Coca: Contrastive captioners are image-text foundation models,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Coca: Contrastive captioners are image-text foundation models,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.028681Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:33.824328Z digest=sha256:2f3e55b549decc32aaf65072c1d5d60631e22221528f7027900d95a59988ce86

Observation 790d6f1e-5c0d-4a5c-9047-606095c7cbd9 · outbound

This paper cites Multimodal few-shot learning with frozen language models,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Multimodal few-shot learning with frozen language models,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.017462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:33.897344Z digest=sha256:3b96ec8966f01348df4b8987352076928bc14f1ea7812d59d5775b413a60983d

Observation e5a0be6c-107b-4935-9338-c4a9398dd94c · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Flamingo: a visual language model for few-shot learning,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:39.007156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:33.976289Z digest=sha256:fd2650da451e7f82910007616ccabeaaadf950f60cecd447f0cbf3eef9f08f7b

Observation f20a7514-0702-4d53-ad7d-f58aaafc9dfa · outbound

This paper cites Visual instruction tuning,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Visual instruction tuning,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.996235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:34.146798Z digest=sha256:df132282eafc0e9d32b9aa2f28aff8a591eaa6de67a5b8b0a124e9ec1d7388f3

Observation d0517812-1c5b-4029-9d48-4d901626aca3 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:34.227206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:34.227206Z digest=sha256:72de06713301b83bd30e4aaab4352e70a464ec9f8073aa80832cb52da611f72c

Observation 5cc71a27-f8e9-43c7-b0f5-e29d26238011 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:34.315551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:34.315551Z digest=sha256:31e547a1e180843ae836006eedfba8c15697669e69d92fed6ed77db8090133c8

Observation 556d3ef5-1f4b-4c51-aee0-77d103f946d8 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:34.409335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:34.409335Z digest=sha256:857b36b68bf4ff3d1612094da58bc2018160d1680df046ae712e6d74c0a14042

Observation 848a5aaa-b4e9-4115-82e6-decb4d529af6 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Evaluating Object Hallucination in Large Vision-Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:34.500395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:34.500395Z digest=sha256:e6ebb44ee5449f616befd3a4dec59d309ef9b2f8fc3076376f22e606b4b343f1

Observation 54725c8f-4675-4ac8-abfb-c2e9f1e25b01 · outbound

This paper cites Multi-modal hallucination control by visual information grounding,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Multi-modal hallucination control by visual information grounding,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.986332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:34.591011Z digest=sha256:f4520a550bf7377c3dc53e3d154711b14d9ada034e5b79bb34e916158b83e903

Observation 5f489965-c74f-4fb8-aa60-306084b23fae · outbound

This paper cites Detecting and preventing hallucinations in large vision language models,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Detecting and preventing hallucinations in large vision language models,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.976233Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:34.723073Z digest=sha256:cd4a2b6ea17e82c2110afa3aff1b60781093fe598a9810d20fe679cb8a255d15

Observation 5b55157d-f698-4213-833c-1e1eefac95f1 · outbound

This paper cites A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor A Survey on Hallucination in Large Language Models: Principles, Taxonomy, Challenges, and Open Questions

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:34.830166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:34.830166Z digest=sha256:0a1aa248b32ad60688e3596a7aa9b0390088185ba9d52b05a2623d3a31d8ab75

Observation 36fe1c95-1620-4561-a3cd-5e6037917e1a · outbound

This paper cites HallE-Control: Controlling Object Hallucination in Large Multimodal Models.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor HallE-Control: Controlling Object Hallucination in Large Multimodal Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:34.896934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:34.896934Z digest=sha256:e8843efa54ae753204cea6d1fe5d46ae433ee5b4e12fb329ae57a4a502f97144

Observation f9bcfcae-e0d3-4335-bc9a-5553952852ae · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:34.983806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:34.983806Z digest=sha256:b96e5ab1f353d3250881b84f01cbd91907d3d74a6edd376d5ec16922dc18d148

Observation ab5b5abd-9b09-4f0f-bccc-629838e11b29 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:35.095890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:35.095890Z digest=sha256:aa321d80a803faf0797d15b78ea6000ae4c1df0afa09189e0dedad1ca52db793

Observation 201ad2db-d2af-4eb2-b99f-c19dc5e6cf4f · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Llava-next: Improved reasoning, ocr, and world knowledge,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:35.206253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:35.206253Z digest=sha256:989663edaf3434c5c166652f3d7d84fd1be086df602b9377952ac1b0be30ff01

Observation 26b9cebe-850f-4bff-b75f-82e4460c81a5 · outbound

This paper cites MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:35.255074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:35.255074Z digest=sha256:612426451c1247ec0f28adcfb9884c430a755089f95898b458d24bbf0df210c6

Observation 1a775c40-908f-4be4-97fc-d183c552ff0e · outbound

This paper cites Improved baselines with visual instruction tuning,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Improved baselines with visual instruction tuning,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:35.415756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:35.415756Z digest=sha256:2d249803800480b81d81cd804e0a985013b9bf6bdfccde35cd6ddc27de2eea58

Observation 4aa7ebc1-2d2f-4923-9cc2-1c696551b0bf · outbound

This paper cites 4m: Massively multimodal masked modeling,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor 4m: Massively multimodal masked modeling,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.952048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:35.541595Z digest=sha256:00555eb66955c6c3a59a923941d13f1815fbc82597c5cfbec917539fb3241d4f

Observation e08a5649-69ea-45a7-8806-6160ae111cfd · outbound

This paper cites 4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor 4M-21: An Any-to-Any Vision Model for Tens of Tasks and Modalities

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:35.658391Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:35.658391Z digest=sha256:becc6fdefcc17217b9831ee8a101ca144aa366f35d9af72700e713e10b41af74

Observation 958be85f-58a5-495e-b952-93836b10ad0a · outbound

This paper cites Llava-plus: Learning to use tools for creating multimodal agents,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Llava-plus: Learning to use tools for creating multimodal agents,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.941704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:35.742782Z digest=sha256:fbbff83f075db2a8e677bfa355e578168510154c036f13f97df52239efac425a

Observation 524a1413-4b8c-4d16-94c5-8e3364a96190 · outbound

This paper cites Visual programming: Compositional visual reasoning without training,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Visual programming: Compositional visual reasoning without training,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.931881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:35.829893Z digest=sha256:1a8cfc035466cce0b67b4788e254803035b3ec35b8ceb1ea2b5e4b719c03061f

Observation cd605569-c973-412d-a5ba-9693f9cea517 · outbound

This paper cites Spatialbot: Precise spatial understanding with vision language models,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Spatialbot: Precise spatial understanding with vision language models,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.921200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:35.917707Z digest=sha256:3b7be19da82dc52d10d503083d42728a9a315331d1358d7aa83b30435f156343

Observation adba8552-2fb7-450e-9f5f-74e68331758f · outbound

This paper cites Your diffusion model is secretly a zero- shot classifier,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Your diffusion model is secretly a zero- shot classifier,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.911539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:35.983026Z digest=sha256:dcfad336d082ea2fa39d76deb224cb16103cfea14982c7d69fd5436e6bf17121

Observation 6958a577-10b3-40c1-bfe4-a581be05c5bb · outbound

This paper cites Diffusion Models Beat GANs on Image Classification.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Diffusion Models Beat GANs on Image Classification

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:36.050247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:36.050247Z digest=sha256:471d43ceaf769238ebaf72e89560ba3931c2d1acfe5399b17a7a1bb8fc44ff4a

Observation 396ca2fd-a5fa-4ad7-9dc9-4f33babf60e7 · outbound

This paper cites Open-Vocabulary Panoptic Segmentation with Text-to-Image Diffusion Models.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Open-Vocabulary Panoptic Segmentation with Text-to-Image Diffusion Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:36.136857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:36.136857Z digest=sha256:73e3d17201bfc20a66a1785c310ac89f8d053dd04cfde5ebbbed7d6ab185cbff

Observation 57b935b2-94bf-41df-af2b-48c18dd48d23 · outbound

This paper cites Diffusion Models for Open-Vocabulary Segmentation.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Diffusion Models for Open-Vocabulary Segmentation

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:52:37.892673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:36.201294Z digest=sha256:ad4e7e35194e4c03134d0d188651f0a90058d82c9300f0f0402a09881ff12b1a

Observation 2801e47e-4379-472f-b257-71088ab13e4b · outbound

This paper cites Not all diffusion model activations have been evaluated as discriminative features,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Not all diffusion model activations have been evaluated as discriminative features,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.901068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:36.301593Z digest=sha256:41bc0861dcfd93ae17e946c72069cb8610763e826bdac6c433b5e77ac00f6a6f

Observation 7d74e8ff-50e6-4bf4-b698-08dacdd59fbb · outbound

This paper cites Dinov2: Learning robust visual features without supervision,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Dinov2: Learning robust visual features without supervision,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:36.377242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:36.377242Z digest=sha256:0f57e72c38bbdf39074826208469a4a42274234a9fdced90039efc6cbed84cc9

Observation 00a9423d-718c-4cc2-8596-7644d85d2322 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.883626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:36.434918Z digest=sha256:6794893c4d617c2431684ab07755e0c40f194af159d4430d474c88ec6358ed06

Observation 3ea49173-02b8-4e6d-9f24-224ee2b21304 · outbound

This paper cites Naturalbench: Evaluating vision-language models on natural adversarial samples,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Naturalbench: Evaluating vision-language models on natural adversarial samples,

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.873717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:36.531122Z digest=sha256:3e2a806ea5ca8a123a3a7b302c27c953be78bd904fe09491a9c3ae493fbfac22

Observation 2f3f7607-0885-427b-932b-be059f4fe581 · outbound

This paper cites Openclip,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Openclip,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.863477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:36.592231Z digest=sha256:1982b3ec66e950b533121654c916ce247d23ab1a51fa47b2faede52e1df7b8f8

Observation 6ffaac83-c65c-4172-b043-419a46750ca0 · outbound

This paper cites Sigmoid loss for language image pre-training,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Sigmoid loss for language image pre-training,

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.853467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:36.661149Z digest=sha256:ebe1b080dec16c0565f3c41faa1438afa5a5be1119cd645240785733c8dcedff

Observation 7e2d64dd-5c70-499f-b96d-4a57f5104283 · outbound

This paper cites Data filtering networks,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Data filtering networks,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.843141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:36.753898Z digest=sha256:fd4288345ac8d04dc422e893e027e3d876734e10ab8e1be0ee8837d890457cca

Observation 298bcc47-91a2-48c3-a73b-f196b9054cd7 · outbound

This paper cites Demystifying clip data,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Demystifying clip data,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.833840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:36.822538Z digest=sha256:fa87c022d4968ee9ea5f34397ec883e425bed7bfe567fddf42395e96c21a5098

Observation f2e55f68-91ff-47ae-b704-e53d070df640 · outbound

This paper cites Eva-clip: Improved training techniques for clip at scale,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Eva-clip: Improved training techniques for clip at scale,

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:36.885560Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:36.885560Z digest=sha256:a7afe21475d7fc7a9657340d9c1508af8b2f0c3ce311e67bd65c4e923c550d54

Observation e10e12ee-90e0-4599-8d23-15f3ac9f1acf · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:36.954496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:36.954496Z digest=sha256:08dc2ab5b8a4be23f7e7b8f71c2899becb4784c20c2e8018934f2c2cae611fdc

Observation 7d47a46a-4416-4d94-b899-9dc929b9eccc · outbound

This paper cites Accurate computation of the log-sum-exp and softmax functions,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Accurate computation of the log-sum-exp and softmax functions,

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.817386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:37.033529Z digest=sha256:517326498e7ae97fe51b8af25c61aa75761c985dac7714a781861ba497b21900

Observation f2f69d76-9387-498b-b3cd-f189c898635f · outbound

This paper cites Winoground: Probing vision and language models for visio-linguistic compositionality,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Winoground: Probing vision and language models for visio-linguistic compositionality,

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.806943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:37.111149Z digest=sha256:59f0f237a2c05a6afdece2ad1128dbe5408d93dbec44ee9dce9bebc7b9fde437

Observation e1eba4e4-6e52-40d7-b4ef-7a107f9ca734 · outbound

This paper cites Similarity of neural network representations revisited,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Similarity of neural network representations revisited,

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.796071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:37.189998Z digest=sha256:c8607b18a40ba36e43f9611bc0deeb7ffb1de52b0efc3d83ff5bb29b6bf77300

Observation a693805c-afd9-4fbe-825d-cd9103233a47 · outbound

This paper cites Cider: Consensus-based image description evaluation,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Cider: Consensus-based image description evaluation,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T18:52:37.250094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:52:37.250094Z digest=sha256:43ebe84644dc9b2b633456c7f0d3079d457449185509e5c3293d086fdee933ba

Observation 9c03cdc7-06dd-4d99-8b0c-27ce0bf0c6ab · outbound

This paper cites Spice: Semantic propositional image caption evaluation,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Spice: Semantic propositional image caption evaluation,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:52:38.741941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:37.333853Z digest=sha256:0ee295da5e551ab6e60e64b56c4cb7f60cda6987b34a980d53afc8409e48a709

Observation 95b8348a-10c1-416b-a666-9ee0fe2a9ff0 · outbound

This paper cites Decoupled weight decay regularization,.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Decoupled weight decay regularization,

Reference 69

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T18:52:37.725447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:37.421346Z digest=sha256:f3f4f590ce9a210233883bd0732cfc8e5f430d3cfe71b39dbb9eb4b2a3969bca

Observation fdb9678a-2a92-4434-9134-38e0b3a60d74 · outbound

This paper cites Frisbees.

Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor Frisbees

Reference 70

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T18:52:38.503754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-06T18:52:37.501612Z digest=sha256:f9970f224a242164e5d1f33ac071b873a209d13e7366960ed7fbc0873917d801

Pith citing papers

Observation 92a96c14-d83f-43c4-b8d8-62cbd1f2168b · inbound

LAP: Fast LAtent Diffusion Planner for Autonomous Driving cites this paper.

LAP: Fast LAtent Diffusion Planner for Autonomous Driving Towards Multimodal Understanding via Stable Diffusion as a Task-Aware Feature Extractor

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T19:30:30.517564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T19:30:30.517564Z digest=sha256:a46d8f81d5bba7472c0da0a3311514babe17755632a855acf833e1d5e1b24f15