Pith. sign in

Paper Citation Record · LEDGER

Diffusion Instruction Tuning

As of 9 August 2026, this Paper Citation Record lists 39 of 39 outbound references and 0 inbound Pith citation observations for arXiv:2502.06814.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.06814 v2

Coverage vector

measured 39 of 39 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T11:20:02.961462Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

39 of 39 outbound references displayed

  • verified exact2
  • verified fuzzy3
  • unresolved34
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d338fb1d-b675-43e3-9579-d667c57762b9 · outbound

This paper cites Quantifying Attention Flow in Transformers.

Diffusion Instruction Tuning Quantifying Attention Flow in Transformers

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.827545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.827545Z digest=sha256:98f7800932a4f21a7ec6d80efc63c1a88cba5424593c2bf87bbb56f22371ed41

Observation 006e4a96-c10a-43d1-a547-73b0e193cf74 · outbound

This paper cites an unresolved cited work.

Diffusion Instruction Tuning Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:20:03.487975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:20:02.954871Z digest=sha256:ea9e26ac58d0244c491fef10a63cd8d46f5fd4250cfd2e0ec2ef6a321d94217a

Observation 753a6847-e0fb-4ca6-bff8-6e4114a0a820 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Diffusion Instruction Tuning Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.839886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.839886Z digest=sha256:14ba09505eaa061959b7a3c62dc4a23a47ef398760778e4148770620ccb07c8a

Observation d3e91e9b-6ceb-4084-807e-582139deaa96 · outbound

This paper cites Locality Alignment Improves Vision-Language Models.

Diffusion Instruction Tuning Locality Alignment Improves Vision-Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.847466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.847466Z digest=sha256:b75af76dfdcc7569fddff1c9e7a61e5da3cee03664a12a5a85b254bb349cafff

Observation 97079387-3cba-4818-9ce4-1de22dc43b31 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Diffusion Instruction Tuning InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.851199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.851199Z digest=sha256:8544deacadcefb85607f2e33f0e9e1b56d73b87006b321496edaa722b38108c8

Observation fd0b7c91-341b-4bd7-b841-e62f05ec2dd8 · outbound

This paper cites DeepSeek-V3 Technical Report.

Diffusion Instruction Tuning DeepSeek-V3 Technical Report

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.854961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.854961Z digest=sha256:b26a2396f5c8899eed03c9887be04a24749d8500d6d6706acfd0126a3c64bde9

Observation bd5dadf8-0c1c-4964-91ca-f7ea139c9adb · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Diffusion Instruction Tuning An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.858643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.858643Z digest=sha256:39aa6e029e7cdd9d16d7c8ce59c201eeacaace214b454bb69116fe9dc145f3b2

Observation 86928694-8f44-4b8e-8ac7-171e879ad086 · outbound

This paper cites The Llama 3 Herd of Models.

Diffusion Instruction Tuning The Llama 3 Herd of Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.861920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.861920Z digest=sha256:7ddc11e17971cea8159de31f11ff048845187baf11f7303419414d7fc8007f4a

Observation 5da7329e-dcb0-4ae5-ab1c-9af495ef7fb5 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Diffusion Instruction Tuning MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.865354Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.865354Z digest=sha256:b0c084d0d87cbacc6e10e20940bf84cff5394b9977251aff49d099980bc3933e

Observation 2ec41dfa-e504-4b23-810b-c228404d5359 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

Diffusion Instruction Tuning LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.868705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.868705Z digest=sha256:f2f7355cf31869b1c1b15420a05b023426baa7e6d097c4c1877abcc69c2f9744

Observation d4cde6a4-3798-4820-b931-7571094cf985 · outbound

This paper cites An Image is Worth Multiple Words: Discovering Object Level Concepts using Multi-Concept Prompt Learning.

Diffusion Instruction Tuning An Image is Worth Multiple Words: Discovering Object Level Concepts using Multi-Concept Prompt Learning

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-09T11:20:03.328169Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:20:02.878657Z digest=sha256:ae6b7517f3188036ada17ff3ca4c26da4f332ef141d036ea9b56548025ed240a

Observation ce78c41b-9411-4a30-9193-69c056eb575f · outbound

This paper cites Yes I’m not able to provide a name for the person in this picture.

Diffusion Instruction Tuning Yes I’m not able to provide a name for the person in this picture

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:20:03.478271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:20:02.958257Z digest=sha256:846697866f01b198ee0fac5699f9acea153b0003bf9b33f5b3963b3e0fe6609c

Observation 29ba56bc-ad06-4e2f-9a4f-2abd4030c8c3 · outbound

This paper cites A diagram is worth a dozen images.

Diffusion Instruction Tuning A diagram is worth a dozen images

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:20:03.513418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:20:02.885151Z digest=sha256:594468cc07666df0d9e2849eccea7592ea3d00a33a1c426674638302160ab1fd

Observation 8cef1745-d6ec-4c29-8da3-e693cbf9c072 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Diffusion Instruction Tuning SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.888308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.888308Z digest=sha256:2e7ed0f5a7e049001ebe64a2edbd898d2de554ce0dd39278a3c76440bb8674aa

Observation 8ea88340-2eda-4ee4-b965-2c0c86a1074c · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

Diffusion Instruction Tuning HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.895321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.895321Z digest=sha256:ab38dd73791f69a67864efb205d36d29132719ef6478434a81ef82a2b080f849

Observation 05b1e9d4-cfdb-4221-b7d0-83d0c64e4274 · outbound

This paper cites WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation.

Diffusion Instruction Tuning WorldMedQA-V: a multilingual, multimodal medical examination dataset for multimodal language models evaluation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.902422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.902422Z digest=sha256:bda59cee48f1b8815f0a09ba7d8c332eb1fcc27b5e2cde635fda140605f7fbb7

Observation 36a53c59-dd2e-4c85-a3f2-2498162596d4 · outbound

This paper cites K., and Chakraborty, A.

Diffusion Instruction Tuning K., and Chakraborty, A

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.905887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.905887Z digest=sha256:74ee882e7efd4b2f1defd715885c2ab81f178ac8a3da8519ef7864b344f2ae84

Observation dedd01ec-d94b-4eb1-82aa-2a9b642f4438 · outbound

This paper cites Null-text Inversion for Editing Real Images using Guided Diffusion Models.

Diffusion Instruction Tuning Null-text Inversion for Editing Real Images using Guided Diffusion Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.909198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.909198Z digest=sha256:4d13105e10387c6204d967e354df6d9a5b40adafab27eea455f9d68d62b68aa7

Observation 533938b4-b2bb-4b3a-bd32-4cf798a355d5 · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

Diffusion Instruction Tuning Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.912656Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.912656Z digest=sha256:6d8b724068662d064cc0300355fae674e436c7994362c7323a39d82b131b5304

Observation ee30dcd7-c7ae-46b0-88c6-2244f47e3098 · outbound

This paper cites LMFusion: Adapting Pretrained Language Models for Multimodal Generation.

Diffusion Instruction Tuning LMFusion: Adapting Pretrained Language Models for Multimodal Generation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.915960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.915960Z digest=sha256:026da7f09be4a08274d82bdffb202f49c6d2f2160d8edb599a4f6e0dc35ff08c

Observation 98d4ebac-f61e-41de-8b99-dea333610fa5 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Diffusion Instruction Tuning Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.920209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.920209Z digest=sha256:cc0faa46d92f61c49e18e2e4797e6b38e012e62a24602980719dbfb330fb7bfc

Observation 4eb3cb22-1e42-4136-b9b9-fa8ed3b8d626 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Diffusion Instruction Tuning Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.923770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.923770Z digest=sha256:b0ab1d563e739b838fc765b15bccccf26c55b0569aef70afecea2db310654025

Observation 831adb08-f292-4032-9ce2-52ce6f47f6c8 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Diffusion Instruction Tuning MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.930797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.930797Z digest=sha256:46bde54e1ba0598bff178e7d5bb678d9a40e7f2eb84b43f318dd1dd7515e7164

Observation 93323bd5-84d2-4465-a886-e599e250c5b5 · outbound

This paper cites Ferret: Refer and Ground Anything Anywhere at Any Granularity.

Diffusion Instruction Tuning Ferret: Refer and Ground Anything Anywhere at Any Granularity

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.934093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.934093Z digest=sha256:24dea05dddb3b866d4195f5533fdbc46bf9a4ffdc033339788a122305b64af8f

Observation fdcb8bb8-c620-4063-b910-8d2c30a36e4f · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Diffusion Instruction Tuning Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.940685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.940685Z digest=sha256:8ed3babbeb4a762c7523cce136cf5fd943e22f3c090e48e502b31ef219e10159

Observation 9b2dabd3-6c42-4bed-840f-89c5436cf7a0 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Diffusion Instruction Tuning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.944165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.944165Z digest=sha256:7a310daf23a6befad296bd1718ccd1b6d50c1ee17440f443ec49c9e98b53d4e6

Observation 282b4b10-ff46-44ee-934c-395677431439 · outbound

This paper cites MoVA: Adapting Mixture of Vision Experts to Multimodal Context.

Diffusion Instruction Tuning MoVA: Adapting Mixture of Vision Experts to Multimodal Context

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.947795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.947795Z digest=sha256:b04be65060ab457db257229cd95e398fef72486b3c2240fbf2d2251d60b0200f

Observation 083baf5f-7022-403b-8cd9-cfb287ed4f02 · outbound

This paper cites Lavender-Llama3.2-11B occasionally refuses to answer questions for privacy reasons, resulting in a FALSE score and reduced performance on MME as shown in Figure.

Diffusion Instruction Tuning Lavender-Llama3.2-11B occasionally refuses to answer questions for privacy reasons, resulting in a FALSE score and reduced performance on MME as shown in Figure

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T11:20:03.468256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:20:02.961462Z digest=sha256:9b6bc51d316d74e67413bc3e74e2e87dd8ef73c4ead42c8a899777c7294f775f

Observation 94208058-f678-4688-94a6-7e46de1a8519 · outbound

This paper cites Rlaif-v: Aligning mllms through open-source ai feed- back for super gpt-4v trustworthiness.

Diffusion Instruction Tuning Rlaif-v: Aligning mllms through open-source ai feed- back for super gpt-4v trustworthiness

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.937592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.937592Z digest=sha256:f208ae13b728f16be37d50680c148107aca6ba32729c59dddf677e8ddc9b864b

Observation ae997530-f339-4175-a6d3-29970b2e5456 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

Diffusion Instruction Tuning Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.843537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.843537Z digest=sha256:77faf177342d8a72f8f71e54991dcf8a787d49b8119e958c90127f01a2fd6b30

Observation 1dcd946e-6b12-49be-819c-8c930699e662 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Diffusion Instruction Tuning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.927306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.927306Z digest=sha256:fc639d49928a0cd80768369f37742d9da453b385aaada044bbab72b420653c8f

Observation 1e2c2531-220b-4ead-ad40-fabf444c2d55 · outbound

This paper cites Generative Visual Instruction Tuning.

Diffusion Instruction Tuning Generative Visual Instruction Tuning

Reference 2018

Resolution
verified exact
local_arxiv, observed 2026-08-09T11:20:03.351075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:20:02.871867Z digest=sha256:9247041e918f4bc3b6105a10473ea2174ecf507418190c73fc64d7da8bccd87f

Observation 9f27da9b-4f64-418e-8d25-7da08b19e408 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

Diffusion Instruction Tuning ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.898927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.898927Z digest=sha256:9d0f3eba69bf33cfc23d3202a26138194cb1d11393dd86460f8699a783ad7e3a

Observation f44f9665-ad84-49ad-a442-55c9e4ea713b · outbound

This paper cites an unresolved cited work.

Diffusion Instruction Tuning Unresolved cited work

Reference 2020

Resolution
unresolved
raw_fallback, observed 2026-08-09T11:20:03.497726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-09T11:20:02.951042Z digest=sha256:a91ee4d09260e8dec85172a0cc4af09ed3e43f34b53a43606a8f57b9fb0b1d92

Observation c7b5b60e-23e2-47cf-953a-30ef807b3ed0 · outbound

This paper cites From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models.

Diffusion Instruction Tuning From CLIP to DINO: Visual Encoders Shout in Multi-modal Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.875384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.875384Z digest=sha256:3279fe44b27188d846c3a30755d5f6157d11e8e8dc591404e41161502479e920

Observation 887f7ad3-bf50-4b40-a7bc-c69fb48be241 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Diffusion Instruction Tuning Evaluating Object Hallucination in Large Vision-Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.891980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.891980Z digest=sha256:21cfcfa4f96eef8924d8b41fb0a70a6fab3567e1f9b8f3c19f53981fc3c825ad

Observation a4ca4871-c7c3-4e9d-bf23-56ed5f6ce608 · outbound

This paper cites Qwen Technical Report.

Diffusion Instruction Tuning Qwen Technical Report

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.836120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.836120Z digest=sha256:140ed44822f85998464f140835a48ec608a604a849f329f26028a7a67e699aff

Observation 02df6427-1317-4c79-bb21-8109ec65a7bd · outbound

This paper cites Pixtral 12B.

Diffusion Instruction Tuning Pixtral 12B

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.832217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.832217Z digest=sha256:ad6a094335073a263cd178edc4afef03647e6423faa6c2d1670d1465716591fd

Observation dae22de9-8599-466a-b913-24ff7e4a6a9d · outbound

This paper cites Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models.

Diffusion Instruction Tuning Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-09T11:20:02.881929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T11:20:02.881929Z digest=sha256:c983c755f4272e156ad20899438a5e17f136f82fe52554ef023c2b6a08c06985

Pith citing papers

No inbound Pith citation observations are available.