Pith. sign in

Paper Citation Record · LEDGER

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA

As of 19 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2505.06356.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.06356 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:47:49.169054Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 663b3fed-252e-4bcd-9f2b-3b31b2a02610 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Flamingo: a Visual Language Model for Few-Shot Learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.579507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:47:49.049971Z digest=sha256:6fe88e0d26136b85b616ab03e56e1c26809556eae1e9e07f24e880b098614627

Observation cbaf101d-231b-4de9-9e3a-656974b9fcbb · outbound

This paper cites Y our vision-language model itself is a strong filter: Toward s high-quality instruction tuning with data selection, 2024.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Y our vision-language model itself is a strong filter: Toward s high-quality instruction tuning with data selection, 2024

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.566146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:47:49.054656Z digest=sha256:e9e2cfef85c327c3b3ecd705bd3c1f96295799247289ade4b950d198f8ce21ad

Observation c8739de5-0843-4a1d-9212-26007803a608 · outbound

This paper cites Comm: A coherent inter- leaved image-text dataset for multimodal understanding an d generation, 2024.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Comm: A coherent inter- leaved image-text dataset for multimodal understanding an d generation, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.553590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:47:49.058646Z digest=sha256:6238679940ea86a570a0e6372495adda117d836effceb9b715e139e429c31e2c

Observation 571a04cb-9f43-4929-ab0f-3c26ca496045 · outbound

This paper cites PaLI: A Jointly-Scaled Multilingual Language-Image Model.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA PaLI: A Jointly-Scaled Multilingual Language-Image Model

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T22:47:49.063146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:47:49.063146Z digest=sha256:f898e373aba08125655f3b5543fade98e1442e04817939e91580a081542303b9

Observation beb628dc-a79c-441d-b713-aeae39b9511f · outbound

This paper cites PaLI-X: On Scaling up a Multilingual Vision and Language Model.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA PaLI-X: On Scaling up a Multilingual Vision and Language Model

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T22:47:49.067724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:47:49.067724Z digest=sha256:44eeb716da29ab3d5ee743687d631b0b55e029f03c79ac3e5b228967ce04a6cc

Observation 4b1fa9a7-9aa9-4813-9bd3-f003bb63ac61 · outbound

This paper cites Command R.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Command R

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.540544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:47:49.072264Z digest=sha256:5867c8ba953f26e5f0b866869297893406a6d0f7911f8779e26445ae74323a6b

Observation 4b29b4f1-c834-4464-a24f-ca449b161f2c · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T22:47:49.076972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:47:49.076972Z digest=sha256:c6da01f67583fe7a3cd91030d7fef05c18dc3d0c4898cb076bd9e86a6037f1c8

Observation 11a0ef41-deff-45d3-befd-cdea1cce37d2 · outbound

This paper cites Detoxify.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Detoxify

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.527155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:47:49.081229Z digest=sha256:6569d60f2d5bf0e2c7ad6d3b63b22725ce8234973bc6972f43b83860db953485

Observation b1b27d42-1677-40a9-80d5-d26b7f492402 · outbound

This paper cites LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA LlavaGuard: An Open VLM-based Framework for Safeguarding Vision Datasets and Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T22:47:49.084993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:47:49.084993Z digest=sha256:3d4262c16a90597cab9d9461abb4f39a73ca2e6f1dbdfb51c93a995eae9dd811

Observation 28400e92-8541-4288-8602-28fc2a78a483 · outbound

This paper cites Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language an d vision-language models, 2024.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Jailbreakzoo: Survey, landscapes, and horizons in jailbreaking large language an d vision-language models, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.513777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:47:49.089163Z digest=sha256:c323981485f2a7ddd3dec67211b218bc7b21e4acb289688db57837004480637b

Observation 72617fd2-c0ba-45ae-8ca4-8d833f511516 · outbound

This paper cites Vhelm: A holistic evaluation of vision language mod- els, 2024.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Vhelm: A holistic evaluation of vision language mod- els, 2024

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.500892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:47:49.092974Z digest=sha256:283ed1dbea147cb1fd7b51cf200497ae0a2a2fb60f0e68d7a843d0b2c3204db6

Observation 0016f5da-6da0-4bd6-832d-0f41ca6cdf4f · outbound

This paper cites Elite: Enhanced language-image toxicity evaluation for safety, 2025.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Elite: Enhanced language-image toxicity evaluation for safety, 2025

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.487453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:47:49.096768Z digest=sha256:2a952358915143dd7b246c6a6ad9427d0ec1a08412ec3210ae85df3fe0a3cd97

Observation 7c52c311-f8aa-44e2-982b-3c03ae64f2f2 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning, 2023.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Improved Baselines with Visual Instruction Tuning, 2023

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.472673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:47:49.100663Z digest=sha256:6f5d2c28af45632796181387858a1b2bc0451c7ce05d3235cf3d10fe86a623a2

Observation 0f0d46d8-289f-45b3-80c1-1269dd77d017 · outbound

This paper cites Visual Instruction Tuning, 2023.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Visual Instruction Tuning, 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.458294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:47:49.104582Z digest=sha256:e20c0c891383d6b9243d2c326fe013a64802b1432c8b0acfd71ed376961be277

Observation 13475d53-5d5d-486c-b399-112d1f2cb2f0 · outbound

This paper cites Mm-safetybench: A benchmark for safety eval- uation of multimodal large language models, 2024.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Mm-safetybench: A benchmark for safety eval- uation of multimodal large language models, 2024

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.445746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:47:49.108332Z digest=sha256:9529254b037e2b9a7cc378908973663711334eea533c2fcb81014a7918d0ebdb

Observation 2a99a062-e29d-44f4-b738-b0a4e56e1c93 · outbound

This paper cites Towards interpreting visual infor - mation processing in vision-language models, 2024.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Towards interpreting visual infor - mation processing in vision-language models, 2024

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.431701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:47:49.112379Z digest=sha256:b728621fe5f657eded30bb86a32c39f5a3b48bd03319151a0131b8f2ce0ce372

Observation 92697fc4-b12a-4c82-99dd-ea89b95b9e8a · outbound

This paper cites Kosmos-G: Generating Images in Context with Multimodal Large Language Models.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Kosmos-G: Generating Images in Context with Multimodal Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T22:47:49.116160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:47:49.116160Z digest=sha256:507e52d387708ae8a9a52b7a36cc98fc8cbaa47589025bb27113ca46f61edc0d

Observation 1647ad66-906c-4a99-940b-0121ea53c54e · outbound

This paper cites Kosmos-2: Grounding Multimodal Large Language Models to the World.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Kosmos-2: Grounding Multimodal Large Language Models to the World

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T22:47:49.120423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:47:49.120423Z digest=sha256:eebb5c8b2d6f45148d4e09edeb36e324cdf4f2f7f2de26e075419ccf2d0a794c

Observation 088b5361-2522-4ad9-b26a-8b90d41c74a5 · outbound

This paper cites Learn- ing Transferable Visual Models From Natural Language Su- pervision.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Learn- ing Transferable Visual Models From Natural Language Su- pervision

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.418176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:47:49.124792Z digest=sha256:20dad08bff82307687457b09533272e8576ce5f31ad5e20e58d40230678df15a

Observation 6543f544-6704-401b-9c6b-d4d5116611c4 · outbound

This paper cites Training-free mitigation of language reasoning degradation after multimodal instruction tuning, 2024.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Training-free mitigation of language reasoning degradation after multimodal instruction tuning, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.404586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:47:49.130340Z digest=sha256:d0eb146f0aaf0812adcedfb83759f7e25e3970174e99debceb2c75992b553c46

Observation 852ddd3e-c245-4c8d-80de-3d3a7fbc7846 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models, 2022.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Laion-5b: An open large-scale dataset for training next generation image-text models, 2022

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.391510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:47:49.134378Z digest=sha256:44daa23483b3205e52eaf25c35c0af2a536c4e5ea7b4707780a47127454aea56

Observation 66fd9b2b-a4bd-4769-a8c1-b8851402f569 · outbound

This paper cites From pixels to prose: A large dataset of dense image cap- tions, 2024.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA From pixels to prose: A large dataset of dense image cap- tions, 2024

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.378068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:47:49.138090Z digest=sha256:346a54db896f640460c7310474991972db6b01168d16fc960336f6f09cd97e0d

Observation 0e45e7a4-cd81-4d42-b2c9-df0cac1f1b5f · outbound

This paper cites RoFormer: Enhanced Transformer with Rotary Position Embedding, 2021.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA RoFormer: Enhanced Transformer with Rotary Position Embedding, 2021

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T22:47:49.141758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:47:49.141758Z digest=sha256:e3ba989a63a54a4b9fe67c34769898af82715f036a1d120996201dbaecf0adda

Observation e120afa9-4647-45ad-afdd-aae5e0c7bdef · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T22:47:49.145359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:47:49.145359Z digest=sha256:2996f2e8faf78e17739feae10838c0ae86069b13a4b90f4cd3673fc9c8fb05b6

Observation 1aeacf78-2ebc-4b92-840c-ed201d70f13a · outbound

This paper cites Florence-2: Advancing a unified representation for a variet y of vision tasks.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Florence-2: Advancing a unified representation for a variet y of vision tasks

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T22:47:49.148914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:47:49.148914Z digest=sha256:21e8fae7fa324bd61c3d9342a810f41b7f47f983f5715a5cee262a9ae57b2635

Observation 3da6cc34-7eb3-48ab-9aca-e841e01b8453 · outbound

This paper cites Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Pangea: A Fully Open Multilingual Multimodal LLM for 39 Languages

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T22:47:49.152993Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T22:47:49.152993Z digest=sha256:bcbe12f1241c7bcee64828a9c8ad685d7507cef49e6bd8dc8fe42569f29f8f43

Observation 4f8a7d9c-5f57-4835-9f93-a60c693cfe9e · outbound

This paper cites Sigmoid Loss for Language Image Pre- Training.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Sigmoid Loss for Language Image Pre- Training

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.348897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:47:49.157150Z digest=sha256:4982ea272c075444cd4a75dfc9dc7f2567457fe9f050a171a527554902d7a92f

Observation 23d66c40-6f80-4e29-a430-85d07664c789 · outbound

This paper cites Spa-vl: A comprehensive safety preference alignment dataset for vi - sion language model, 2025.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Spa-vl: A comprehensive safety preference alignment dataset for vi - sion language model, 2025

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.336125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:47:49.161173Z digest=sha256:fcb241bf3a23843135bc6abcf645a2b3cda3f198da23273f052e1f944fb1afc9

Observation aeeabb02-1824-46e6-aa80-75e3ffa4eb70 · outbound

This paper cites Zero-shot defense against toxic images via inherent multimodal alignment in lvlms, 2025.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Zero-shot defense against toxic images via inherent multimodal alignment in lvlms, 2025

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.322869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:47:49.165134Z digest=sha256:2a05d8985c2f29b73f0ae7482f5cf623aaddce39bf62a352dcfa4e369c9396bf

Observation 53243062-5057-4e78-81f9-1e3cb69f06cd · outbound

This paper cites Un- derstanding and rectifying safety perception distortion i n vlms, 2025.

Understanding and Mitigating Toxicity in Image-Text Pretraining Datasets: A Case Study on LLaVA Un- derstanding and rectifying safety perception distortion i n vlms, 2025

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T22:47:49.309450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-15T22:47:49.169054Z digest=sha256:a3865a8d73008f1f6ae1f27f45dfb3e6466254e5c6417ee8c434f59420e79bbb

Pith citing papers

No inbound Pith citation observations are available.