Pith. sign in

Paper Citation Record · LEDGER

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization

As of 17 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 2 inbound Pith citation observations for arXiv:2504.15619.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.15619 v1

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:27:36.967307Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T09:32:15.287630Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T06:54:20.297814Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved24
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 093899e6-7095-4e10-b223-4ddb7bcf3d23 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.775609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.775609Z digest=sha256:25ca8bd89bb00020516803ddb5769bc1b69a1405269b427a0c57d9e7350afa9c

Observation 92150601-f9d1-4c86-bdd9-6a97ca3e59e7 · outbound

This paper cites Rank analysis of incomplete block designs: I.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Rank analysis of incomplete block designs: I

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.781271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.781271Z digest=sha256:a445c22984eadf4f3dccc9c376e5256d77effe1b4f86afde1f1b8d33b18ef1f7

Observation 719192e5-89df-4588-bac2-e43d57f8e384 · outbound

This paper cites End-to- end object detection with transformers.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization End-to- end object detection with transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.786299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.786299Z digest=sha256:a4a16aba228aed6cbb3727b68dad0281ddfa6c0c87a8d504a1de559696b44282

Observation d81c4bbe-ef1f-4e95-a20b-9f4be29fe3b7 · outbound

This paper cites On Softmax Direct Preference Optimization for Recommendation.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization On Softmax Direct Preference Optimization for Recommendation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.791252Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.791252Z digest=sha256:a232a3c757c384c92010778b7e44792be5d9d1d2c673b1018e3d68d4fe207162

Observation bb9c2a54-d787-4175-94d9-5ca5817a5c73 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.795638Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.795638Z digest=sha256:ab0d310b88d38c86687037acd7db957e6bb6a6fa7bb2d411285f0b8c72cc0839

Observation d04cb186-8fa6-46a8-80e5-ebef887297ef · outbound

This paper cites Fine-Grained Verifiers: Preference Modeling as Next-token Prediction in Vision-Language Alignment.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Fine-Grained Verifiers: Preference Modeling as Next-token Prediction in Vision-Language Alignment

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.799901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.799901Z digest=sha256:7c5666ef0da3e7751f85118d8bfd2d7743ed63e2fea2d4430330681f545365a8

Observation 3a17869f-2166-42cf-8746-51fd6f8a1e80 · outbound

This paper cites The Llama 3 Herd of Models.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization The Llama 3 Herd of Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.804737Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.804737Z digest=sha256:49dc72013696b55be25b1ca96fbdae53bb13b2012a8365b522c797ccc5d44476

Observation 95945a37-ba21-4a3d-baa5-fdab24ce670e · outbound

This paper cites Token pref- erence optimization with self-calibrated visual-anchored rewards for hallucination mitigation.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Token pref- erence optimization with self-calibrated visual-anchored rewards for hallucination mitigation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.809824Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.809824Z digest=sha256:35c1b36c7a077ac186e0ac83d902e814bbd6e6ded9e79fe40d22f9133a21705a

Observation a6d22f44-da87-4636-b64c-3c08af0b715a · outbound

This paper cites Opera: Alleviating hallucination in multi- modal large language models via over-trust penalty and retrospection-allocation.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Opera: Alleviating hallucination in multi- modal large language models via over-trust penalty and retrospection-allocation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.650484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:27:36.814777Z digest=sha256:16df2811d8493e21fbc2ed6d0ced024ed183824f9951f4ed04c92acb3097da17

Observation d8808e51-ce19-4561-ba05-82f0bd35a36f · outbound

This paper cites Vcoder: Ver- satile vision encoders for multimodal large language models.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Vcoder: Ver- satile vision encoders for multimodal large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.636582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:27:36.819185Z digest=sha256:b73d507104eccf46d26cdaaace4468eb8c48a2aa86ca4019eab2be00c69b9cd3

Observation eed9db34-3573-4c79-9481-0aebebfa17fd · outbound

This paper cites Segment any- thing.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Segment any- thing

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.823599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.823599Z digest=sha256:2ef52440a5393c95d35b56510e4c1b7a0db26962abecd1e43b568962ec15fcad

Observation 7a39659f-7ac2-4425-9003-a4517094f08f · outbound

This paper cites Mitigating object hal- lucinations in large vision-language models through visual contrastive decoding.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Mitigating object hal- lucinations in large vision-language models through visual contrastive decoding

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.613634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:27:36.828357Z digest=sha256:8ec7b8218046a0d758c876ae5e41090d415638f29dad8d541db2f7565f1dc210

Observation 5a629912-dca5-4e7f-9a44-ec6495c14280 · outbound

This paper cites Mitigating hallucination in large multi-modal models via robust instruction tuning.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Mitigating hallucination in large multi-modal models via robust instruction tuning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.600403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:27:36.832906Z digest=sha256:56452b34f90b6d3cf3b14c90fdba7aa595d46d57084a875c68c3d2acb552b79e

Observation 4934d308-8fd1-44a2-b4de-e18754222cd3 · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.586286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:27:36.837390Z digest=sha256:93930e580ae129910ccc7b8b0fd87e6de695d57e8ff2e2ddf9d1d0a220a2f734

Observation 82e1dca7-42f8-4e1b-a925-a5a10ba7aaf6 · outbound

This paper cites Visual instruction tuning.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Visual instruction tuning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.842748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.842748Z digest=sha256:efcc5a8f766cf10ebd5088964e70837f44a6dc6192ab11c2d5188d1091035dac

Observation 424ef249-6884-4c4d-884f-fb2f98b8c06c · outbound

This paper cites A Survey on Hallucination in Large Vision-Language Models.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization A Survey on Hallucination in Large Vision-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.846971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.846971Z digest=sha256:053079659c78c905971a2ab6e01c1fbc00543cd368fa545ef20ae7e787dfbf63

Observation 130569f9-1b16-4e9e-8e42-4c1bcff895e6 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.851574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.851574Z digest=sha256:c426eeaa581dee2e63d12145c1fa82d3909c3ff66edceb3c5e41740ffbd2fb93

Observation 5f4a199b-33f3-482a-8863-026fdecfdc87 · outbound

This paper cites DAMA: Data- and Model-aware Alignment of Multi-modal LLMs.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization DAMA: Data- and Model-aware Alignment of Multi-modal LLMs

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.855648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.855648Z digest=sha256:f430d7e1d90bc0c8e9407a0149c6a78bf4b6d2867513df9ddd5c6c494a67dc1a

Observation 125a8a78-1c1b-4aac-af6b-4f56387f5b13 · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.859960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.859960Z digest=sha256:b50e389501bc4fef68577ababf4513c8cbddda2f66e854d680e4210fb745fc1f

Observation 23502d6e-736e-4c68-87de-ac85f7769eb8 · outbound

This paper cites Training language models to follow instructions with human feedback.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Training language models to follow instructions with human feedback

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.552804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:27:36.865016Z digest=sha256:75475828feea52b31760cd3b4bb2f4062466d9891c530c035381bdca4142c98a

Observation adad47df-3b0a-4954-9301-9987268c5d02 · outbound

This paper cites The analysis of permutations.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization The analysis of permutations

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.538273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:27:36.868685Z digest=sha256:8dde0b0e9de7495b49b7333d98f5795eaf2c6a2d09d50515ed14495538964e78

Observation dd965f93-ac32-4fd8-8348-4841393b0271 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Learning transferable visual models from natural language supervi- sion

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.872489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.872489Z digest=sha256:5d9b98931f6c75a858ea50fceac44c0eec8d478aa31a4d93a85bab679eb32f92

Observation e26f40fc-0a66-49a1-8279-52db45feb442 · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Direct preference optimization: Your language model is secretly a reward model

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.515321Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:27:36.876456Z digest=sha256:329aae515f301ab7667638bb2a27e192c4fd89f586df59d6be05370572fb5214

Observation e19d75d0-46dc-46b5-b2e6-f4718afbb906 · outbound

This paper cites Object Hallucination in Image Captioning.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Object Hallucination in Image Captioning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.881296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.881296Z digest=sha256:7243495a622bf448763727da64ccb3a52f7f10a2f54ae23b26d185ff279a1a2b

Observation 3f03910b-e67d-4962-94b5-a4567d505901 · outbound

This paper cites Proximal Policy Optimization Algorithms.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Proximal Policy Optimization Algorithms

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.885833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.885833Z digest=sha256:f222e9d84b61182065831521935d8eccb082efe9f8c8101caa647fb1416048c4

Observation 37f7f5e2-deb2-4bf3-a4a3-a7ad16d48696 · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Objects365: A large-scale, high-quality dataset for object detection

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.501824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:27:36.890601Z digest=sha256:60484a30433d7352b84296ea97e5d6f7433d7e970cc4038489010d8b85023111

Observation 1214fcfc-d56e-4ed2-9e45-9a34c45751bf · outbound

This paper cites Aligning large multi- modal models with factually augmented rlhf.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Aligning large multi- modal models with factually augmented rlhf

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.487560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:27:36.895020Z digest=sha256:eac7532330f604391f0f27330d90fe0dd040c153031b84b60f65cdf8f1aeb8b2

Observation ee3e42fc-198b-4fbc-baa8-8b677aebf1ee · outbound

This paper cites Resolution-robust large mask inpainting with fourier convolutions.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Resolution-robust large mask inpainting with fourier convolutions

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.473590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:27:36.899747Z digest=sha256:24557ba1c8f552c003c14fd0b08ee2a136e6a274262782e1ea45f9535f4db029

Observation c1c23691-359e-4934-86a2-7da63f6ed46c · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.459135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:27:36.904069Z digest=sha256:6b74922a9cb9c42805f2ae6068a31e35f4ca9fb8a53fa31f941b9bad9f7a9457

Observation 2d99a1ee-41d2-498b-9578-1562c7ca1bd3 · outbound

This paper cites mDPO: Conditional Preference Optimization for Multimodal Large Language Models.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization mDPO: Conditional Preference Optimization for Multimodal Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.908291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.908291Z digest=sha256:6e1d08dcf9662b797eb787cd9195e77231d5ef9df5a39af0a5e3d9e9a367a4be

Observation 7de03321-57d7-460a-8924-106a1a52e5c4 · outbound

This paper cites AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization AMBER: An LLM-free Multi-dimensional Benchmark for MLLMs Hallucination Evaluation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.912762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.912762Z digest=sha256:09391be50c9d87e87fc45afcc41fda395a606711ab2a40c25093fa73fd5f1d46

Observation 5f2e13fd-6b91-40ca-b3e9-fb4316cd8b43 · outbound

This paper cites V-dpo: Mitigating hallucination in large vision language models via vision-guided direct preference optimization.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization V-dpo: Mitigating hallucination in large vision language models via vision-guided direct preference optimization

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.445952Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:27:36.917543Z digest=sha256:ccf1b4307dc2403108bcbd53be93897fbb471b224cd4529a2afd9779afb66fd1

Observation b1706a88-95c2-4dc8-91d8-c7a31db28773 · outbound

This paper cites Miti- gating object hallucination via concentric causal attention.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Miti- gating object hallucination via concentric causal attention

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.433369Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:27:36.921589Z digest=sha256:878bc96ac7dfe8d7251223a5beed188004d40ad6e8396cc1f8097252e7b1a4f3

Observation 543e52e2-b383-42e1-b31c-521ecae710f5 · outbound

This paper cites Hallucidoctor: Mitigating hallucinatory toxicity in visual instruction data.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Hallucidoctor: Mitigating hallucinatory toxicity in visual instruction data

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.419164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:27:36.925145Z digest=sha256:7eb31025ce2fd36e86806b92652e74980140a99db1e132b2c5b54670e8d25e3f

Observation 1b7b53f2-0ecd-44f9-8737-94a9d76baf2e · outbound

This paper cites Gradient surgery for multi-task learning.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Gradient surgery for multi-task learning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.403148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:27:36.928913Z digest=sha256:434d9e5bb11a87158db079fb0a1129ffa17e2541b9dab1745c99ce8288f4dba5

Observation 6ac5f39a-6808-4e98-a437-78bec96e61a2 · outbound

This paper cites Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional hu- man feedback.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional hu- man feedback

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.389025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:27:36.932671Z digest=sha256:b84aa7b3f5642653a1805169b52cd7e82692642cc68ccf7e6397e833aba4bb60

Observation a5b3b0f2-d9a5-4cd6-b1e4-0139ba3583ec · outbound

This paper cites Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Rlaif-v: Aligning mllms through open-source ai feedback for super gpt-4v trustworthiness

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.936265Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.936265Z digest=sha256:521c9b7a19d8260b6104310824ce786882b6108c1675cffbd4c3c3ebd30698d0

Observation 73ad28ea-13da-4df4-b350-6544a29966c6 · outbound

This paper cites Less is more: Mitigat- ing multimodal hallucination from an eos decision perspec- tive.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Less is more: Mitigat- ing multimodal hallucination from an eos decision perspec- tive

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.373451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:27:36.940132Z digest=sha256:7be24ede89ec05442b1d9119df6a20babd634a78dfd44692048e7a085f8f65e7

Observation 45faf425-0cfa-479f-8119-e2b3bd570cd4 · outbound

This paper cites Dino: Detr with improved denoising anchor boxes for end-to-end object de- tection.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Dino: Detr with improved denoising anchor boxes for end-to-end object de- tection

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.359000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:27:36.943904Z digest=sha256:1bd6b659cd8af86c09a6cfd5c35792b83b572e00c97305c250e206614bfe2d73

Observation 01843850-9619-4c58-b617-ebe5e29f0e00 · outbound

This paper cites Automated multi-level preference for mllms.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Automated multi-level preference for mllms

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:27:37.344886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-16T11:27:36.948595Z digest=sha256:5f5fa6fe8c9a7d256a7f77f46b806a480569b064cfca332f5eab5d3b1cfb0182

Observation 44b884ca-911e-4888-bdc4-d7fd8c394d63 · outbound

This paper cites Recognize Anything: A Strong Image Tagging Model.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Recognize Anything: A Strong Image Tagging Model

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.953580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.953580Z digest=sha256:bce97806793aa8d3ae9db80b0535a189cf7f60b588b00b1348b17623a1a6e485

Observation 692784a0-41d9-449d-ae1a-f18255bc8761 · outbound

This paper cites Mitigating Object Hallucination in Large Vision-Language Models via Image-Grounded Guidance.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Mitigating Object Hallucination in Large Vision-Language Models via Image-Grounded Guidance

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.958272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.958272Z digest=sha256:926c33249c64f47672b0be4ff5015dc405c72c3b6ac1e998e203511d0bf0faed

Observation e37eccb4-077b-490d-8871-adb8cb90c6ee · outbound

This paper cites Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.962572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.962572Z digest=sha256:a475add35ed090863de8b4b2b972121ed5c8892eccfbe0f6b4b19f445d76047d

Observation 332e1d13-7462-41bc-a4f0-d5d64fe41616 · outbound

This paper cites Aligning Modalities in Vision Large Language Models via Preference Fine-tuning.

AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization Aligning Modalities in Vision Large Language Models via Preference Fine-tuning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-16T11:27:36.967307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:27:36.967307Z digest=sha256:9c45edfd68ea0c64d7c5f917d00bd6ce7980c5054da0a73d31da2b3643120db6

Pith citing papers

Observation 96fb79cb-15f2-49a7-a315-a5403f276793 · inbound

Experience Augmented Policy Optimization for LLM Reasoning cites this paper.

Experience Augmented Policy Optimization for LLM Reasoning AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-06-30T06:54:20.299776Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-06-30T06:52:55.368187Z digest=sha256:061063e8e5e8479721595bf9a2c4b4ea8ec9faba57c78e7ca95cc06fd7c23df4

Observation d05ecd30-59cf-4ae0-81dc-2c38046ca202 · inbound

Experience Augmented Policy Optimization for LLM Reasoning cites this paper.

Experience Augmented Policy Optimization for LLM Reasoning AdaViP: Aligning Multi-modal LLMs via Adaptive Vision-enhanced Preference Optimization

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-02T09:32:15.287630Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:32:15.287630Z digest=sha256:dd0e1c7f7ef4cef829bff5f88b650defebcb0340260eb66f9121f85a9a09abd0