Pith. sign in

Paper Citation Record · LEDGER

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning

As of 14 August 2026, this Paper Citation Record lists 66 of 66 outbound references and 0 inbound Pith citation observations for arXiv:2506.07227.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07227 v1

Coverage vector

measured 66 of 66 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:43:48.054205Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

66 of 66 outbound references displayed

  • verified exact1
  • verified fuzzy32
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2de8716a-b311-4ff6-96fe-0cdd660016d0 · outbound

This paper cites URL https://api.semanticscholar.org/CorpusID:276612236.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning URL https://api.semanticscholar.org/CorpusID:276612236

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:59.507971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:38.608937Z digest=sha256:5aa949d5eb384dd003374dbaa197a23821b89e7b357710018a83c3756d7ac33d

Observation bca26de9-f16f-4a11-a98c-8a064ec72ebc · outbound

This paper cites GPT-4 Technical Report.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:38.711206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:38.711206Z digest=sha256:524e3971dbcef201dfb7530ca56cfee266649db6746e17e6d75fd65299bfdfa3

Observation d6ad0563-42e0-4b3e-a964-2444405eca28 · outbound

This paper cites Qwen2.5-VL Technical Report.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:38.884823Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:38.884823Z digest=sha256:ed00a5633b47e14e1037c84a9d9e583ce2afd28198f0e5c207e4ad2b814d49d8

Observation 2a3bf250-33e3-4e51-84b1-e5994c2a80f7 · outbound

This paper cites Hallucination of multimodal large language models: A survey .CoRR, 2024.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Hallucination of multimodal large language models: A survey .CoRR, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:59.171317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:39.041996Z digest=sha256:53581aafb4fd4872aa38bf7af774df8978d96a9614f57e6163deefcbb2e352cc

Observation 423e27ac-0282-4dda-9410-db59986b3274 · outbound

This paper cites Flux.1 fill [dev].

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Flux.1 fill [dev]

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:58.861034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:39.254218Z digest=sha256:bba9f2dbf0bd38881feaea6568ec88884cc2f1c18be066a84815a817fb590770

Observation 5ecf6e1c-d64d-4ff5-ad5e-c9aecd056bb3 · outbound

This paper cites Ledits++: Limitless image editing using text-to-image models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Ledits++: Limitless image editing using text-to-image models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:58.484131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:39.375131Z digest=sha256:f9ca790347f9f5b9ad8fe6d385484682e807a2263c371105688c9b556ec8a770

Observation 662d1733-67bc-49c9-991b-986ec3bad760 · outbound

This paper cites Instructpix2pix: Learning to follow image editing instructions.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Instructpix2pix: Learning to follow image editing instructions

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:58.126442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:39.449044Z digest=sha256:39e8b7268c2c6d216cdf570fb13a3c84f87d67ff27db3594859d446c20804ed4

Observation c0ed1ce1-9362-47b4-9348-0c32bc4c05d5 · outbound

This paper cites The revolution of multimodal large language models: a survey .arXiv preprint arXiv:2402.12451, 2024.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning The revolution of multimodal large language models: a survey .arXiv preprint arXiv:2402.12451, 2024

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:39.543224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:39.543224Z digest=sha256:00d53a777b8a1cbc8c0d3ae1c5c1e3c01f1e2273ac9157b0c8c7440ebab526dd

Observation 06591c43-d4d1-42e2-bec5-d356c2e0dd67 · outbound

This paper cites Vlmimic: Vision language models are visual imitation learner for fine-grained actions.Advances in Neural Information Processing Systems, 37:77860–77887, 2024.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Vlmimic: Vision language models are visual imitation learner for fine-grained actions.Advances in Neural Information Processing Systems, 37:77860–77887, 2024

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:57.810907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:39.598014Z digest=sha256:768c7d31c035b7c42bbd0de13bbd010e43ec58c35c065cdc22bf83e1109385a1

Observation e6a69b11-b329-4a2a-83ce-f3fd8ebd5ac0 · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:39.736678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:39.736678Z digest=sha256:9b79faec656068184aff3f6fdc276cd78c29b111d5932471bf8261aa0684ffa2

Observation c46b15e7-cf14-4543-9487-ea9126ad7257 · outbound

This paper cites Anydoor: Zero-shot object-level image customization.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Anydoor: Zero-shot object-level image customization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:57.415831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:39.854427Z digest=sha256:d1ed06646a2eff7ecb932711b7a20aed106bc7af0053efd5f74132ef8323332d

Observation 30ddf113-d876-428b-b26e-898c2a7e035f · outbound

This paper cites UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning UniReal: Universal Image Generation and Editing via Learning Real-world Dynamics

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:40.028036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:40.028036Z digest=sha256:18529182f0b7c7668269f026bc2e67ec46da9fb161931846bedcb4f822302f75

Observation 9c6eee60-6280-4765-813f-2a5265c515b8 · outbound

This paper cites Diffusion self-guidance for controllable image generation.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Diffusion self-guidance for controllable image generation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:40.204373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:40.204373Z digest=sha256:b8c3722651c52d232cff398780920252b7f30ee0c93bd17f66487b414e63b3c1

Observation ceaae5c7-6498-4941-8dce-05f664adf7a8 · outbound

This paper cites MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning MME-Survey: A Comprehensive Survey on Evaluation of Multimodal LLMs

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:40.351303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:40.351303Z digest=sha256:5a8cc0bdc2b21e5341ea28f9800c4513e567b63cdaa347291ed7f6d00873ed52

Observation 2dd79992-184b-4994-89f0-007c48796424 · outbound

This paper cites TLDR: Token-Level Detective Reward Model for Large Vision Language Models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning TLDR: Token-Level Detective Reward Model for Large Vision Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:43:48.602722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:40.500892Z digest=sha256:4bf8d8b44e02d83f1d26d8c35194607fd87a62b75406193ffeeb851a227cd754

Observation 76f01343-d522-4ef2-82d9-fee5a2b553c0 · outbound

This paper cites Towards a benchmark of multimodal large language models for industrial engineering.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Towards a benchmark of multimodal large language models for industrial engineering

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:57.107153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:40.648100Z digest=sha256:bea1883a2e03021c1a88ce5d6985d8eed86674a4a7a63abc01c9a8c8c4cda875

Observation 76a6500e-94ff-48d3-be82-8f5016ad95ea · outbound

This paper cites Generative adversarial networks.Communications of the ACM, 63(11):139–144, 2020.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Generative adversarial networks.Communications of the ACM, 63(11):139–144, 2020

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:56.719981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:40.743202Z digest=sha256:1450bbe4bc953ba24149fdd42bf8a23eeef0f9e24b0a2e8be620cb90522f11a5

Observation a92043ad-40b2-42c7-8553-5c2f0f6a01c3 · outbound

This paper cites Gemini 2.5 flash.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Gemini 2.5 flash

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:56.397687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:40.943773Z digest=sha256:a7c102c42f42ec2303a3f6b42f736cbb4e3f254d3a93845e21a2d4291b7c2afa

Observation dee98651-ea66-4bc2-986c-1b7b5327c82f · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answering.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Making the v in vqa matter: Elevating the role of image understanding in visual question answering

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:56.065207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:41.114382Z digest=sha256:b27127f74cd3f8630a831df765b89b5919c2069e28fcfcdc48a20afadbb41536

Observation e6543063-4d93-461e-a0b3-51ece3c23a4e · outbound

This paper cites The Llama 3 Herd of Models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning The Llama 3 Herd of Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:41.230598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:41.230598Z digest=sha256:761898ccbc0255f35025df56848839ad1b012ca106be0a571a1df388f0ed9f19

Observation d1a55e4f-cf1d-44ad-b42f-285b3a768821 · outbound

This paper cites Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Denoising diffusion probabilistic models.Advances in neural information processing systems, 33:6840–6851, 2020

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:55.619250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:41.386418Z digest=sha256:3b376300663529cc7aa747d73754744b2cf37709756f30fce3aa1abecefeb8b5

Observation 6ac500f9-93a6-4298-a548-a3f564b985ed · outbound

This paper cites Diffusion Model-Based Image Editing: A Survey.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Diffusion Model-Based Image Editing: A Survey

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:41.528747Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:41.528747Z digest=sha256:533ff12925862dd877e45e3603e8971ce6799ecfd38245e16f9f79912038762d

Observation 047225bb-e58e-49e4-ad4b-298873d44d7f · outbound

This paper cites Smartedit: Exploring complex instruction-based image editing with multimodal large language models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Smartedit: Exploring complex instruction-based image editing with multimodal large language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:55.263948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:41.681296Z digest=sha256:ef2d3e668f9f44e502da5481c395975afbf22c924ae52d1b1f88d9d95aac67d4

Observation d8922cbf-9b68-4df8-81f1-1ee09e87f00a · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:54.928422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:41.824325Z digest=sha256:73894909f30fa6e903d4b3f9f607465a173dab290f4b1690ce22a4ede4727298

Observation 5b9c44cb-39a4-488a-9d8b-3b679c278984 · outbound

This paper cites GPT-4o System Card.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning GPT-4o System Card

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:42.006009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:42.006009Z digest=sha256:2b29ec54f7ee12633020ec7e521659dd7c0280cf465c4872e2d0b6722f62ffd0

Observation d0a575a1-2eee-43b7-9f09-029c7d9a1366 · outbound

This paper cites A style-based generator architecture for generative adversarial networks.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning A style-based generator architecture for generative adversarial networks

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:54.551988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:42.126473Z digest=sha256:6583028096972fb3206da4e07f0cd1547b84ce32c411c173015fe4155c2574f5

Observation 555ec626-1b0d-4dc4-915a-bc6c28fc7849 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Visual genome: Connecting language and vision using crowdsourced dense image annotations

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:42.310898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:42.310898Z digest=sha256:55704be32f3fc1f222113bdb1006d4d58f04f27287eb37603afcfc43b8657ad2

Observation 2e22816e-3bbe-49f3-a114-f95cd95da233 · outbound

This paper cites Evaluating and improving compositional text-to-visual generation.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Evaluating and improving compositional text-to-visual generation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:54.203404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:42.448452Z digest=sha256:06e8ee523f8e3162549e00a0845545592408ff1a90f0c4d0eda0992cd4f6423f

Observation d0a127ad-0c9a-430a-b980-f740fee99f64 · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Evaluating Object Hallucination in Large Vision-Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:42.594952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:42.594952Z digest=sha256:2ca80a3d32cec819acae9b2a7760cc5b16d16de9d96f7ae54777574d73a1e1f6

Observation e5c16a2e-1828-48e6-a089-e465466d583e · outbound

This paper cites Evaluating text-to-visual generation with image-to-text generation.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Evaluating text-to-visual generation with image-to-text generation

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:53.827596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:42.736228Z digest=sha256:5ae30e3c46fe5c2bf749025305d1b34d5522e6c3a42fc20779b5116ff643a59a

Observation 01af5ce7-54ed-4e9c-807b-54fd8acf05e5 · outbound

This paper cites Flow Matching for Generative Modeling.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Flow Matching for Generative Modeling

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:42.883509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:42.883509Z digest=sha256:58cf59abb9cd44f497fc989ba493752be1ee4ba134ad3dab4b9070fbbb606492

Observation 27be5de5-b3ff-4872-a905-992872b5b4d5 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Improved baselines with visual instruction tuning, 2023

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:43.066529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:43.066529Z digest=sha256:c3254f31171d828c40e33f4964905ba733752c9bfab4cb833d326050572fc8b7

Observation 848a6b9b-f94d-49a4-889d-173ff694dbe8 · outbound

This paper cites Step1X-Edit: A Practical Framework for General Image Editing.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Step1X-Edit: A Practical Framework for General Image Editing

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:43.205553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:43.205553Z digest=sha256:4cf4614dc31ae83eff9228fe8873c4ad4849ffac8df7231acb0be8a0ded92d10

Observation af6e4322-9d20-45ca-b28b-218f0f494c27 · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:43.306278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:43.306278Z digest=sha256:378525d4edb5648714c4e5e097f8b740375ce4cd7870463a92d403a206ec0ceb

Observation fefda47f-026b-4584-8273-f95ab4f3f7a0 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pp.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pp

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:53.466271Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:43.488660Z digest=sha256:9b092499f1c2cc0ecfbdeedc76ea15f208b44245d11048055d28c5e18ad49bdd

Observation cd9cbec6-efba-4ef9-8ee6-195b3fe737c3 · outbound

This paper cites Customizable image synthesis with multiple subjects.Advances in neural information processing systems, 36:57500–57519, 2023.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Customizable image synthesis with multiple subjects.Advances in neural information processing systems, 36:57500–57519, 2023

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:53.106964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:43.656077Z digest=sha256:756f2cdac47f86c62294ac368fcfd963432f8036f18525f6011368ad666ea470

Observation ad8ec765-f698-4aa8-8ebc-9c83b96732ad · outbound

This paper cites MagicQuill: An Intelligent Interactive Image Editing System.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning MagicQuill: An Intelligent Interactive Image Editing System

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:43.759971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:43.759971Z digest=sha256:a0e31c6a731802469e5a9e04831a7a1983585f4a89572557b7c74ae8c58da9f3

Observation d76e4595-fe58-4982-b974-0dcafeb317c7 · outbound

This paper cites Fine-grained Image Editing by Pixel-wise Guidance Using Diffusion Models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Fine-grained Image Editing by Pixel-wise Guidance Using Diffusion Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:43.858271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:43.858271Z digest=sha256:d785334c18a416e7727fcc148667eb2c7304c9a28f578d27c8d933e159255de1

Observation a609cc02-e81c-44b7-99a2-3fd8f5996f94 · outbound

This paper cites Mm1: methods, analysis and insights from multimodal llm pre-training.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Mm1: methods, analysis and insights from multimodal llm pre-training

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:52.759445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:44.006676Z digest=sha256:27da8c761fa802f8fe48a294a1762e656c89ed6c15a2edef8cb45b0059695a71

Observation 4d5993ce-9291-4d6a-abb7-e2b9c4a38b10 · outbound

This paper cites SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning SDEdit: Guided Image Synthesis and Editing with Stochastic Differential Equations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:44.132494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:44.132494Z digest=sha256:2837c0f26a6ddf41b7dbb0ccf717adb719e2f72275a1e03a8146127bf7bf3c3f

Observation 568f2b0c-7928-450d-8e16-dcbddcbffc46 · outbound

This paper cites DragonDiffusion: Enabling Drag-style Manipulation on Diffusion Models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning DragonDiffusion: Enabling Drag-style Manipulation on Diffusion Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:44.281953Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:44.281953Z digest=sha256:da79dfb466bf6e609bba0664ddfe55794762d3c32c4471bdf5419669b1272857

Observation 2fc74a18-437b-4103-9df5-adaac60d1a1c · outbound

This paper cites Docci: Descriptions of connected and contrasting images.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Docci: Descriptions of connected and contrasting images

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:52.419574Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:44.426315Z digest=sha256:c16d79bfdbff1700eb9f7e06b39341d99fa9511501c43e03868dbe9638d10857

Observation 01206368-d747-4e59-9a64-aa0084941692 · outbound

This paper cites Drag your gan: Interactive point-based manipulation on the generative image manifold.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Drag your gan: Interactive point-based manipulation on the generative image manifold

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:52.139918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:44.593681Z digest=sha256:d1659f8f8fda2e06a6d854f41cce457ca222d5256906c584eee0dcdfd301e1fc

Observation 7d1111be-3583-4d87-91b0-2f76161e47d2 · outbound

This paper cites Evaluating LLM -- Generated Multimodal Diagnosis from Medical Images and Symptom Analysis.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Evaluating LLM -- Generated Multimodal Diagnosis from Medical Images and Symptom Analysis

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:44.803144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:44.803144Z digest=sha256:c9ea04774c2fd5eca0f00b1531971bf3f521febddfd5e148945a0587e1c2cef5

Observation adc2cbee-e611-4ec5-87e2-4bb6e832c22b · outbound

This paper cites Synthesize diagnose and optimize: Towards fine-grained vision-language understanding.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Synthesize diagnose and optimize: Towards fine-grained vision-language understanding

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:51.859033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:44.947642Z digest=sha256:db933e27141cde4518e0e78e1c0866de6a309a1b93ab511f3926bbfc2b1bad79

Observation 8550cc06-d2f2-421a-b074-ece6b8b6db09 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Learning transferable visual models from natural language supervision

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:51.523948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:45.089647Z digest=sha256:0f40db3ada87def217903ceb7e872d181b8a0e979752a53ff463f84c30fd2995

Observation 2d389b6c-0f67-4f42-af7d-3ab502b2dc24 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning High-resolution image synthesis with latent diffusion models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:51.197125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:45.285969Z digest=sha256:2e8ef50a3183b21961f3d0d83b8095e6973fa3773f1488eaf5506175f59d9660

Observation d7148344-4a87-4cec-9daa-1b2ce1559118 · outbound

This paper cites Towards vqa models that can read.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Towards vqa models that can read

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:50.871131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:45.403062Z digest=sha256:c9a948caa5060de2e7f9aaf56d72ca8f4b5dfdb88b4e11c6190469de0bead1ee

Observation 96e3fc7b-ebf8-4fd2-b77a-5842f29604d0 · outbound

This paper cites Denoising Diffusion Implicit Models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Denoising Diffusion Implicit Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:45.549277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:45.549277Z digest=sha256:45b19d7a4be46dea66f36501079ccb0ac10db4cb3748baef719858bf390d2a96

Observation 044c6697-aa23-4e2f-943f-064832cfdae0 · outbound

This paper cites Emu: Generative Pretraining in Multimodality.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Emu: Generative Pretraining in Multimodality

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:45.643297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:45.643297Z digest=sha256:edd900b4e65a1d92705b5e507f1cc9d1891b401176d27c587ff8b44c2efca2fb

Observation af8ef4b5-8b9f-4203-bcd4-7f186fdc9c5e · outbound

This paper cites Generative multimodal models are in-context learners.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Generative multimodal models are in-context learners

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:50.613669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:45.803662Z digest=sha256:9d04ac3d67f2d53941110a785800ecb6b2933e0dfb2dd69ac89727c7d5063827

Observation c5c494f0-50f1-4823-b2b6-bad93ebce212 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:45.997544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:45.997544Z digest=sha256:e0dae47fa87140ed3491f0fd3a55615e797367733f07ec4bf9ed740271e269f5

Observation a0777960-5a47-4dc7-89ae-a04082045c34 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:50.258784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:46.166326Z digest=sha256:db14cada88c800a65765c712e04b4f0728d69d3e28407b5501ba736b89f0be46

Observation 00dbf96f-a28d-4884-b1fd-07ac49ee689a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:46.340293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:46.340293Z digest=sha256:4d8e8373379dd4e57fb94c7eca36bada04805649b3ea8c1f80f4980fa80d1918

Observation 70ba53b7-68aa-4836-99a8-2ce4865f6875 · outbound

This paper cites Image inpainting with external-internal learning and monochromic bottleneck.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Image inpainting with external-internal learning and monochromic bottleneck

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:49.954470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:46.490587Z digest=sha256:9a6ce14908c4a8008c12c0765407e51b87a5f2549b831390891f7bd4f0876bfb

Observation 6e62a470-9110-4471-9613-977118ae06a6 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Emu3: Next-Token Prediction is All You Need

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:46.605395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:46.605395Z digest=sha256:7743ab75c58478a637210e87e0896457bb7e7dec23580b367e91fbbf333bca9c

Observation 8655a29c-b26d-4aa5-8575-aa726a637913 · outbound

This paper cites Uni-paint: A unified framework for multimodal image inpainting with pre- trained diffusion model.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Uni-paint: A unified framework for multimodal image inpainting with pre- trained diffusion model

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:49.620544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:46.748089Z digest=sha256:689c480a93013bc8717be60bc0b4ee51c8b89cdf9bb1a5cc5d12785cb2297625

Observation 0faa7296-ebbd-4f28-8003-dc62ddc19aa5 · outbound

This paper cites Imagebrush: Learning visual in-context instructions for exemplar-based image manipulation.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Imagebrush: Learning visual in-context instructions for exemplar-based image manipulation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:49.316942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:46.871389Z digest=sha256:9a33c88141a31a1f9341dd559099e4023c4f97f68f6ea158ecd1317b738243c8

Observation 89789c16-fd0a-4782-a737-d1b9b93d86ca · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:47.046094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:47.046094Z digest=sha256:32fd6b9d8f283be4b8ff25fc8e0af58b4d6cb744ba1cd6b5ef5dff5f7538f0de

Observation f38df3e5-9889-4c8b-a62a-44c5a6f99045 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:47.201274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:47.201274Z digest=sha256:1dd8964a193949d2dd88b743ace2c227bf8e420ca1ab0b39a11af8cc762ec66f

Observation bcbbe6df-eef7-4337-be5b-744b37dbf1b9 · outbound

This paper cites MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning MM-Vet v2: A Challenging Benchmark to Evaluate Large Multimodal Models for Integrated Capabilities

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:47.364932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:47.364932Z digest=sha256:2d85764b0e85b3e9a7d77ecbdcd19364cd191a999575e3fcc0520834c601b3bd

Observation 1b62c9c2-7faa-47de-b423-b5ce211a1ec8 · outbound

This paper cites WalkVLM:Aid Visually Impaired People Walking by Vision Language Model.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning WalkVLM:Aid Visually Impaired People Walking by Vision Language Model

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:47.468600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:47.468600Z digest=sha256:8b74a1ef836b3c6d36d6bdca6d35e1e1ed2c16922e8aa08feca7a7a5c948c41a

Observation 71db210a-92bf-4f4d-8565-0b4df8778aea · outbound

This paper cites Mmvp: A multimodal mocap dataset with vision and pressure sensors.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Mmvp: A multimodal mocap dataset with vision and pressure sensors

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:43:48.999053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:43:47.607411Z digest=sha256:7053ddfc2e3a67e555d5544644052e3e87259b6e2062001033028c25cfd48ba4

Observation 91ac0d10-9a62-47b6-8b17-530ea3d2b2dc · outbound

This paper cites MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning MLLMs Know Where to Look: Training-free Perception of Small Visual Details with Multimodal LLMs

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:47.780118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:47.780118Z digest=sha256:6c1867cebfbfe52ec6029ac6242e887e385b84f45178e24adf9c73b329334e52

Observation c9513c2c-afe2-4653-ae9d-c04995f7b64d · outbound

This paper cites Magicbrush: A manually annotated dataset for instruction-guided image editing.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning Magicbrush: A manually annotated dataset for instruction-guided image editing

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:47.935836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:47.935836Z digest=sha256:dfd6ebede0ce838d75106e73be67451ca4c6e8e69505e010699e43399537b36a

Observation 47a0b8e6-6264-4127-9748-95e7cd8a76be · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Hallucination at a Glance: Controlled Visual Edits and Fine-Grained Multimodal Learning LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T05:43:48.054205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:43:48.054205Z digest=sha256:426f3c92418b3d84831f183ce81b27261d87b6d0e61b2d2789f32fa48bf8bb1c

Pith citing papers

No inbound Pith citation observations are available.