Pith. sign in

Paper Citation Record · LEDGER

Detecting Text Manipulation in Images using Vision Language Models

As of 9 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 1 inbound Pith citation observation for arXiv:2509.10278.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.10278 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T18:04:20.574759Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-03T20:02:56.057102Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:08:55.138666Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 70313364-3ac7-4635-bfda-c39daa62fee4 · outbound

This paper cites Qwen2.5-VL Technical Report.

Detecting Text Manipulation in Images using Vision Language Models Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:18.054749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:18.054749Z digest=sha256:cc42080fda5159140afc02a6f8f93251429561a42569ac0da75a7c17d90b822f

Observation 9c311dd9-9dc4-4572-9b57-9db1271ac223 · outbound

This paper cites Textdiffuser- 2: Unleashing the power of language models for text rendering.

Detecting Text Manipulation in Images using Vision Language Models Textdiffuser- 2: Unleashing the power of language models for text rendering

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:18.254757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:18.254757Z digest=sha256:90ae737bd6b9000bcdf08288c74c1954054c6e24e6891bf84bfe9e0a9a7210a3

Observation 2053abec-4f97-4f26-b31d-20405ddcceb2 · outbound

This paper cites Image manipulation detection by multi-view multi-scale supervision.

Detecting Text Manipulation in Images using Vision Language Models Image manipulation detection by multi-view multi-scale supervision

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:18.425371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:18.425371Z digest=sha256:b7d1780fb6ad7d00244781214f64429961551ae86a48f4d4ee6a1a7d06daebbc

Observation af2f30ca-4554-4e2a-9bc9-d1c5b424d0b3 · outbound

This paper cites Chatbot arena: An open platform for evaluating llms by human preference.

Detecting Text Manipulation in Images using Vision Language Models Chatbot arena: An open platform for evaluating llms by human preference

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:18.518832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:18.518832Z digest=sha256:e4a19ff0c969b801240beaf6290c03641b32bcdb252040dc090ce34b275c3c67

Observation 67b6ed52-900b-46f3-8cb4-77cb1ff70fbf · outbound

This paper cites On the detection of digital face manipulation.

Detecting Text Manipulation in Images using Vision Language Models On the detection of digital face manipulation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:18.632298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:18.632298Z digest=sha256:13be95620a8691ff6ee8563c6e0a9237eae77fcc7c388784ff5e15e0e28221ad

Observation f10857ed-5151-46ae-9b21-d3ca7fa3f3a6 · outbound

This paper cites Mvss-net: Multi-view multi- scale supervised networks for image manipulation detection.IEEE Transactions on Pattern Anal- ysis and Machine Intelligence, 45(3):3539–3553, 2022.

Detecting Text Manipulation in Images using Vision Language Models Mvss-net: Multi-view multi- scale supervised networks for image manipulation detection.IEEE Transactions on Pattern Anal- ysis and Machine Intelligence, 45(3):3539–3553, 2022

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:18.712784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:18.712784Z digest=sha256:68af6824d09071f845488ab25fafd1c985e77371dc3b818ee3e9f8d1ba9df638

Observation 9d389dfa-c654-422d-acde-acf72991fc14 · outbound

This paper cites Casia image tampering detection evaluation database.

Detecting Text Manipulation in Images using Vision Language Models Casia image tampering detection evaluation database

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:19.004846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:19.004846Z digest=sha256:6ca36006a7d98c5009fb1f588e1ac45353cad808b4d8ae715bf58e4be6130bcf

Observation ee526ebb-f2ea-4c8b-9b16-f86c92c165db · outbound

This paper cites AMMeBa: A Large-Scale Survey and Dataset of Media-Based Misinformation In-The-Wild.

Detecting Text Manipulation in Images using Vision Language Models AMMeBa: A Large-Scale Survey and Dataset of Media-Based Misinformation In-The-Wild

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:19.084829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:19.084829Z digest=sha256:543e3af65765c5879243a2d792c338ee2b6106fa76b6cc9cfb424628df3f9b67

Observation 3be553d2-7ed7-42da-8b78-37e74bb50334 · outbound

This paper cites The Llama 3 Herd of Models.

Detecting Text Manipulation in Images using Vision Language Models The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:19.213306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:19.213306Z digest=sha256:a97d09d73806471fecffeeafa8431cacbee05e697ff85f59b33fe357f2a6f09f

Observation 364d1e11-953a-4e0f-8f92-19b467d019fd · outbound

This paper cites Tru- for: Leveraging all-round clues for trustworthy image forgery detection and localization.

Detecting Text Manipulation in Images using Vision Language Models Tru- for: Leveraging all-round clues for trustworthy image forgery detection and localization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:19.444751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:19.444751Z digest=sha256:a0836a29a95184746b9ac43c9d6a5f21eb912f950ae466e306a1319dd48d0012

Observation de5a0f2b-6b45-472e-8df7-e002ff5e1642 · outbound

This paper cites Hier- archical fine-grained image forgery detection and localization.

Detecting Text Manipulation in Images using Vision Language Models Hier- archical fine-grained image forgery detection and localization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:19.577394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:19.577394Z digest=sha256:3533f88256be76c70a27d12ad93f789d5043e817a4ff481c60c365282b18ddd4

Observation 8bc15881-3022-4427-8dcc-7fea0464ccfd · outbound

This paper cites SIDA: Social Media Image Deepfake Detection, Localization and Explanation with Large Multimodal Model.

Detecting Text Manipulation in Images using Vision Language Models SIDA: Social Media Image Deepfake Detection, Localization and Explanation with Large Multimodal Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:19.694757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:19.694757Z digest=sha256:8bec33ae824d2f8aaa6cc619c2d5fd45b4523e81b26a204a8393abed40846614

Observation 319047f5-1dad-4a24-b3e9-5dafa75af0af · outbound

This paper cites The point where reality meets fan- tasy: Mixed adversarial generators for image splice detection.Advances in neural information processing systems, 32, 2019.

Detecting Text Manipulation in Images using Vision Language Models The point where reality meets fan- tasy: Mixed adversarial generators for image splice detection.Advances in neural information processing systems, 32, 2019

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:19.870008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:19.870008Z digest=sha256:c9153c14a052f9cc35dd022dfd262838356c6d9021a7b13272d3b700496070ad

Observation 06f360f2-0ac5-4bdd-8eb3-c4d96ac76b3d · outbound

This paper cites Exploring chatgpt for face presentation attack detection in zero and few-shot in-context learning.

Detecting Text Manipulation in Images using Vision Language Models Exploring chatgpt for face presentation attack detection in zero and few-shot in-context learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:20.064836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:20.064836Z digest=sha256:6d3d0aaa6994d3244c06df80b0937e48f5e9eb72e79b7c5edc2a86affedcaa43

Observation fb3976a8-f6d0-4f6d-a82f-60e8bfa54103 · outbound

This paper cites FantasyID: A dataset for detecting digital manipulations of ID-documents.

Detecting Text Manipulation in Images using Vision Language Models FantasyID: A dataset for detecting digital manipulations of ID-documents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:20.314744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:20.314744Z digest=sha256:94ffd7fd0d0644654c707d29a196dbe99aae8d0b2201f6f56ee6aa91cae3db99

Observation 99a6db89-cbf4-4995-9191-9f09dc58bc09 · outbound

This paper cites Cat-net: Compression artifact tracing network for detection and localization of image splicing.

Detecting Text Manipulation in Images using Vision Language Models Cat-net: Compression artifact tracing network for detection and localization of image splicing

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:20.437130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:20.437130Z digest=sha256:7a34b8d446c24ef3b4b26f44952b4d34100bc4b27229f16353ba27e42f3bbdab

Observation 81f44cdf-b203-4b89-aef2-6b9571f8938c · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

Detecting Text Manipulation in Images using Vision Language Models Lisa: Reasoning segmentation via large language model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:20.574759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:20.574759Z digest=sha256:12a62ab708e68271089c7c9c6ece4ec01627d3da26b1b84c823f683300585634

Pith citing papers

Observation 26fe4bc7-e59a-4e4e-acee-314a42cff1f3 · inbound

From Forgeries to Foundation Models: A Systematic Survey of Identity Document Attack and Detection cites this paper.

From Forgeries to Foundation Models: A Systematic Survey of Identity Document Attack and Detection Detecting Text Manipulation in Images using Vision Language Models

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:08:55.141276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-03T20:02:56.057102Z digest=sha256:23638e7cf76187dab5051241980534ccf9a985cc3c3c84c6d10c633c9d30d48b