Pith. sign in

Paper Citation Record · LEDGER

Detecting Text Manipulation in Images using Vision Language Models

As of 22 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 1 inbound Pith citation observation for arXiv:2509.10278.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.10278 v1

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T18:04:20.574759Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-03T20:02:56.057102Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T20:08:55.138666Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 70313364-3ac7-4635-bfda-c39daa62fee4 · outbound

This paper cites Qwen2.5-VL Technical Report.

Detecting Text Manipulation in Images using Vision Language Models Qwen2.5-VL Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:18.054749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:18.054749Z digest=sha256:97164e2513ffb337c9a74a220b87fa4d4ee2902b21e1b02c7e775f2046ce660f

Observation 9c311dd9-9dc4-4572-9b57-9db1271ac223 · outbound

This paper cites Textdiffuser- 2: Unleashing the power of language models for text rendering.

Detecting Text Manipulation in Images using Vision Language Models Textdiffuser- 2: Unleashing the power of language models for text rendering

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:18.254757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:18.254757Z digest=sha256:1cf875b916646c7a1a1a658f40c55f68c5b097080568de14c6b1f6569e9f3c5b

Observation 2053abec-4f97-4f26-b31d-20405ddcceb2 · outbound

This paper cites Image manipulation detection by multi-view multi-scale supervision.

Detecting Text Manipulation in Images using Vision Language Models Image manipulation detection by multi-view multi-scale supervision

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:18.425371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:18.425371Z digest=sha256:d74c249581c4b825e1f568f584f83b19b214e724bbaf16d30e8f0f5c9c1741e2

Observation af2f30ca-4554-4e2a-9bc9-d1c5b424d0b3 · outbound

This paper cites Chatbot arena: An open platform for evaluating llms by human preference.

Detecting Text Manipulation in Images using Vision Language Models Chatbot arena: An open platform for evaluating llms by human preference

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:18.518832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:18.518832Z digest=sha256:8a01749d67c86d214cbda44120efe136a51047414ed2c0c37e268547706f38d9

Observation 67b6ed52-900b-46f3-8cb4-77cb1ff70fbf · outbound

This paper cites On the detection of digital face manipulation.

Detecting Text Manipulation in Images using Vision Language Models On the detection of digital face manipulation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:18.632298Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:18.632298Z digest=sha256:f4d532fbc0b2023e44da33d69fc812d7164a2bb915e1cf4d7177eecfd8095cb9

Observation f10857ed-5151-46ae-9b21-d3ca7fa3f3a6 · outbound

This paper cites Mvss-net: Multi-view multi- scale supervised networks for image manipulation detection.IEEE Transactions on Pattern Anal- ysis and Machine Intelligence, 45(3):3539–3553, 2022.

Detecting Text Manipulation in Images using Vision Language Models Mvss-net: Multi-view multi- scale supervised networks for image manipulation detection.IEEE Transactions on Pattern Anal- ysis and Machine Intelligence, 45(3):3539–3553, 2022

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:18.712784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:18.712784Z digest=sha256:6d82f1e8dfab39d0c705095d96486b1b9660ff21148407e314e9a6e0a28075b2

Observation 9d389dfa-c654-422d-acde-acf72991fc14 · outbound

This paper cites Casia image tampering detection evaluation database.

Detecting Text Manipulation in Images using Vision Language Models Casia image tampering detection evaluation database

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:19.004846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:19.004846Z digest=sha256:58eca8afc981773388c5cae6876c9ab94f455b5f7dfe1254f4cd885a35b2d3f8

Observation ee526ebb-f2ea-4c8b-9b16-f86c92c165db · outbound

This paper cites AMMeBa: A Large-Scale Survey and Dataset of Media-Based Misinformation In-The-Wild.

Detecting Text Manipulation in Images using Vision Language Models AMMeBa: A Large-Scale Survey and Dataset of Media-Based Misinformation In-The-Wild

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:19.084829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:19.084829Z digest=sha256:a369d1784536c081a363dbb68ce2c4dfc80d3db3a17e0644b7c043498fd5fc01

Observation 3be553d2-7ed7-42da-8b78-37e74bb50334 · outbound

This paper cites The Llama 3 Herd of Models.

Detecting Text Manipulation in Images using Vision Language Models The Llama 3 Herd of Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:19.213306Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:19.213306Z digest=sha256:85af672b8372ce1417fb85a4cfff1588dfa77aac89bea43829f7fb27cc63f525

Observation 364d1e11-953a-4e0f-8f92-19b467d019fd · outbound

This paper cites Tru- for: Leveraging all-round clues for trustworthy image forgery detection and localization.

Detecting Text Manipulation in Images using Vision Language Models Tru- for: Leveraging all-round clues for trustworthy image forgery detection and localization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:19.444751Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:19.444751Z digest=sha256:2c590ddd8059fcb363d3c6f08534d6e3a773c9ae0b046b765b511332dfc8ab11

Observation de5a0f2b-6b45-472e-8df7-e002ff5e1642 · outbound

This paper cites Hier- archical fine-grained image forgery detection and localization.

Detecting Text Manipulation in Images using Vision Language Models Hier- archical fine-grained image forgery detection and localization

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:19.577394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:19.577394Z digest=sha256:746b288873be143c735183b1d0d571f34520a6771ff6858ed9f8db632bf2c12c

Observation 8bc15881-3022-4427-8dcc-7fea0464ccfd · outbound

This paper cites SIDA: Social Media Image Deepfake Detection, Localization and Explanation with Large Multimodal Model.

Detecting Text Manipulation in Images using Vision Language Models SIDA: Social Media Image Deepfake Detection, Localization and Explanation with Large Multimodal Model

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:19.694757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:19.694757Z digest=sha256:03ac266310eb58995c2b9a3997b5d98029d9341a62fdc34ca87635a1bf69810f

Observation 319047f5-1dad-4a24-b3e9-5dafa75af0af · outbound

This paper cites The point where reality meets fan- tasy: Mixed adversarial generators for image splice detection.Advances in neural information processing systems, 32, 2019.

Detecting Text Manipulation in Images using Vision Language Models The point where reality meets fan- tasy: Mixed adversarial generators for image splice detection.Advances in neural information processing systems, 32, 2019

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:19.870008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:19.870008Z digest=sha256:7809d2c45e464a149db1b879e083a3fb3ed13eb7b189395ff0eb9e5ca1eec961

Observation 06f360f2-0ac5-4bdd-8eb3-c4d96ac76b3d · outbound

This paper cites Exploring chatgpt for face presentation attack detection in zero and few-shot in-context learning.

Detecting Text Manipulation in Images using Vision Language Models Exploring chatgpt for face presentation attack detection in zero and few-shot in-context learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:20.064836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:20.064836Z digest=sha256:ec8b3a5dc105f6afa52633b570230bf89c643cb0a47ebdf2c361262209cbc918

Observation fb3976a8-f6d0-4f6d-a82f-60e8bfa54103 · outbound

This paper cites FantasyID: A dataset for detecting digital manipulations of ID-documents.

Detecting Text Manipulation in Images using Vision Language Models FantasyID: A dataset for detecting digital manipulations of ID-documents

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:20.314744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:20.314744Z digest=sha256:ef57f718cdf084f943ec13d71cce8028d0ba6d8e87e52c01c6a27749e4e7bbfa

Observation 99a6db89-cbf4-4995-9191-9f09dc58bc09 · outbound

This paper cites Cat-net: Compression artifact tracing network for detection and localization of image splicing.

Detecting Text Manipulation in Images using Vision Language Models Cat-net: Compression artifact tracing network for detection and localization of image splicing

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:20.437130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:20.437130Z digest=sha256:edce7ca9b8165810955d028545fbfd6eb7facc62f03123f5c54fc3e6ebcb89d3

Observation 81f44cdf-b203-4b89-aef2-6b9571f8938c · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

Detecting Text Manipulation in Images using Vision Language Models Lisa: Reasoning segmentation via large language model

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T18:04:20.574759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T18:04:20.574759Z digest=sha256:cca7d0be1bf13368bf7b5263ff9273889499528c698c61a16e5e6f54ceb0cfb7

Pith citing papers

Observation 26fe4bc7-e59a-4e4e-acee-314a42cff1f3 · inbound

From Forgeries to Foundation Models: A Systematic Survey of Identity Document Attack and Detection cites this paper.

From Forgeries to Foundation Models: A Systematic Survey of Identity Document Attack and Detection Detecting Text Manipulation in Images using Vision Language Models

Reference 98

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:08:55.141276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-07-03T20:02:56.057102Z digest=sha256:1118af2f5c0e5bb34e7b289db98aefdda44aef8d2d2935333657ceb1de0d667d