Pith. sign in

Paper Citation Record · LEDGER

EchoSight: Advancing Visual-Language Models with Wiki Knowledge

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2407.12735.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2407.12735 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T14:33:43.532433Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:59:46.886481Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d3ef9275-f502-49c9-ab57-ffb0919c85a8 · inbound

Augmented Vision-Language Models: A Systematic Review cites this paper.

Augmented Vision-Language Models: A Systematic Review EchoSight: Advancing Visual-Language Models with Wiki Knowledge

Reference 121

Resolution
unresolved
no resolver link, observed 2026-08-06T14:33:43.532433Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:33:43.532433Z digest=sha256:9c1d3b5332f8c6f5a32de9dc75489a9dfae30dcf0f856f69f63cf9fb0081765a

Observation f7871e68-32ff-4d91-81aa-d9d2307a3a28 · inbound

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey cites this paper.

Empowering Multimodal LLMs with External Tools: A Comprehensive Survey EchoSight: Advancing Visual-Language Models with Wiki Knowledge

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-05T20:28:51.739526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:28:51.739526Z digest=sha256:7d27989042dc0693e3e8d32f8e52c7ae79f26b0484501f1feb3158f11ef76e7f

Observation 909a70fa-6c91-4148-aec9-d57a7798f601 · inbound

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum cites this paper.

Wiki-R1: Incentivizing Multimodal Reasoning for Knowledge-based VQA via Data and Sampling Curriculum EchoSight: Advancing Visual-Language Models with Wiki Knowledge

Reference 35

Resolution
unresolved
no resolver link, observed 2026-07-15T14:43:21.866055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T14:43:21.866055Z digest=sha256:dc8d7325679d0eb10cbef1580d2dde18bb62e593920cda298356eadb1e15ca2e

Observation b105937d-eadc-428e-bd23-551f0953ac4b · inbound

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition cites this paper.

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition EchoSight: Advancing Visual-Language Models with Wiki Knowledge

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T13:15:50.573740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-15T13:11:54.384284Z digest=sha256:3d6bcc26a67772d06c314079ef2ceb72ab95263d125e7db1cb230ede854c7087

Observation e30d155d-b208-49da-850f-831c8dcfda72 · inbound

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition cites this paper.

WikiCLIP: An Efficient Contrastive Baseline for Open-domain Visual Entity Recognition EchoSight: Advancing Visual-Language Models with Wiki Knowledge

Reference 43

Resolution
unresolved
no resolver link, observed 2026-07-14T23:55:24.006436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T23:55:24.006436Z digest=sha256:755c3fc6f761ff06ba99dfc1838cdbb109490640a9d5772c3ae9bd614a171a64

Observation 2b7e3776-9f5f-4795-8230-af053ee6119e · inbound

MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval cites this paper.

MCERF: Advancing Multimodal LLM Evaluation of Engineering Documentation with Enhanced Retrieval EchoSight: Advancing Visual-Language Models with Wiki Knowledge

Reference 33

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T09:30:48.227529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T09:29:32.250418Z digest=sha256:0ccf45f1c651a0e9c7f3957c5a29872a7721eb039913169c95a0f5dc8aaea068

Observation be4e461e-93f0-4222-a7a9-83751cd35bc2 · inbound

WikiVQABench: A Knowledge-Grounded Visual Question Answering Benchmark from Wikipedia and Wikidata cites this paper.

WikiVQABench: A Knowledge-Grounded Visual Question Answering Benchmark from Wikipedia and Wikidata EchoSight: Advancing Visual-Language Models with Wiki Knowledge

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-21T04:43:58.727343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-21T04:41:06.779529Z digest=sha256:94838de6e92536e2611b9adf1e31463cf7c3913b905cc3484e97988f6bd042bc

Observation e42bb48c-b815-47bc-a373-7db31c3079ff · inbound

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning cites this paper.

MathVis-Fine: Aligning Visual Supervision with Necessity via Progressive Dependency-Guided Training for Multimodal Mathematical Reasoning EchoSight: Advancing Visual-Language Models with Wiki Knowledge

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-07-03T20:18:57.617210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-27T01:23:40.564561Z digest=sha256:0ef422f73e5948bcb16e27790d0a625beaed412383c60c50ecdb0b96a805ff97

Observation 61ffe059-d1ae-4d04-950f-47add1214d65 · inbound

Ground Then Rank: Revisiting Knowledge-Based VQA with Training-Free Entity Identification cites this paper.

Ground Then Rank: Revisiting Knowledge-Based VQA with Training-Free Entity Identification EchoSight: Advancing Visual-Language Models with Wiki Knowledge

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:59:46.888382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-06-26T08:12:14.829556Z digest=sha256:a3d66013c7ce4902facb20f74eb2d2cedc772f0ae7592399b3c3c8d218560154

Observation 27be4eae-a4ff-43a8-8cbe-3dece275ac9f · inbound

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG cites this paper.

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG EchoSight: Advancing Visual-Language Models with Wiki Knowledge

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-02T10:20:50.309084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:20:50.309084Z digest=sha256:2508b9766256bd86598a9128f0e6e059bba72b83d76a8aa5cf07a5bb45a3d257

Observation d3e5af33-0afa-4b71-ba0c-f91d25a5e579 · inbound

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering cites this paper.

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering EchoSight: Advancing Visual-Language Models with Wiki Knowledge

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T00:31:17.396355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:31:17.396355Z digest=sha256:af0f5d037c974d9476a858edafc30a78b277a9035fb5ade34c00142ccfcb781e

Observation bd2f5435-1540-49da-8431-66d651fd345d · inbound

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering cites this paper.

UniHEAR: Unified Heterogeneous-Source Attentive Retrieval for Knowledge-Based Visual Question Answering EchoSight: Advancing Visual-Language Models with Wiki Knowledge

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T04:19:53.199025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:19:53.199025Z digest=sha256:bfbf47ce27b0c4b25d5d40167693c5bf34c67470f41ac54200124ce11d5a5c79