Pith. sign in

Paper Citation Record · LEDGER

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models

As of 11 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 0 inbound Pith citation observations for arXiv:2501.12418.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.12418 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T18:16:25.745948Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

43 of 43 outbound references displayed

  • verified exact3
  • verified fuzzy0
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d3f1bcd6-c70b-4d87-898f-ed34590ff693 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.478505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.478505Z digest=sha256:7c7d3e38670f19aaee3067eca19cd4d0ec0cb8a98a5a2797e94a9ff28480f534

Observation 78ad0596-1cf7-42be-9a48-27e086ca04b9 · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.489908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.489908Z digest=sha256:07936c521f54bc12528966e50daf4824a8f3d17232dc09ff53e279ac71d6b2e7

Observation fd44d442-f1f2-4410-9704-411d6df1e24c · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.855782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.500593Z digest=sha256:d9220c60e9bc719016dae8659c26b0b067a2fd6c10aa5eeec84680a574aca1d0

Observation 5d661999-f36f-4037-b99f-76c5e0f1400f · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.836411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.506459Z digest=sha256:2a3cb0f4ffd64334b88d4610498b64a364334485d2e8a05ce0c3c5c10a5d871d

Observation 6cfd4a4f-804c-41d2-a245-989a647ecb83 · outbound

This paper cites Learning to Plan and Generate Text with Citations.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Learning to Plan and Generate Text with Citations

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.512501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.512501Z digest=sha256:1686c19212101dace8a8823f030f05e50ef00502399e51587033806fa1eed613

Observation 3ecb4581-1e0e-4db6-9bf9-3adb572429e4 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.518335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.518335Z digest=sha256:4ec6338a86942ef719fefbe57a1642f23bcdf186a9a0d43ce956cc379df32250

Observation ca1e3466-f52a-470b-b095-c64f46e697cc · outbound

This paper cites Enabling Large Language Models to Generate Text with Citations.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Enabling Large Language Models to Generate Text with Citations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.525389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.525389Z digest=sha256:6614f62a70afa6710e53a4ee95c264b3bc40320a146b4c98a1d407943101dbf6

Observation b3931a67-f4e3-4821-96e8-8b3a7ffa821d · outbound

This paper cites Retrieval-Augmented Generation for Large Language Models: A Survey.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Retrieval-Augmented Generation for Large Language Models: A Survey

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.531582Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.531582Z digest=sha256:f82519bb6bfecbb677df04379a19eb92e997ad07e32999d94029c8cdaa5f529d

Observation cd25219e-df64-40b9-b348-54153b77cf7f · outbound

This paper cites A Comprehensive Survey of Retrieval-Augmented Generation (RAG): Evolution, Current Landscape and Future Directions.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models A Comprehensive Survey of Retrieval-Augmented Generation (RAG): Evolution, Current Landscape and Future Directions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.537683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.537683Z digest=sha256:5d5a104712d1d95c7b5ec67df664b8b54dae71d209a86e7c0fd452d156267566

Observation 2440a93c-3f7d-491a-a1f3-68822f06b681 · outbound

This paper cites Towards Verifiable Text Generation with Symbolic References.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Towards Verifiable Text Generation with Symbolic References

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.543464Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.543464Z digest=sha256:76112c14785ac8b10fc305d32c7afce31589a78e36f2cea429f3d9f7f3a43ba8

Observation 61773371-c9bc-4a60-bf0b-9c0035504e97 · outbound

This paper cites Training Language Models to Generate Text with Citations via Fine-grained Rewards.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Training Language Models to Generate Text with Citations via Fine-grained Rewards

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.549812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.549812Z digest=sha256:ba63c140eea000cf1f6fd693ce2960986aac060a94026fda30674fbd60c63ac0

Observation c2001045-7c55-4b96-a412-b82bd13399a4 · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 12

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.819478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.555646Z digest=sha256:cd658813c70c8a0883c1081b09fbc3d7c38c13c5d776e856c978ade9df2e608d

Observation 3942e530-2209-421a-bc6d-ed3cbdb689c9 · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.800976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.561384Z digest=sha256:d39cddcfb54f62b7606db0f4b97378417e33759710579b91166a46d528d4c4ec

Observation c9148c76-5f53-48ab-a03c-9e23f89f2945 · outbound

This paper cites GPT-4o System Card.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models GPT-4o System Card

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.567338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.567338Z digest=sha256:0c761abf079f9942056815b0f51bd36f59470034a67679c8f95a105aee0dce79

Observation 1b394101-d2d2-4ddb-bec2-f56c560aded5 · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.779772Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.573600Z digest=sha256:a7f719a3ed9aa20a359263ee04dc459d2999e3176d01e4c4c551cb37244319aa

Observation 089bacb6-b49e-45d5-9fcb-a5911443e629 · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.755031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.579592Z digest=sha256:17c174b2625f144dd24a0355a67a126a185727b56030cd91e8421db85dc7b19f

Observation a8484ff9-f8e0-4ab7-a617-943a4ead9f4a · outbound

This paper cites Towards Reliable and Fluent Large Language Models: Incorporating Feedback Learning Loops in QA Systems.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Towards Reliable and Fluent Large Language Models: Incorporating Feedback Learning Loops in QA Systems

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-10T18:16:26.120152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.585379Z digest=sha256:064b8074b1ca803538da123e61d9de728c1e6e17112d3bfd60574b75d62e9cca

Observation 649546f3-a34b-40fe-ac77-b628926cbb85 · outbound

This paper cites Improving Attributed Text Generation of Large Language Models via Preference Learning.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Improving Attributed Text Generation of Large Language Models via Preference Learning

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-10T18:16:26.089678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.591478Z digest=sha256:a801696a816c5bf00f7890b6fe377be472db16879795c2135abe65068ee4d62e

Observation 97b966fe-596c-4582-a2d5-ff7f850a019a · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.734247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.598412Z digest=sha256:1b898a542f36c287848442f16fd0ce4b7230684086a1493796600635cdac4fe2

Observation d097dc70-7d00-4245-b8a4-ab888f08bf8f · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.714210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.604077Z digest=sha256:94ed632f5f9aadd26b55bedf9547e2864c8a165a42be9a7b7cb67b249c7dac7f

Observation 4799352b-a7d1-4f1b-b085-ad4609652a67 · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.690333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.610224Z digest=sha256:5a61ac938d87ae4a78d9ef8cd5c63ed9b0a3c40c47c13ddfb33e5b8b3b8b2d77

Observation fc911385-58c2-4d29-8960-28a2efed4a5f · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.668979Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.615354Z digest=sha256:ab4bccf0538335aa3ef0f133e45738b7e7e4459d0f8faf84282721e56120335d

Observation 7957f7e0-13a0-49d0-bc50-6a40f7e65fa1 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.621128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.621128Z digest=sha256:03b3e0b9f8595d501db60adf7d972b70024f8667601301da6b1d2defb61fcb1b

Observation 0ba157aa-1050-4032-b4de-ad8a13b995d8 · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.644069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.627439Z digest=sha256:3ac4630cfbb34db92fa5af33137eef5178fb6482dbfb5b5750b5ce66cf5d23ce

Observation 5484146a-643a-449c-898a-9a95aadc7e39 · outbound

This paper cites DyKnow: Dynamically Verifying Time-Sensitive Factual Knowledge in LLMs.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models DyKnow: Dynamically Verifying Time-Sensitive Factual Knowledge in LLMs

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.633276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.633276Z digest=sha256:7eff74e606e74471538a2fbb201a4308d58a518b75c08f52ee590d69c7181d0d

Observation 39cb3076-c777-471f-b7a5-f00248c93f6f · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.620189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.640027Z digest=sha256:3c20da92022a4d46452574e2e37a68f77d79cedb1919d377dbe4e44dd63dc34f

Observation 87af93bd-6f40-4b73-bba1-372d988aa146 · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.587531Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.646855Z digest=sha256:606207eddb4b6dce51e8ae05e329761651cb986a16760efcf1901296a9d6a7ad

Observation 7e2fdcbe-a844-4989-97bc-092846acb64e · outbound

This paper cites Synchromesh: Reliable code generation from pre-trained language models.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Synchromesh: Reliable code generation from pre-trained language models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.653670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.653670Z digest=sha256:4dd29bdda859dc22565e4761237e3170eb1256b5e5f405471f41a6be06d302dd

Observation 58397e66-d170-454b-926a-d1997761ea6a · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.560725Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.659144Z digest=sha256:1f7f0d6d83da2e15c85e70eecef0e79be5de54b193012eb30eef8dc3d6b4eda3

Observation e253b605-890b-4671-9ba9-84ef2227c51d · outbound

This paper cites Citekit: A Modular Toolkit for Large Language Model Citation Generation.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Citekit: A Modular Toolkit for Large Language Model Citation Generation

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-08-10T18:16:25.974877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.665773Z digest=sha256:5ace4aca06c87de986c80d4901a39088e6599cff9b6e8fec98492e66afb3a19b

Observation 6b93e68f-306c-4521-9ca8-fa02621dbc03 · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.532493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.671292Z digest=sha256:fd862130c88f85d7b1796c29ed753210f689245ae8bf2e39626fc89233e3c686

Observation b53eb951-6531-4e0b-95ed-faa89f2851c2 · outbound

This paper cites A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models A Comprehensive Survey of Hallucination Mitigation Techniques in Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.676939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.676939Z digest=sha256:d3a5b38e7c9f99cb22d9208752f034b8545c50b8eece70f75c46df155bd2d834

Observation 344f8f6f-a051-4bd0-83d1-85b5b879be5f · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.511687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.682144Z digest=sha256:7ecf22b00837c7fc6bee725f32c3087ae04552c54692a0231024dbd731ce30fd

Observation 019bafe7-4c6c-4911-b4a1-5ae5d81c959a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.687214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.687214Z digest=sha256:8c42fb765011eb63e5805ec04a83cc4045314361abe6cac3438acf40a097f17c

Observation e402273f-a06f-44c4-bf4d-67f400b16324 · outbound

This paper cites Hallucination is Inevitable: An Innate Limitation of Large Language Models.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Hallucination is Inevitable: An Innate Limitation of Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.692816Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.692816Z digest=sha256:aa30d3dbd7f5ecd7d8e457b4b2b6ddcf63cb46b3eaec7cfd993a50c7ff345602

Observation 241b834d-f3d9-4120-84ee-13065d37bbd8 · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.490873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.698440Z digest=sha256:f6d961946011e9e96a3227741e3f88a7778577d93780bf32002478f43822e765

Observation c812ddfb-7440-4723-a8db-be9867f96b6b · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.704108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.704108Z digest=sha256:4480ee6bae88d139a0d6d1c71555cf2df8574d27693a37bafc8befdca72dbdf1

Observation 4fc75928-77d6-4d36-98e4-d9bab9f44392 · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-10T18:16:26.469151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-08-10T18:16:25.710476Z digest=sha256:a62a376060c7e0ccabcd3154d6019b705080a1729e6f6b5cb8f6754bc8d68572

Observation 1bfce028-bca9-48bf-b799-e1a9dbccfb10 · outbound

This paper cites LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QA.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models LongCite: Enabling LLMs to Generate Fine-grained Citations in Long-context QA

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.717571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.717571Z digest=sha256:99ce5628a35c5f0765c38d67b5b0924325ef5865d216759d1b679fdd7e99bd10

Observation fe6e9f42-aa12-431a-8146-69411c6c6d09 · outbound

This paper cites Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Verifiable by Design: Aligning Language Models to Quote from Pre-Training Data

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.725604Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.725604Z digest=sha256:f6276436693cee1118e4368fc6f5020af44e2310b3b665cda7245ec812050aa7

Observation 3c686bcb-82a8-49ce-9c05-5c9128ca258b · outbound

This paper cites an unresolved cited work.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.731564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.731564Z digest=sha256:8bb3d034c3f2e19efbab92c4c2a106af9ee5fe8594d293e73d7a7ce303cb6ad0

Observation 2f5fddca-8002-424b-8c90-b1eb12485805 · outbound

This paper cites online" 'onlinestring :=.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models online" 'onlinestring :=

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.737702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.737702Z digest=sha256:6bc6c9ffd698f4a1f00c13acd9a5dec4b55d694d3d05d990fce4156ce7f7ee87

Observation 30148be3-d543-49a2-8f81-e7d975819348 · outbound

This paper cites write newline.

ImageRef-VL: Enabling Contextual Image Referencing in Vision-Language Models write newline

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T18:16:25.745948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T18:16:25.745948Z digest=sha256:eaf973cab0621e801510e216b26fd678ce0b3b21d93da51c69b59340bef9d526

Pith citing papers

No inbound Pith citation observations are available.