Pith. sign in

Paper Citation Record · LEDGER

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models

As of 4 August 2026, this Paper Citation Record lists 43 of 43 outbound references and 3 inbound Pith citation observations for arXiv:2604.04172.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.04172 v1

Coverage vector

measured 43 of 43 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-13T16:45:17.022062Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T16:12:23.659005Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-30T12:16:12.909640Z

Reference resolution

43 of 43 outbound references displayed

  • verified exact21
  • verified fuzzy2
  • unresolved19
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3aad4050-c983-46f0-9cdb-2582f8018c3e · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.872301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:912c0f57f8aeac4f94a91c18f306885defe5a1fc2ea17881d92900c92ab624fd

Observation 29923d39-5047-4c56-a629-c4fddb464879 · outbound

This paper cites SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models SridBench: Benchmark of Scientific Research Illustration Drawing of Image Generation Model

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.127435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:ea939ae16ef352eb9903c9806491784e9ec69b9c842e152a27500019169169aa

Observation 637a6ca9-398d-4514-b9a2-aa414119ae6e · outbound

This paper cites Figure Captioning with Reasoning and Sequence-Level Training.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Figure Captioning with Reasoning and Sequence-Level Training

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.153640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:21cf178182e2260bb1d02d20185cce606c2df7edd0d490427989c9159bd01bec

Observation 3c45bc01-ce71-4cf9-90a9-a5378ebe2afe · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.856423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:eb037e0e6a27dc3450c739e6116ade4e99e3d61b2843fa07e1472fbfc2e70d33

Observation 4906e079-76c2-4fdb-bc59-428e16635311 · outbound

This paper cites TIAM -- A Metric for Evaluating Alignment in Text-to-Image Generation.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models TIAM -- A Metric for Evaluating Alignment in Text-to-Image Generation

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.134881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:cc2c01aad6f83cc1455bf36a4b051bca135adebceff9b3246a3828dfddd81170

Observation 9ac896bd-787d-49b9-8aa1-0384152d2871 · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.853550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:8ad3b73e551c9a44a1d2b3d0ef48676fc62cac6ee0ce7dd497d03dd0a5657cc4

Observation aa87a366-a820-4db3-9be1-bdd36d720a5b · outbound

This paper cites SciCap: Generating Captions for Scientific Figures.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models SciCap: Generating Captions for Scientific Figures

Reference 7

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.157019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:26c2c594521f6d2f962646f2d74325742d8338b2a1f208d4052e665719d76e47

Observation 1a8df55f-d682-43a3-9150-19411977c5c1 · outbound

This paper cites mPLUG-PaperOwl: Scientific Diagram Analysis with the Multimodal Large Language Model.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models mPLUG-PaperOwl: Scientific Diagram Analysis with the Multimodal Large Language Model

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.131525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:a6afe613f86d12b865a5a006ba23df6d1e4051897cfc5f0e60c6ac07fc3f7ce2

Observation 1a663a3f-b1ac-4b2e-aa3d-8ca67ed77482 · outbound

This paper cites GPT-4o System Card.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models GPT-4o System Card

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-13T16:48:03.142195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:ba484643c72f4c2a8df572e75a894cc35e77f6dc690ad2c966c9c47e4dcf2337

Observation 6ea16a10-3e8f-4df2-98b2-5816654ae642 · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.859100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:67818ce0b8cab16e7dec609b4f0a76ea3a73ad8ff029f576883f08dd7a1749d9

Observation f9e1b439-4b0a-440e-8fe7-509dc4e9a0e0 · outbound

This paper cites DVQA: Understanding Data Visualizations via Question Answering.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models DVQA: Understanding Data Visualizations via Question Answering

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.145919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:4858918d01271e1a0b2107dfe04c4e675e1f8447949b5a0727bc764addfbc391

Observation 57e0cd0d-85d4-4dc5-8cb7-eb0fa457a6e5 · outbound

This paper cites FigureQA: An Annotated Figure Dataset for Visual Reasoning.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models FigureQA: An Annotated Figure Dataset for Visual Reasoning

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.120684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:4cfc65d0c55e0cf9666d5ff87385e027438ba59ee1b2358d6146e5dd6088f9be

Observation 36d86ebd-6a76-4eb8-bf23-5d56f63093bb · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.894565Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:3f2b599feccd24d5fc980488ec3f5b5e9b78b83c61688cd51d2fde38ab871025

Observation 5ab9584a-fef9-4d1f-ba7b-4957d8e3abe5 · outbound

This paper cites Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Multimodal ArXiv: A Dataset for Improving Scientific Comprehension of Large Vision-Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.177074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:a34d9db18078285e3913dee62a8032a0a962cf0f26b6b01ae68071ade6bdb8f2

Observation 4241a15d-582b-4c18-aafa-a8dd0b191c77 · outbound

This paper cites MMSci: A Dataset for Graduate-Level Multi-Discipline Multimodal Scientific Understanding.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models MMSci: A Dataset for Graduate-Level Multi-Discipline Multimodal Scientific Understanding

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.187477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:cf16912034dd5f185fd6b987db09dccc5b4e306e91e6c75ae94fe17613ee9b73

Observation 728949d0-95c7-4b12-8945-2e66611f446e · outbound

This paper cites Aligning Large Language Models with Human Preferences through Representation Engineering.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Aligning Large Language Models with Human Preferences through Representation Engineering

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.173387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:c498576dc0af5086a5889496f156a3dd87528a05a5a5b8a29a3e00d43a1bf197

Observation a68b9895-21cb-4fc4-bbda-60bd4415c2d0 · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.889617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:7fdc6601ac943fbf76e2b7506f16cd87ffce6d610462b4d53eca8cfe3f761cf2

Observation c1c364f7-8a10-4067-b9ba-964d218cacd4 · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.904571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:a9cb0228ef4e1d609d295880aa6ddc3d6dc5a78a937f1340111d7dd350f182da

Observation bd49d785-8948-4577-a935-5beb3f527ee1 · outbound

This paper cites GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-13T16:48:03.160299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:850390e6a7bff2160c882e7b4d73f0588260da1a5faea862bb6ddc631786145c

Observation f1cfd909-86c2-4065-8e3c-e448d9621a5c · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models DINOv2: Learning Robust Visual Features without Supervision

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-13T16:48:03.169996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:e4a4a916eb4321aa1cfd7ff766243fe03d3b46e15cc7336d330f05e57baf5248

Observation 8d6ccf26-197f-4b07-a06d-602920319e59 · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.902116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:b002d757faeb1175dc7cf4fd2da515ce6883330ae07e1a7ecaf21fe182d7b9e7

Observation 670a2808-f012-4078-851e-45ada37ec5d8 · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.897160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:a178bfb53101ed91bf3668b979fa5674fd57835bcb94644be5c3f74119702c18

Observation ffe3a154-5aaf-4027-80b5-847567de1cfd · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.899686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:bb3d32008487aec64b3d909699f789fa5ddaf03138bb0bdb8d58b85287154bd2

Observation 6a1702af-bec8-460d-b80e-41888bc8672b · outbound

This paper cites SciFIBench: Benchmarking Large Multimodal Models for Scientific Figure Interpretation.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models SciFIBench: Benchmarking Large Multimodal Models for Scientific Figure Interpretation

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.163678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:0cf66fc7b652fa5d76d97681d56331983b772b6b57da746f99ac4ca93e298b35

Observation 99457d57-9e9b-488e-9493-2aa71d9cee3b · outbound

This paper cites StarVector: Generating Scalable Vector Graphics Code from Images and Text.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models StarVector: Generating Scalable Vector Graphics Code from Images and Text

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T16:48:03.166902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:3c0ee4e2e4383c93e586717d0ef1f42a0ceef3f9caa021cd695f2fdb2277ff0b

Observation 87f3003e-6cf7-43a5-9a1f-829d500ca49d · outbound

This paper cites OCR-VQGAN: Taming Text-within-Image Generation.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models OCR-VQGAN: Taming Text-within-Image Generation

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.180740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:91273f1bb85d6063e2fbc876ab77ea61bab949bf8f33e79a7070cd4d319a1105

Observation a2666c8e-65e4-4d7e-bc8c-bf595c51b326 · outbound

This paper cites FigGen: Text to Scientific Figure Generation.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models FigGen: Text to Scientific Figure Generation

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.184022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:04f9065ade74a3504ba141b0e40935c8d2cbdd74e3315a20e3e9e807d34ac5ed

Observation 8fcb09b1-f074-4784-a7d7-675820ef6211 · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.891982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:6ab4c8e5a295fe93d30e2b5911088c29235e1e684f8aef5a28d6bfd5975f34a6

Observation 77b4005e-1dd4-46f5-b1b1-8be724f3edf7 · outbound

This paper cites Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Photorealistic Text-to-Image Diffusion Models with Deep Language Understanding

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-13T16:48:03.149583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:99ab30039975e724120584d5b528f9a5aee5fa265096712edac9a7986887681c

Observation a725aeec-e143-4fca-a19f-070a6fa01b0b · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.867052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:0443caadc2bb80c60267e0fd43413cdcd072d5ccd168c77385fe01d5114469fa

Observation 8aeb5649-b9f0-406c-a534-cf6ff8eb1f1b · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.861534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:59256709a1bf483fbcca520ad4504cffa72ca8c26c5a8f8909f506f0913faf50

Observation 841567f4-5b03-4a87-a5bf-e8314b9bf0e4 · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.884429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:0810a08b8fae5dc7b7e0ac16d56a07894acd8a7cd17591eac874ce0a12c21a2e

Observation a2252e7b-776c-41c1-a6b9-0db6ed906358 · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.875150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:75179129057c9a1ae8693ca2fb81d8b9b2b25b638ec30c5f16c851be7f008a01

Observation 6d1625b7-d971-470e-97a3-00b0d46a3d21 · outbound

This paper cites CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models CharXiv: Charting Gaps in Realistic Chart Understanding in Multimodal LLMs

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.138494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:15ddcf5e418cdeec60bf5214bbab99c255119ed226178456b04234de7bba9f61

Observation e977d784-24ad-476a-aef0-f7058ec1b1fd · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.887186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:ce32364e943c9db7490ffcc3750270dc575152a60a1694cd31a7e6c8769b7086

Observation 2098109e-c795-4eb8-a2f0-e990eeeb7b7c · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.869763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:eb4f0a23cc9d48ba8a22531324ce8250d2397c07072f5e479bfdcba6cd7cca4c

Observation 73f37f0d-40e9-4022-b32b-978c601e84b5 · outbound

This paper cites SciCap+: A Knowledge Augmented Dataset to Study the Challenges of Scientific Figure Captioning.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models SciCap+: A Knowledge Augmented Dataset to Study the Challenges of Scientific Figure Captioning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.123892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:4ef5cd470350edb7289b168e265c0333f2929b26dedef79c1286b0ea9cbbd248

Observation 0b982cf7-55e5-4ecf-82bd-c1a2c83e607e · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.864256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:6fee593dc2d744618ad2a0a7f197542a34f74b23389b4d424932880b5809debc

Observation d9f239f0-318a-4107-9de3-8ec44bafc90b · outbound

This paper cites an unresolved cited work.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-05-14T05:58:14.877915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:25b5c35ce944a237b56603802bbcc6b35aa052b2f4fcfb6d79fdee66857f9db6

Observation 83ad4aca-ed07-4ca4-be35-c520170bfa20 · outbound

This paper cites You are an expert on translating academic writing to visual specification.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models You are an expert on translating academic writing to visual specification

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.115963Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:50e92443924f8f9d7e1231bf7f4559e1070557ef2dd96c88c66c5ce3740001e9

Observation 2487533b-bfc4-4ebc-8995-757631700859 · outbound

This paper cites Autofigure: Generating and refining publication-ready scientific illustrations.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models Autofigure: Generating and refining publication-ready scientific illustrations

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-13T16:48:03.111073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:5d19b3b9e7b1e3b9de143f1de54966c029d8ea202ade59feca59a3832c01caa3

Observation aaddee79-f288-4f75-8b22-6b81650df147 · outbound

This paper cites online" 'onlinestring :=.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models online" 'onlinestring :=

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T05:58:14.881515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:a32968b4a3fd7e7ae0cf2aa5bff25f07d9935105e7d4470c81ecf9b0cd6cb0bf

Observation f235db4d-e6cf-4d13-9104-83fca1482a0d · outbound

This paper cites write newline.

GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models write newline

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-05-14T05:58:14.850483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-05-13T16:45:17.022062Z digest=sha256:afea4628cdc5535ad28ed82c8014031784f24e62fd1644ceffe9989aac1b41e0

Pith citing papers

Observation 079c2a9e-6acc-4a07-b591-945fb24f61b2 · inbound

SciForma: Structure-Faithful Generation of Scientific Diagrams cites this paper.

SciForma: Structure-Faithful Generation of Scientific Diagrams GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-01T16:12:23.659005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T16:12:23.659005Z digest=sha256:271e74e9cfd68af7ab12c9e8980025d36be949d5f8c9bdda7e88310e49a98454

Observation 87fa53f8-e43c-4cdb-ae0d-e5e9fbe6fa81 · inbound

SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence cites this paper.

SciFigAlign: Scoring Scientific Figures by Fine-tuned Alignment of Visuals with Manuscript Evidence GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-30T12:14:15.738599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-30T12:14:15.738599Z digest=sha256:8a05cb8e91307b570fd3a701fc00b6fc881cb8d229c348a660c54b77bd7a0201

Observation 39b48dcb-1e7a-4f9f-9980-a454bbd1f682 · inbound

SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context cites this paper.

SciFigQual-Bench: A Benchmark for Scientific Figure Quality Assessment with Full-Manuscript Context GENFIG1: Visual Summaries of Scholarly Work as a Challenge for Vision-Language Models

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-30T11:41:21.736732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-04T06:34:03.388597+00:00.

source=arxiv_source observed=2026-07-30T11:37:27.095506Z digest=sha256:1bd52f5081255600207636338cd8ceb8a40e7fd5ff2d9bfc0e72dcb8449ea8a9