Pith. sign in

Paper Citation Record · LEDGER

J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM

As of 15 August 2026, this Paper Citation Record lists 21 of 21 outbound references and 2 inbound Pith citation observations for arXiv:2412.15574.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.15574 v1

Coverage vector

measured 21 of 21 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T11:19:55.660432Z

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-06-30T23:07:21.558834Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-01T13:35:46.054815Z

Reference resolution

21 of 21 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d537423b-c6ac-41e1-acf4-6d91d72cf7c2 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T11:19:55.539076Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:19:55.539076Z digest=sha256:dc0a8785d87a4167824808211268d8ceabb645d6973798b123980132fa859ec7

Observation 92c676d1-9330-41f3-af57-c0829452c97a · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T11:19:55.544988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:19:55.544988Z digest=sha256:f13717939f903c910a36ae3b106d454ca92db4c1d3cb989cbdf549de44204ed8

Observation 0dbc1604-abb1-492b-9577-fe133f3d88c7 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T11:19:55.550086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:19:55.550086Z digest=sha256:e7b57f2965c153ca358186a42e5220bab8ce70be1fa219b28c531378329b9ee0

Observation a11d6dd5-d448-46ba-a9c5-5c2ec1d2f13f · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM Improved Baselines with Visual Instruction Tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T11:19:55.555858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:19:55.555858Z digest=sha256:4b57c8eb54d506fa2229d038c13392b3f03f723aa89f380efbccb28b5002fe8c

Observation e2b7b835-0f67-4478-9862-97133b946902 · outbound

This paper cites VILA: On Pre-training for Visual Language Models.

J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM VILA: On Pre-training for Visual Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T11:19:55.561562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:19:55.561562Z digest=sha256:de7ac91dfdcf0f311ee07c670c1fcf5b94e7361e2ecfc054bb7843caf44d3799

Observation b164f4d5-3d5e-4551-8311-461005726f74 · outbound

This paper cites Towards VQA Models That Can Read.

J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM Towards VQA Models That Can Read

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T11:19:55.566918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:19:55.566918Z digest=sha256:0c74eee7c18454ae05b9d02b8b213b3772eb71c934dfa0d8778ecc68f0f7065d

Observation 9c804421-ee55-4c21-b32e-d065cf3b9bf9 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T11:19:55.572606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:19:55.572606Z digest=sha256:427c8c8ceb9b92283eb7b2d51e5eaed698b1f15ff0cd5d8648e61173180f91ed

Observation 52090fd6-6981-47d3-977a-0d6b209b1977 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T11:19:55.578987Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:19:55.578987Z digest=sha256:193b91c35c2944b2f26217bf28902049fcea62e55a3f384f1f02fba64d172faa

Observation 37bb9b46-26fb-4f5a-b339-e0e82f0ed43a · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T11:19:55.584128Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:19:55.584128Z digest=sha256:4f393ddc3cc505864ea1b6207d58e89c04f684171c01a8f7a3f464d898993dbb

Observation f1f298b3-084b-498a-ad5d-cd3856b114c3 · outbound

This paper cites GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering.

J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM GQA: A New Dataset for Real-World Visual Reasoning and Compositional Question Answering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T11:19:55.590272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:19:55.590272Z digest=sha256:49c5d55c34fac93835e654f7f05a56eeb31194d370035aa1164fa65f0afcf070

Observation e13a3f53-9826-477a-bb6f-f16d829690da · outbound

This paper cites LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models.

J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM LVLM-eHub: A Comprehensive Evaluation Benchmark for Large Vision-Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T11:19:55.596858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:19:55.596858Z digest=sha256:eee739fadf198dd9e5adcbc037cee7008542983e1fe83c8986fd88544a2bcff9

Observation 8171a99e-f033-4905-880f-20e6cc1afa22 · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T11:19:55.602716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:19:55.602716Z digest=sha256:cb387b454aebf567b6d55b845b07246f8bde994038b913b2b98841b96c57aa26

Observation ceed74dc-d043-49e9-82ab-63a82947bb50 · outbound

This paper cites GAIA: a benchmark for General AI Assistants.

J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM GAIA: a benchmark for General AI Assistants

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T11:19:55.610958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:19:55.610958Z digest=sha256:173cd39124ff795e431451fe29484d7cbba0b008e8a7072ced440383e5b63670

Observation c43b7dc4-e864-4441-90c3-98f940ce2eef · outbound

This paper cites LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and Benchmark.

J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM LAMM: Language-Assisted Multi-Modal Instruction-Tuning Dataset, Framework, and Benchmark

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T11:19:55.616921Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:19:55.616921Z digest=sha256:bdd5f5905ea1521a58fabb543294b4447f3160184e861617aef48b34cc906cd9

Observation 7b93a2db-f74a-4028-ae7a-d3742f483373 · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T11:19:55.624220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:19:55.624220Z digest=sha256:b0d4e5f6e311c428335878deadb71afb73d0f5260b8c26077c1bcece356229fe

Observation 298c5c78-70e8-464c-aa83-3959799fe7ca · outbound

This paper cites VATEX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research.

J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM VATEX: A Large-Scale, High-Quality Multilingual Dataset for Video-and-Language Research

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T11:19:55.631307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:19:55.631307Z digest=sha256:8c5d4348a62772195034e8ec5300a053d5e3465b8ce701c0f9d42ac7d4c26b74

Observation 3630f9cd-e051-4639-8f7b-db90f12e0eb9 · outbound

This paper cites Perception Test: A Diagnostic Benchmark for Multimodal Video Models.

J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM Perception Test: A Diagnostic Benchmark for Multimodal Video Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-11T11:19:55.636847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:19:55.636847Z digest=sha256:0b0b60dc6053c6dc4a744efdae06b74504c4d69ef01d0b0b3e702cd18427e7e0

Observation a3f8adc9-9e04-453b-bbad-62d03cbc9531 · outbound

This paper cites Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering.

J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T11:19:55.643473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:19:55.643473Z digest=sha256:ee2d31f757de78edfee4cc08d8c01bf2cd101bf8f96805fe04fcc1c0e32bbb19

Observation 857cc4c2-e511-49b9-a51e-46bb6b214275 · outbound

This paper cites Heron-Bench: A Benchmark for Evaluating Vision Language Models in Japanese.

J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM Heron-Bench: A Benchmark for Evaluating Vision Language Models in Japanese

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T11:19:55.648610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:19:55.648610Z digest=sha256:2e938d4deeb5005278cee83e1f39b2a515b7e3b8c14df09f9c7bee4e9b11b34a

Observation b4c762fd-04a0-46d5-981f-c0bcca5d4c38 · outbound

This paper cites Evolutionary Optimization of Model Merging Recipes.

J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM Evolutionary Optimization of Model Merging Recipes

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T11:19:55.653912Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:19:55.653912Z digest=sha256:43fc95addc06159f3e0c8caa64491685d689a4453f2f186fd623577bce2a50eb

Observation 7ba94a51-1a69-444b-83cd-954f656fe79d · outbound

This paper cites JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation.

J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM JMMMU: A Japanese Massive Multi-discipline Multimodal Understanding Benchmark for Culture-aware Evaluation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T11:19:55.660432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T11:19:55.660432Z digest=sha256:0a9e6103532105432874c368d32dc85de4b1d3751a3c1c2b25537f38e4d16463

Pith citing papers

Observation 5295aa94-fba0-4600-a95d-3766c4bb0ecb · inbound

Earth Science Foundation Models: From Perception to Reasoning and Discovery cites this paper.

Earth Science Foundation Models: From Perception to Reasoning and Discovery J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM

Reference 193

Resolution
verified exact
arxiv_id, observed 2026-05-14T22:08:03.377193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-14T22:07:40.242567Z digest=sha256:2fa54416e6cfcb6fd4916ce9aeac850dc5927d5cc8738fc7dfe8c57f99850c98

Observation 88efef0f-03fa-47f4-ab0d-13a387965924 · inbound

Earth Science Foundation Models: From Perception to Reasoning and Discovery cites this paper.

Earth Science Foundation Models: From Perception to Reasoning and Discovery J-EDI QA: Benchmark for deep-sea organism-specific multimodal LLM

Reference 193

Resolution
verified exact
arxiv_id, observed 2026-07-01T13:35:46.056613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T23:07:21.558834Z digest=sha256:7ef1e10c0fba668c1bfad523a38a9186ab7b4508dcb90a48756b22b100fa4b25