Pith. sign in

Paper Citation Record · LEDGER

MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models

As of 8 August 2026, this Paper Citation Record lists 4 of 4 outbound references and 0 inbound Pith citation observations for arXiv:2508.13938.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.13938 v1

Coverage vector

measured 4 of 4 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T18:53:05.214761Z

measured 4 of 4 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

4 of 4 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7464e93f-84fc-44e9-a514-0b3aa0c65612 · outbound

This paper cites PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models.

MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models PRMBench: A Fine-grained and Challenging Benchmark for Process-Level Reward Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T18:53:05.114132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:53:05.114132Z digest=sha256:ea1e48976e87ba3f321ae241ee7a5811d184971fb2ae8c3b781143ca8a4f55a0

Observation dac95634-5e48-4fae-897d-b95e50e70bcc · outbound

This paper cites Evaluating the Performance of Large Language Models on GAOKAO Benchmark.

MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-05T18:53:05.214761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:53:05.214761Z digest=sha256:f8162b6b2e8a5a93dd77fa1bd8c6450c6ae6bfe77cb54cfcc2fe4749a3604ad5

Observation 84f8d49f-094c-4fea-85e1-5baa35d2c0d5 · outbound

This paper cites A Survey on LLM-as-a-Judge.

MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models A Survey on LLM-as-a-Judge

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-05T18:53:04.972885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:53:04.972885Z digest=sha256:6a7686da4548d933370ebd510456f77c22e67310940f58f0a09041f6057fdabd

Observation e2498c15-1ae6-4456-91ac-4b9828478eed · outbound

This paper cites R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation.

MME-SCI: A Comprehensive and Challenging Science Benchmark for Multimodal Large Language Models R-Bench: Graduate-level Multi-disciplinary Benchmarks for LLM & MLLM Complex Reasoning Evaluation

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-05T18:53:05.031094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:53:05.031094Z digest=sha256:f0de154ea87fac5000aef917aaacde097ee84eba8871c3718ddf791a225c3640

Pith citing papers

No inbound Pith citation observations are available.