Pith. sign in

Paper Citation Record · LEDGER

OpenCompass: A Universal Evaluation Platform for Large Language Models

As of 19 August 2026, this Paper Citation Record lists 23 of 23 outbound references and 1 inbound Pith citation observation for arXiv:2605.19276.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.19276 v3

Coverage vector

measured 23 of 23 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-30T18:41:34.878542Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T01:59:49.787962Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

23 of 23 outbound references displayed

  • verified exact10
  • verified fuzzy11
  • unresolved0
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d592d2f6-3689-4cd5-a00e-b686cf03132e · outbound

This paper cites Longbench: A bilingual, multitask benchmark for long context understanding.

OpenCompass: A Universal Evaluation Platform for Large Language Models Longbench: A bilingual, multitask benchmark for long context understanding

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:34:32.842784Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:81f1786c4f1296f132d45ea06bbaca7464ef7da5d3cc4b51c843a49a1ee4593a

Observation 71843c89-91ee-46e1-b745-5a7c8035d78a · outbound

This paper cites ARC Prize 2024: Technical Report.

OpenCompass: A Universal Evaluation Platform for Large Language Models ARC Prize 2024: Technical Report

Reference 2

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:45:00.502117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:844c17fb2c566dac2247b2004733362727e4ce8712f9d76d7d089624b62843b4

Observation 75a1b657-6649-4020-9ce1-65c629aeccc8 · outbound

This paper cites Lmdeploy: A toolkit for compressing, deploying, and serving llm.https: //github.com/InternLM/lmdeploy.

OpenCompass: A Universal Evaluation Platform for Large Language Models Lmdeploy: A toolkit for compressing, deploying, and serving llm.https: //github.com/InternLM/lmdeploy

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:34:32.840887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:c34d4142fddad60206dbcd0c9272fd93dcd5caa1de35ff593b1110448812183e

Observation 0a0d9803-7ce2-4924-b3d8-ac51f9fb4e28 · outbound

This paper cites MMEngine: Openmmlab foundational library for training deep learning models.

OpenCompass: A Universal Evaluation Platform for Large Language Models MMEngine: Openmmlab foundational library for training deep learning models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:34:32.831285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:5f1b495502d2d9f0f292f24470d18a0edc9c80ebbd2bd3a99c89f789a5f5a156

Observation ae30e3c5-e26f-4805-b60b-d8e64ac5360a · outbound

This paper cites Physics: Benchmarking foundation models on university-level physics problem solving.

OpenCompass: A Universal Evaluation Platform for Large Language Models Physics: Benchmarking foundation models on university-level physics problem solving

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:34:32.838966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:82cf89c33de113a7252d09ea9c6a96a8a090233b40e35b0413bc35bb864356e8

Observation f71df103-5455-45a5-b596-6a14c78bacfb · outbound

This paper cites Measuring Massive Multitask Language Understanding.

OpenCompass: A Universal Evaluation Platform for Large Language Models Measuring Massive Multitask Language Understanding

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:45:00.508507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:852b39f1b1b3e156030c4b8e64ba9937271d738c97658b91dcd07fd76f38ce2d

Observation 70f6b27c-8ebd-45ff-9323-15ff9bc55762 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

OpenCompass: A Universal Evaluation Platform for Large Language Models Measuring mathematical problem solving with the math dataset

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:34:32.822755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:a41a97a8a0d68a60ee1c87e7027c3defcecf5d532a54fe43aa60809ad8df2193

Observation fd4e5296-2d2d-4cfc-a779-b68347ca344a · outbound

This paper cites RULER: What's the Real Context Size of Your Long-Context Language Models?.

OpenCompass: A Universal Evaluation Platform for Large Language Models RULER: What's the Real Context Size of Your Long-Context Language Models?

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:45:00.505348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:ee8f15bcc81e8013b6968b48c22d16be291ff2e6c05dd46cc352b7d7fd45e7f2

Observation 84eeaa5b-4d63-4434-92fd-42e42ae12c6d · outbound

This paper cites LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code.

OpenCompass: A Universal Evaluation Platform for Large Language Models LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:45:00.513123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:c9324ba7457e579e0f785b6cb49ebfee8f98e0bfc8c5d975ebf5df2ededb0cad

Observation d4a2c6f4-899c-40a1-9da6-64e7ebe254a9 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

OpenCompass: A Universal Evaluation Platform for Large Language Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:34:32.837020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:9831422bcd77a451fff0b7c066643bc0f9726bcf2433efc5fcc278bb2cf8ff14

Observation 8fa6604c-aee4-4044-a3d5-dc3a1c4a6f20 · outbound

This paper cites ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models.

OpenCompass: A Universal Evaluation Platform for Large Language Models ClimaQA: An Automated Evaluation Framework for Climate Question Answering Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-06-30T18:45:00.510578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:84145fc3c6305b780cf6878427cab974d1489b3954e9dd02c6c2ee7de5d5094c

Observation 9b8e936d-3c4c-4ce4-bb92-ef402ad1dd3c · outbound

This paper cites Humanity's Last Exam.

OpenCompass: A Universal Evaluation Platform for Large Language Models Humanity's Last Exam

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:45:00.519875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:5c8a440aaeef8b22d7e1402fe38801b0de78c8ec4d41db73d839f12a927f9653

Observation 04cfaca0-f851-48e4-a4ea-333a348a2415 · outbound

This paper cites Generalizing Verifiable Instruction Following.

OpenCompass: A Universal Evaluation Platform for Large Language Models Generalizing Verifiable Instruction Following

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:45:00.522603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:4ab361d37d180cfce55f47fe5334dc003a1a2ff4905b78db6ea9ff5f04fa9489

Observation c8e155bb-869b-40e6-a463-0ed768a3a19c · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

OpenCompass: A Universal Evaluation Platform for Large Language Models Gpqa: A graduate-level google-proof q&a benchmark

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:34:32.844694Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:d95dce05ad0450ffc262af4f29f2fd4d7e0daafab59ce1571c9e9966d79ead42

Observation 00e1dd4e-1e7b-4a83-a767-4a376a74bd9e · outbound

This paper cites Challenging big-bench tasks and whether chain-of-thought can solve them.

OpenCompass: A Universal Evaluation Platform for Large Language Models Challenging big-bench tasks and whether chain-of-thought can solve them

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:34:32.835142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:18a8ba35b3faa72835b16a0e7704a164b6bf412be0921abbb7840374901f9f45

Observation f8ec25ec-0c30-4789-b8ab-79d0c0058f19 · outbound

This paper cites Measuring short-form factuality in large language models.

OpenCompass: A Universal Evaluation Platform for Large Language Models Measuring short-form factuality in large language models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:45:00.504164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:97b4e341917fd7d924484ad8b9c55cedcbf8445d5bcbb3ada82c9531df514b84

Observation 69ed31f7-26d3-426e-a763-6ef3bf3c226b · outbound

This paper cites Openicl: An open-source framework for in-context learning.

OpenCompass: A Universal Evaluation Platform for Large Language Models Openicl: An open-source framework for in-context learning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:34:32.827268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:cab189a13f50041992db89ce44df6f24f63b1688066a2ced816437034577eff4

Observation 79721af4-5f87-4c63-b9df-a67a3c703a6a · outbound

This paper cites LlaSMol: Advancing Large Language Models for Chemistry with a Large-Scale, Comprehensive, High-Quality Instruction Tuning Dataset.

OpenCompass: A Universal Evaluation Platform for Large Language Models LlaSMol: Advancing Large Language Models for Chemistry with a Large-Scale, Comprehensive, High-Quality Instruction Tuning Dataset

Reference 19

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T18:45:00.507659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:87c50881a4f9790c75e21fd64e985068c64ddb82241afbe31a01bb515aa524da

Observation 4342b2be-848a-402f-a365-b2b45a0d70d9 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL).

OpenCompass: A Universal Evaluation Platform for Large Language Models Hellaswag: Can a machine really finish your sentence? InProceedings of the 61st Annual Meeting of the Association for Computational Linguistics (ACL)

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:34:32.824937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:f26fc61c28dfabe91c932c78d5bdffbb7edc97b632a3e92712dd0c3a2255ed90

Observation a0f67ffe-b832-4361-b884-a0bb14748663 · outbound

This paper cites P-mmeval: A parallel multilingual multitask benchmark for consistent evaluation of llms.

OpenCompass: A Universal Evaluation Platform for Large Language Models P-mmeval: A parallel multilingual multitask benchmark for consistent evaluation of llms

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-07-08T03:34:32.829275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:735c4151b857f426c917a484ac12138bb464dbb7d039a4d8d0f0861c2e4dcdef

Observation 0652a02d-29a8-4341-afe0-b8e1c143abd8 · outbound

This paper cites Instruction-Following Evaluation for Large Language Models.

OpenCompass: A Universal Evaluation Platform for Large Language Models Instruction-Following Evaluation for Large Language Models

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:45:00.527674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:cacf2dbcb4ac7ca89db8146fc6ebed3f4514f03258d81641e9a29c260d37cfda

Observation b523fd04-a9c8-4360-9e61-e2e1b58b08f1 · outbound

This paper cites BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions.

OpenCompass: A Universal Evaluation Platform for Large Language Models BigCodeBench: Benchmarking Code Generation with Diverse Function Calls and Complex Instructions

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-06-30T18:45:00.525070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:3d5cc1be0c4e0e08d9d00a1facbd93e0771fa1d8744a308097fdf4ff17668c43

Observation 49ce5992-0ec6-407e-a9eb-ef2370e8097e · outbound

This paper cites an unresolved cited work.

OpenCompass: A Universal Evaluation Platform for Large Language Models Unresolved cited work

Reference 24

Resolution
malformed identifier
raw_fallback, observed 2026-07-08T03:34:32.833117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-30T18:41:34.878542Z digest=sha256:c8e7324730a16d6cdaa74070d0e7e309605ca5070b5a6103c9015290f991bc24

Pith citing papers

Observation d9c2fd53-3dc8-4f36-93bc-8a52f7067c9a · inbound

MemSFT: Mitigating Alignment Tax with an External Parametric Memory cites this paper.

MemSFT: Mitigating Alignment Tax with an External Parametric Memory OpenCompass: A Universal Evaluation Platform for Large Language Models

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-01T01:59:49.787962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T01:59:49.787962Z digest=sha256:565c2392b1f2ff87b8cd662ffcce8e424164d8a9ad06550b5d5ff02eaba5ad10