Pith. sign in

Paper Citation Record · LEDGER

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs?

As of 13 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2412.10056.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.10056 v1

Coverage vector

measured 54 of 54 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T16:30:34.819197Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

54 of 54 outbound references displayed

  • verified exact1
  • verified fuzzy17
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation f9f34630-1069-48fd-a18f-d4f3e35f5d54 · outbound

This paper cites write newline.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.527170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.527170Z digest=sha256:46efc509560e47f4be9f04cd671b3a79d955b6c8ffc3c33f7538f043bcb4ce24

Observation 1beebf1a-d41d-4f68-ac14-b04e49435939 · outbound

This paper cites an unresolved cited work.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.534370Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.534370Z digest=sha256:d556d403f28a3ae0ce28fcf001dccd2f8c218dbb1c6fa253a3237e4ad802aa4d

Observation d06d8d1f-4f0e-4276-aa29-b4a44dac2429 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.539948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.539948Z digest=sha256:ea101cc133ecbe21f5c280d622fd2d3706e07d8c39b339552bb536d3f050fb62

Observation ebbe3dba-72e6-4765-a642-9988a64a8cb9 · outbound

This paper cites a is b" fail to learn.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? a is b" fail to learn

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:30:35.927480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T16:30:34.545544Z digest=sha256:d3c57065a3e1c073768948d19cef6cd5241b7676441f20d88b310f4250fc4939

Observation 238b180c-a372-431d-81a7-b42414758e1b · outbound

This paper cites Applying the rasch model: fundamental measurement in the human sciences, 2007.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Applying the rasch model: fundamental measurement in the human sciences, 2007

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:30:35.908255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T16:30:34.551170Z digest=sha256:cf524618fad0aef0c22f27fb3fd7e5fd26a3abb904cd4ca93df5b4b9ce593af7

Observation 6ae448e2-4f2d-43ef-b19d-74d80cf957bd · outbound

This paper cites Boone and Amity Noltemeyer.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Boone and Amity Noltemeyer

Reference 6

Resolution
metadata mismatch
raw_fallback, observed 2026-08-11T16:30:35.367589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T16:30:34.556401Z digest=sha256:870bb3f293c20e4a28681a72b6c7ff89313c9e41132ef579ca0d06c532c960ae

Observation 720f12fb-7884-4046-8d3d-c849f7bb9c8b · outbound

This paper cites Internlm2 technical report, 2024.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Internlm2 technical report, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.561380Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.561380Z digest=sha256:92f822d2c28d68b8ab70b80830d2e52adf450f9b55cea1c651d39b7f1775409a

Observation df7ad01c-87f9-45a1-9afe-1dd4a27d19dc · outbound

This paper cites M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.567451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.567451Z digest=sha256:050923aa8f38d4f39216dbcab4293b9ac424d8d9d141faae32c5e78604e610bb

Observation 2a341b58-3540-4bf4-b95a-95630a3b55c8 · outbound

This paper cites See What LLMs Cannot Answer: A Self-Challenge Framework for Uncovering LLM Weaknesses.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? See What LLMs Cannot Answer: A Self-Challenge Framework for Uncovering LLM Weaknesses

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.573319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.573319Z digest=sha256:db79d3642f69c64632a08555e66244d7185280566f7d73720f82ad2e0c702200

Observation 0680fa2f-fecc-4b3b-af2a-2c72670adfba · outbound

This paper cites Flash A ttention-2: Faster attention with better parallelism and work partitioning.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Flash A ttention-2: Faster attention with better parallelism and work partitioning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.582861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.582861Z digest=sha256:323a7377203c112d90ef89bbf381ae3e29f6ed7b0c03f2094635abdb13687f73

Observation 3a118e8d-5923-451d-ae2d-a38037b5c9f7 · outbound

This paper cites Large Language Models of Code Fail at Completing Code with Potential Bugs.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Large Language Models of Code Fail at Completing Code with Potential Bugs

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.588170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.588170Z digest=sha256:6dda80b9f2cbd08a3b1a23d3888b9d2c926350a6616eb9fbeb8d076b5f56367a

Observation 812d37ec-fa00-4e6a-8a5a-ed9a8814733b · outbound

This paper cites The Llama 3 Herd of Models.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? The Llama 3 Herd of Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.593851Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.593851Z digest=sha256:2a674dbd7d63001f1848321c5e4c39d1e6e64d1f2b4ae374477e9073b8d8a6d9

Observation 3243e408-c118-4e2f-a9e5-379ce6eab065 · outbound

This paper cites Chatglm: A family of large language models from glm-130b to glm-4 all tools, 2024.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Chatglm: A family of large language models from glm-130b to glm-4 all tools, 2024

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.599529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.599529Z digest=sha256:9971d38c2bbbcbc419addfb03cd7dd6a4c4367a34c771275d8d7bafc9d45e68f

Observation ea5a6b41-6cc2-4f67-9e9c-6d29285933cb · outbound

This paper cites Measuring massive multitask language understanding.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Measuring massive multitask language understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.604432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.604432Z digest=sha256:8e01b8163ae71a5d9d213b9edfcc8956b7e455b4ab1f371d2a2bedee033d0eb8

Observation d8de0acd-576a-40d2-846b-b7107dce6c1e · outbound

This paper cites Measuring massive multitask language understanding, January 2021 b.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Measuring massive multitask language understanding, January 2021 b

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:30:35.840345Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T16:30:34.611093Z digest=sha256:6d1a172dd468490d824b94415a97299656670c67037440572dc1d0d9f197a761

Observation f9c0d431-1110-40f5-a877-8be71f50a559 · outbound

This paper cites Measuring mathematical problem solving with the math dataset.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Measuring mathematical problem solving with the math dataset

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.615932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.615932Z digest=sha256:ce2dd2cea52d18f8d3ca3c39ff71b8938fcf16e72086f2a37337ef6a7d37f20c

Observation bb6e4b87-812c-471b-83af-6a2bf5457a75 · outbound

This paper cites C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:30:35.805435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T16:30:34.620810Z digest=sha256:bd4810c5b2d45ee351a27e8980972792dacdb095082c7add5e8a8bf0d3270544

Observation b2bea7fb-0b39-46d5-a9a7-5182f0f61443 · outbound

This paper cites New Ontology and Knowledge Graph for University Curriculum Recommendation.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? New Ontology and Knowledge Graph for University Curriculum Recommendation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:30:35.787252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T16:30:34.627029Z digest=sha256:64279daa53fe799870993573cfec5bfe3ac23e7cc8f85156cb4aeba468dbceba

Observation 699aae21-23b2-4c57-9c14-495595825d32 · outbound

This paper cites Large language models and simple, stupid bugs.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Large language models and simple, stupid bugs

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:30:35.768425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T16:30:34.631971Z digest=sha256:1dc586fa8b64c70fc4d16a15d1414f160ebd9e398746b24ea40cc07fe2fa3de0

Observation ab75a756-39e4-42ed-b87e-7a1b909235b4 · outbound

This paper cites FigureQA : An annotated figure dataset for visual reasoning, February 2018.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? FigureQA : An annotated figure dataset for visual reasoning, February 2018

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:30:35.750702Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T16:30:34.636984Z digest=sha256:a19af29a31e3fa8c3dd8c37d89e8d96fa68392bb6befeffc32613079ba453a66

Observation acbf3732-4f48-4231-acbd-b8cf6b609d78 · outbound

This paper cites A diagram is worth a dozen images, March 2016.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? A diagram is worth a dozen images, March 2016

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:30:35.731918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T16:30:34.642021Z digest=sha256:5e1436e162bc0e8a966ac4ab6a3db5baee7c7d2ac21b04f0912f717440a76cb6

Observation 50d3fac8-5555-4968-b1b9-419893fedc70 · outbound

This paper cites Rasch Measurement: Applications in Quantitative Educational Research, volume 1 of Education.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Rasch Measurement: Applications in Quantitative Educational Research, volume 1 of Education

Reference 22

Resolution
verified exact
doi, observed 2026-08-11T16:30:34.924315Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T16:30:34.646882Z digest=sha256:0a827795acebc6f39a228d60d2e13e6470535f37c99529e01032c55386ccf241

Observation c2524c55-333c-49c2-81fa-21fcdde0b44c · outbound

This paper cites Cmmlu: Measuring massive multitask language understanding in chinese, 2023.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Cmmlu: Measuring massive multitask language understanding in chinese, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.652187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.652187Z digest=sha256:c563bd6dcf181e182f2cb9859684f1a750d325430b532c52585af4eb43a4a0e3

Observation 56a7f43f-c67d-4643-9704-7c7807fcb0d4 · outbound

This paper cites CMMLU: measuring massive multitask language understanding in chinese.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? CMMLU: measuring massive multitask language understanding in chinese

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:30:35.700355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T16:30:34.657090Z digest=sha256:41596cf98798130156dd9bfd13b5aa9c115a1626a14d5cf70d8b62e9e8119ac4

Observation 862c2ae7-19d3-41f8-8a04-45789ea16988 · outbound

This paper cites Truthfulqa: Measuring how models mimic human falsehoods.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Truthfulqa: Measuring how models mimic human falsehoods

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.661779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.661779Z digest=sha256:db4288629355f15a358bc5e33c8c424031829ec5ec0f13ecb997195f2fe1cdc8

Observation e502e362-2694-42b3-b138-58a17e22547e · outbound

This paper cites MMBench : Is your multi-modal model an all-around player?, August 2024.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? MMBench : Is your multi-modal model an all-around player?, August 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:30:35.682459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T16:30:34.666335Z digest=sha256:c64a055984829e2e9c55de6f75bede1d163d09346a0f096f1c222e4aa16177e8

Observation ab262c57-8b40-40e5-b69e-1991f1982407 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering, October 2022 a.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Learn to explain: Multimodal reasoning via thought chains for science question answering, October 2022 a

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:30:35.664577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T16:30:34.670919Z digest=sha256:aa28bd2ba0ed84b340e163aa7eff4a7f65cfb6c7febf5ee102fd4c121260df11

Observation 3edc29f7-0156-48a2-8c94-6fb23322fa05 · outbound

This paper cites IconQA : A new benchmark for abstract diagram understanding and visual language reasoning, July 2022 b.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? IconQA : A new benchmark for abstract diagram understanding and visual language reasoning, July 2022 b

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:30:35.647385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T16:30:34.675623Z digest=sha256:c2bc26824a20e844db55c22ebecca5b71f79505ef8b0940fe6dd12345d1b65a1

Observation 1f8d3dba-57da-4986-99c8-fe07a68f3bb9 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.680434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.680434Z digest=sha256:67ab2e904ad045bdb29af39b20add34b842f3ce39ca3027fb267256e9a1e6cb4

Observation 8c248b91-8ab9-4a96-a7f0-663ef60f835a · outbound

This paper cites OK-VQA : A visual question answering benchmark requiring external knowledge, September 2019.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? OK-VQA : A visual question answering benchmark requiring external knowledge, September 2019

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:30:35.615582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T16:30:34.685200Z digest=sha256:270ba841bf77c18c61dc5ae83af958dcf270f7b67cb6ee0fcccf22600f3534f2

Observation b4c773d5-cbb0-400c-8c23-35915940be86 · outbound

This paper cites Mistral Large 2: Designed for Single-Node Inference with Long-Context.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Mistral Large 2: Designed for Single-Node Inference with Long-Context

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:30:35.594683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T16:30:34.690073Z digest=sha256:b2e2d14ce40c7d67c4059aecb7f3d7d7382e5e1594e4a1a2f7436699d59bc51c

Observation 17a8cc32-8cd0-499f-9d89-3ab18a11bf20 · outbound

This paper cites Training on the Benchmark Is Not All You Need.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Training on the Benchmark Is Not All You Need

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.695025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.695025Z digest=sha256:36c0d644d731462b1068b87ae23b4f22e6c171517bc42b7c2cdc920eaf6fd903

Observation dc5ec2c7-f3f5-437f-b7e4-f7fd51ea6349 · outbound

This paper cites GPT-4 Technical Report.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? GPT-4 Technical Report

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.702726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.702726Z digest=sha256:70ee8e5d0cf86227c13aa35ff2c30b030c8304237c73fd63bd4f8cb30677a914

Observation cfc35231-87ab-431a-b617-2423d282763e · outbound

This paper cites We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? We-Math: Does Your Large Multimodal Model Achieve Human-like Mathematical Reasoning?

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.708594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.708594Z digest=sha256:1cace9ffa76e7032ece14304981d0dbb5639634decb48497c3a7950a93236eee

Observation 586c737e-b5d0-44fa-b697-59adf5aaeac8 · outbound

This paper cites Probabilistic models for some intelligence and attainment tests.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Probabilistic models for some intelligence and attainment tests

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.713913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.713913Z digest=sha256:18fd27e6641b743153a69829c80a08c97a6c5d890dae6859723d11ce6880e79a

Observation 84fef0db-e386-495d-86d6-3aad81f94a37 · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Winogrande: An adversarial winograd schema challenge at scale

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.718989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.718989Z digest=sha256:2f17e60030d29c68f19951ec9316be1ae5f23ef3cc8e220331fc8f9a5f2f0a78

Observation 41d8d7db-ae83-4c73-b371-b63728d92ebb · outbound

This paper cites Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Math-LLaVA: Bootstrapping Mathematical Reasoning for Multimodal Large Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.724510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.724510Z digest=sha256:178ab3f1573f40e7ecd844aee8f5667a8ce1bc58fa12801967de5edfefaf6843

Observation ef9a2e91-6618-47de-a8f3-eadb8d6e1dd0 · outbound

This paper cites Assessing programming task difficulty for efficient evaluation of large language models.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Assessing programming task difficulty for efficient evaluation of large language models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.729872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.729872Z digest=sha256:a2996d2bdd7dbf424f417389f9c28ddfd91c37f1e149dbcd6d97d07687309e1a

Observation eadb0a71-d0e0-4b1e-922a-da1e1c8867d3 · outbound

This paper cites Internlm: A multilingual language model with progressively enhanced capabilities.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Internlm: A multilingual language model with progressively enhanced capabilities

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.734760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.734760Z digest=sha256:9e7d8f12d086dfaa134f60870d7d96b0414420907d8f549038339ad82368b76a

Observation 382f101f-6b49-47e6-8b35-2688d437a679 · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark, 2024.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Mmlu-pro: A more robust and challenging multi-task language understanding benchmark, 2024

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.739856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.739856Z digest=sha256:1b5073dbfb373e917d6093b06d582d1a6286c9b946db583bdc86ea115a7c0645

Observation cddee14a-bbc2-48d8-b72c-7f28dd03e332 · outbound

This paper cites Reasoning or reciting? exploring the capabilities and limitations of language models through counterfactual tasks.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Reasoning or reciting? exploring the capabilities and limitations of language models through counterfactual tasks

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:30:35.530830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T16:30:34.745082Z digest=sha256:8c7e402c70d9f37c895bc63b32e2780aaf39e6922e6d3d9f72f78accad5ab5da

Observation 981a2628-7dfc-43d5-8c5c-7073ee72c713 · outbound

This paper cites Qwen2 Technical Report.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Qwen2 Technical Report

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.752029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.752029Z digest=sha256:304c367260de46986380a2d9bc3acef9b8c21f127e5c15891f23110397a5f01b

Observation a546928d-c373-40bd-9018-3262832ed61d · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.757147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.757147Z digest=sha256:0858a0c39498d29ece3156004283ca1aae1970bbcc518415c694a17878428c7d

Observation b024bb3d-431b-4fd1-996f-03166bc670d0 · outbound

This paper cites Hellaswag: Can a machine really finish your sentence? In Anna Korhonen, David R.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Hellaswag: Can a machine really finish your sentence? In Anna Korhonen, David R

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.762725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.762725Z digest=sha256:9e81c83b04819d785d29b0f6c72f515b41ce22e88c112e61759c277cf7d040bc

Observation 66afb50f-3d1f-46ca-9669-24b7d6baeaa3 · outbound

This paper cites Evaluating the Performance of Large Language Models on GAOKAO Benchmark.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Evaluating the Performance of Large Language Models on GAOKAO Benchmark

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.768686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.768686Z digest=sha256:1f0c5096fe99847f59da92f7c433fb314b3b8ad83ff29c6578ed7e62b4647cb3

Observation 327bc6d7-4af6-4a84-b3e3-d2b9490da6e7 · outbound

This paper cites Can llm replace stack overflow? a study on robustness and reliability of large language model code generation.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Can llm replace stack overflow? a study on robustness and reliability of large language model code generation

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:30:35.500400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T16:30:34.774577Z digest=sha256:ec17cd8ec87485bf001bdf073d86280b25d181b05786933d1dd73564d4fd9a62

Observation e9f9fb6f-e8ce-4b77-8a6a-5e870ea1c6ee · outbound

This paper cites Don't Make Your LLM an Evaluation Benchmark Cheater.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Don't Make Your LLM an Evaluation Benchmark Cheater

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.780367Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.780367Z digest=sha256:dce4d2cb31e83fd4fd203ad68c1888bee50cee7979578dde69685d0652192f33

Observation a8208fe8-0b2f-4917-b2fa-3a15216799d2 · outbound

This paper cites Larger and more instructable language models become less reliable.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Larger and more instructable language models become less reliable

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.786183Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.786183Z digest=sha256:a4de0962a4b220b9a0b7aeeecd098915b6931c547f34468387571bc684a8af7a

Observation 03664885-0d05-47c9-8ada-3b3656cf283b · outbound

This paper cites Gaokao-mm: A chinese human-level benchmark for multimodal models evaluation, 2024.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Gaokao-mm: A chinese human-level benchmark for multimodal models evaluation, 2024

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T16:30:35.481241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T16:30:34.791177Z digest=sha256:aa284d19853b75d8df18262b9c01f36ab67b13b68347eedd1ba5466273972197

Observation ec80e9f8-c573-412e-893c-6516311357aa · outbound

This paper cites @esa (Ref.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? @esa (Ref

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.796305Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.796305Z digest=sha256:9cdaf86ca5c3220c2da62cf3a630052089438ead0797c0fd9527e2238beab9ce

Observation 5b9a984e-8ad1-4bc1-ab2a-e77f2536960a · outbound

This paper cites an unresolved cited work.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Unresolved cited work

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.802740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.802740Z digest=sha256:adecdbe639ae80973e429683b88ab3ca53bf49b32b248dda3fc2ba699ef87fdb

Observation 7a33b7fe-d241-4b37-ac08-1263e9b29ec0 · outbound

This paper cites an unresolved cited work.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-11T16:30:35.437386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T16:30:34.807826Z digest=sha256:113e543ad8533aba299a87851f3d177de36aaf98bd6e253dbad620d2540f7e47

Observation f213a072-b012-4ecd-8f93-9642034b8ab2 · outbound

This paper cites , " * write output.state after.block = add.period write.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? , " * write output.state after.block = add.period write

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.814439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.814439Z digest=sha256:89d528c7b00102c5be2695f9d38e4a3542a47e4affadd5309a0f0d969135643a

Observation 6af2e208-bf5f-4412-88c1-cf0aa85a9c7e · outbound

This paper cites write newline.

GAOKAO-Eval: Does high scores truly reflect strong capabilities in LLMs? write newline

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-11T16:30:34.819197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T16:30:34.819197Z digest=sha256:249fa6a13e630d09d435a13bb9c548692616d4170b117d8a95ded9e07f71567d

Pith citing papers

No inbound Pith citation observations are available.