Pith. sign in

Paper Citation Record · LEDGER

Benchmarking the Pedagogical Knowledge of Large Language Models

As of 21 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 2 inbound Pith citation observations for arXiv:2506.18710.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.18710 v3

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:20:19.995889Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:17:25.648174Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T22:20:49.429444Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 24bb187b-f29c-4585-9952-440af2095de3 · outbound

This paper cites an unresolved cited work.

Benchmarking the Pedagogical Knowledge of Large Language Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:20:23.826962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:20:17.479564Z digest=sha256:7729b5aa343b8132042182a76e42731af0beb0367fa55aab7996bff0654aee95

Observation cc8a224a-e939-4fd1-b83f-b674ead897e9 · outbound

This paper cites Lessons from the Trenches on Reproducible Evaluation of Language Models.

Benchmarking the Pedagogical Knowledge of Large Language Models Lessons from the Trenches on Reproducible Evaluation of Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:17.558649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:17.558649Z digest=sha256:2b3bc6714d338da771bbe4b4979d49a70351763930ab0fa8316d6ec9e104177e

Observation cb00ef13-f91c-4a82-ac67-dd4dde68a7a1 · outbound

This paper cites Distractor generation for multiple-choice questions with predictive prompting and large language models.

Benchmarking the Pedagogical Knowledge of Large Language Models Distractor generation for multiple-choice questions with predictive prompting and large language models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:17.679991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:17.679991Z digest=sha256:45c3989417e59e264451c6918a20a19ff8b45b5dc07da20c024b581058c585ea

Observation bea6d80c-6936-4123-94e9-6be154b57933 · outbound

This paper cites Sastry, A.

Benchmarking the Pedagogical Knowledge of Large Language Models Sastry, A

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:23.535094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:20:17.780784Z digest=sha256:62ed3b179fb70badee47e702cbfa3ea3e0a4b53a29d0c9ddca94bc5ea002c197

Observation 1624f01f-6c72-46a0-89fb-34ca3ca107d9 · outbound

This paper cites Chang, X.

Benchmarking the Pedagogical Knowledge of Large Language Models Chang, X

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:23.277396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:20:17.859253Z digest=sha256:b3bfdb6d40dedb255804c6fa914c2e1020ecc1fb2911e439ea93e22b93a14b17

Observation c9655ea1-32ec-4989-b546-a732b26c6c9d · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Benchmarking the Pedagogical Knowledge of Large Language Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:17.947382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:17.947382Z digest=sha256:4c2f6213ca43d818019b50a23ca4152a4370a319bd1fd884dce1fb4a6c50a14c

Observation ab6aa314-1541-4c9d-94fb-38a70c8ef7fc · outbound

This paper cites an unresolved cited work.

Benchmarking the Pedagogical Knowledge of Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:20:23.093985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:20:18.025882Z digest=sha256:68031c718a68f77b66a28a198e01463b0fe864a6cfb56170e6058ae98d640fd2

Observation b0f66b31-b1ae-4858-8d25-f3f518ecbbc3 · outbound

This paper cites Evaluating Large Language Models: A Comprehensive Survey.

Benchmarking the Pedagogical Knowledge of Large Language Models Evaluating Large Language Models: A Comprehensive Survey

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.098093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.098093Z digest=sha256:9270f64d56648c3b271fe7a19c8a6c87ee65cdc488adee5ed9ce70db2fe1278d

Observation b6d1f93e-e3dd-4009-9000-70e8e9a4973f · outbound

This paper cites Changing Answer Order Can Decrease MMLU Accuracy.

Benchmarking the Pedagogical Knowledge of Large Language Models Changing Answer Order Can Decrease MMLU Accuracy

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.169890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.169890Z digest=sha256:9d20f3be417453bd45fee8fd6f8342fcba4433cb0fb72288f1e1b34851c8637e

Observation f814db51-67bd-4305-8466-5d994260db4c · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Benchmarking the Pedagogical Knowledge of Large Language Models Measuring Massive Multitask Language Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.291469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.291469Z digest=sha256:d6a297d01bdf0ea8cc6fd413728e3547243c1ffd17e4fd601e5af19f3aafce59

Observation a7977192-8b60-4100-8ec4-7b10630d5d89 · outbound

This paper cites Kasenberg, A.

Benchmarking the Pedagogical Knowledge of Large Language Models Kasenberg, A

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.383691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.383691Z digest=sha256:6f21036c006ebc6130e2d162e682ae2c056918f82712b6b336313c7104091e30

Observation a825aa2b-cc49-4c4d-bcc5-922451871eec · outbound

This paper cites MinorBench: A hand-built benchmark for content-based risks for children.

Benchmarking the Pedagogical Knowledge of Large Language Models MinorBench: A hand-built benchmark for content-based risks for children

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.494890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.494890Z digest=sha256:19580a67aca6571fb7fff796cdf50fe2307f74ae3edc182b0bf63663e19e5020

Observation 5ebea77e-c4d2-444b-ba25-befa5559ed01 · outbound

This paper cites an unresolved cited work.

Benchmarking the Pedagogical Knowledge of Large Language Models Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:20:22.906863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:20:18.550027Z digest=sha256:82d260da7f30ec5232dca26a8d509884bd9e9ba10847514a832d8a545de8a69a

Observation c44b9237-e4a4-420a-a847-6a0db5c20830 · outbound

This paper cites Macina, N.

Benchmarking the Pedagogical Knowledge of Large Language Models Macina, N

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.624386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.624386Z digest=sha256:29a3c541dc8c5da34aaa582b5305195612faf1245018eb87bbc193f967fe4f8c

Observation 2f1c3676-d795-4007-a5ca-e4cf5528221d · outbound

This paper cites Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence.

Benchmarking the Pedagogical Knowledge of Large Language Models Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.697366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.697366Z digest=sha256:e93b9bb9515727ad707bceb4dc5723ca320f3cdb06dc1c24df714938c1f63ff4

Observation 2ec590f1-abd5-4724-a830-c2b6fc3b6637 · outbound

This paper cites Miller and K.

Benchmarking the Pedagogical Knowledge of Large Language Models Miller and K

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:22.542405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:20:18.788322Z digest=sha256:730c81b612aaf509150d95063a9d2b8952a004c0ea4a7272d95035b3784ab31b

Observation 66338182-1251-4ef6-9821-3d7cb07ade13 · outbound

This paper cites Nancy, D.

Benchmarking the Pedagogical Knowledge of Large Language Models Nancy, D

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:22.232468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:20:18.858576Z digest=sha256:efa1c6af094ac47d7cbf70bd49be66a60aca61d34eb01eeb34827c389be12fc0

Observation dcd55ea8-78ac-492d-8336-9986fa188871 · outbound

This paper cites Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions.

Benchmarking the Pedagogical Knowledge of Large Language Models Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.894389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.894389Z digest=sha256:2c1646fad308ba92be89acac88542cf2ae12d4540759a8220160948b235447a8

Observation f8c1cade-0438-4809-ab45-d0ef468f386b · outbound

This paper cites an unresolved cited work.

Benchmarking the Pedagogical Knowledge of Large Language Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:20:21.989970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:20:18.941933Z digest=sha256:5e40862d97ad2f24f9bfaa21bfc544f342567a2deee1b7c81a653fa83f704a91

Observation 79513115-4e9f-4085-9aa8-aa8a4158768d · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Benchmarking the Pedagogical Knowledge of Large Language Models GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:19.015163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:19.015163Z digest=sha256:cc895cc1635ee00f354846bc64c5850a38feabef7b0af692fcb521b97b5d9b5b

Observation 2f14bc13-72fa-4de4-9d86-c1a26c846bcf · outbound

This paper cites an unresolved cited work.

Benchmarking the Pedagogical Knowledge of Large Language Models Unresolved cited work

Reference 21

Resolution
verified exact
doi, observed 2026-08-06T23:20:20.188179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:20:19.083382Z digest=sha256:b97936b93fd5a36692e98ae00ac437487847bc9faf9a4a20066aa463c299d889

Observation f8914a9b-4d96-4faf-8417-c8c5e8d0b485 · outbound

This paper cites LearnLM: Improving Gemini for Learning.

Benchmarking the Pedagogical Knowledge of Large Language Models LearnLM: Improving Gemini for Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:19.149082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:19.149082Z digest=sha256:48c4014b58c22bbb3165d076a245aebbb961c3ba66c63d19ed7fae62f83694fb

Observation 80662d06-63a6-4f2f-bbd1-0c5299a439b6 · outbound

This paper cites Evaluating Gemini in an arena for learning.

Benchmarking the Pedagogical Knowledge of Large Language Models Evaluating Gemini in an arena for learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:19.262891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:19.262891Z digest=sha256:4fb5edafdf1545fa304f653dc845348884538bdda572e76040e5684e6d717692

Observation 256efb3c-ade7-419c-8204-c485a7f13395 · outbound

This paper cites Large Language Models are not Fair Evaluators.

Benchmarking the Pedagogical Knowledge of Large Language Models Large Language Models are not Fair Evaluators

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:19.346836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:19.346836Z digest=sha256:8e7a4cc5f183a5b828555358a4bbdf9f36986bb5a79235b5b3f7e09133f8e46c

Observation 2210cf71-da4c-40ff-9f41-bf3ab00f527c · outbound

This paper cites Large Language Models for Education: A Survey and Outlook.

Benchmarking the Pedagogical Knowledge of Large Language Models Large Language Models for Education: A Survey and Outlook

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:19.392025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:19.392025Z digest=sha256:18367efce631fa0355ea71cd19ae49bb7caff5c13e5fbbef848687781a2c00c0

Observation c0bf046d-b888-472b-94a8-13ab792af7b1 · outbound

This paper cites "My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models.

Benchmarking the Pedagogical Knowledge of Large Language Models "My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:19.476553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:19.476553Z digest=sha256:0a4b39ea787374be35bb7f5d88db9da98c47203815794c15004f2c520f875b7f

Observation 77f20ff9-76b7-42cc-85f2-bcabae8501c1 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Benchmarking the Pedagogical Knowledge of Large Language Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:19.533268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:19.533268Z digest=sha256:6a5f13db728228862ccfc71a121f62cf2beb4bd820ffe0c42fe749107dd9f6ba

Observation d81164cd-622c-4e6f-8ba6-db336f662f38 · outbound

This paper cites A Survey of Large Language Models.

Benchmarking the Pedagogical Knowledge of Large Language Models A Survey of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:19.618219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:19.618219Z digest=sha256:d1ca6647ea8849778ac12ff406fbfc2a473e65261a3f1078335307f481b8c60d

Observation 6d599a70-0eef-4cc6-9f62-cbc8081f806e · outbound

This paper cites Zheng, H.

Benchmarking the Pedagogical Knowledge of Large Language Models Zheng, H

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:21.782219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:20:19.682163Z digest=sha256:34611d7a4640cfb9673c229c97ee4ad7ede96c284287291aa015df3eff603753

Observation fe32bd1a-bcbb-4f4d-8740-b4e7ef7995f3 · outbound

This paper cites Wheredoyouthinkthereismoreplasticine?.

Benchmarking the Pedagogical Knowledge of Large Language Models Wheredoyouthinkthereismoreplasticine?

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:21.493176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:20:19.753332Z digest=sha256:c5f53129af85cd2d3bf88d3ee98060f7f4dc2f697fc15b7ee4418d63a955f30b

Observation 0846e8bf-2909-4ddd-af7e-d10ca4ffaa55 · outbound

This paper cites Question.

Benchmarking the Pedagogical Knowledge of Large Language Models Question

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:21.120052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:20:19.816947Z digest=sha256:b0d46fc4a7bd4335a54974c71231d33c7148267f13dd8eb7c6c6dbb11c9b664e

Observation a8232885-ba20-49c7-8ce3-03c1bca6a1c9 · outbound

This paper cites The text is as follows: —– paragraph —–.

Benchmarking the Pedagogical Knowledge of Large Language Models The text is as follows: —– paragraph —–

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:20.945282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:20:19.868065Z digest=sha256:75c98b715457fb2407bf656c03c8eb424f915dc1802e9f68c5bec7e53bbee3dd

Observation 99315814-de92-45c3-9951-84e64ec29d9f · outbound

This paper cites question.

Benchmarking the Pedagogical Knowledge of Large Language Models question

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:20.827125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:20:19.914280Z digest=sha256:3c40d1983cff799e3e6ed61198f9b8e12706e9e102634ef60d684cb933b9f1e6

Observation 7278ccc6-e8be-452c-8a1c-6b0307cd8019 · outbound

This paper cites question.

Benchmarking the Pedagogical Knowledge of Large Language Models question

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:20.698192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T23:20:19.995889Z digest=sha256:ab7e3fd8bb6bd1bff5e857e81a6309f0cee40ac3d36d329e414a21d6db0d49a4

Pith citing papers

Observation b77a933a-aeb1-469a-ac90-b173740e58cf · inbound

Application-Driven Pedagogical Knowledge Optimization of Open-Source LLMs via Reinforcement Learning and Supervised Fine-Tuning cites this paper.

Application-Driven Pedagogical Knowledge Optimization of Open-Source LLMs via Reinforcement Learning and Supervised Fine-Tuning Benchmarking the Pedagogical Knowledge of Large Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:20:49.432408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-10T19:55:17.077281Z digest=sha256:d33bacdb792d4c7b4aaf5148578fce011bf4b23667372dca40f21f472551a42c

Observation da3fb9dd-9cfb-47b3-acdd-2272bc508f34 · inbound

Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education cites this paper.

Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education Benchmarking the Pedagogical Knowledge of Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T04:17:25.648174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:17:25.648174Z digest=sha256:8e2c2dbed1bea024d42f66e542edde81de5b87e48b05e588ac0550d910835dce