Pith. sign in

Paper Citation Record · LEDGER

Benchmarking the Pedagogical Knowledge of Large Language Models

As of 21 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 2 inbound Pith citation observations for arXiv:2506.18710.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.18710 v3

Coverage vector

measured 34 of 34 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:20:19.995889Z

measured 36 of 36 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T04:17:25.648174Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T22:20:49.429444Z

Reference resolution

34 of 34 outbound references displayed

  • verified exact1
  • verified fuzzy10
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 24bb187b-f29c-4585-9952-440af2095de3 · outbound

This paper cites an unresolved cited work.

Benchmarking the Pedagogical Knowledge of Large Language Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:20:23.826962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:20:17.479564Z digest=sha256:6045a8dca0ae08951c78520805e2b56e0b529dac37c2550f70ea7269a8c39d8a

Observation cc8a224a-e939-4fd1-b83f-b674ead897e9 · outbound

This paper cites Lessons from the Trenches on Reproducible Evaluation of Language Models.

Benchmarking the Pedagogical Knowledge of Large Language Models Lessons from the Trenches on Reproducible Evaluation of Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:17.558649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:17.558649Z digest=sha256:2b3bc6714d338da771bbe4b4979d49a70351763930ab0fa8316d6ec9e104177e

Observation cb00ef13-f91c-4a82-ac67-dd4dde68a7a1 · outbound

This paper cites Distractor generation for multiple-choice questions with predictive prompting and large language models.

Benchmarking the Pedagogical Knowledge of Large Language Models Distractor generation for multiple-choice questions with predictive prompting and large language models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:17.679991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:17.679991Z digest=sha256:45c3989417e59e264451c6918a20a19ff8b45b5dc07da20c024b581058c585ea

Observation bea6d80c-6936-4123-94e9-6be154b57933 · outbound

This paper cites Sastry, A.

Benchmarking the Pedagogical Knowledge of Large Language Models Sastry, A

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:23.535094Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:20:17.780784Z digest=sha256:feec91cda7c951b75f11c43709744f8e6d4d6f85eaf8ab19c0cfb691d7cb0c07

Observation 1624f01f-6c72-46a0-89fb-34ca3ca107d9 · outbound

This paper cites Chang, X.

Benchmarking the Pedagogical Knowledge of Large Language Models Chang, X

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:23.277396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:20:17.859253Z digest=sha256:de9ef7134faa6390709c09524c17166cec110427f12530a9ff3a9a9fa034dc64

Observation c9655ea1-32ec-4989-b546-a732b26c6c9d · outbound

This paper cites Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference.

Benchmarking the Pedagogical Knowledge of Large Language Models Chatbot Arena: An Open Platform for Evaluating LLMs by Human Preference

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:17.947382Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:17.947382Z digest=sha256:4c2f6213ca43d818019b50a23ca4152a4370a319bd1fd884dce1fb4a6c50a14c

Observation ab6aa314-1541-4c9d-94fb-38a70c8ef7fc · outbound

This paper cites an unresolved cited work.

Benchmarking the Pedagogical Knowledge of Large Language Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:20:23.093985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:20:18.025882Z digest=sha256:8d44c9ad1966dad13f80be30b90d2385571cb0a02684914c10c3862b63c59a72

Observation b0f66b31-b1ae-4858-8d25-f3f518ecbbc3 · outbound

This paper cites Evaluating Large Language Models: A Comprehensive Survey.

Benchmarking the Pedagogical Knowledge of Large Language Models Evaluating Large Language Models: A Comprehensive Survey

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.098093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.098093Z digest=sha256:9270f64d56648c3b271fe7a19c8a6c87ee65cdc488adee5ed9ce70db2fe1278d

Observation b6d1f93e-e3dd-4009-9000-70e8e9a4973f · outbound

This paper cites Changing Answer Order Can Decrease MMLU Accuracy.

Benchmarking the Pedagogical Knowledge of Large Language Models Changing Answer Order Can Decrease MMLU Accuracy

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.169890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.169890Z digest=sha256:9d20f3be417453bd45fee8fd6f8342fcba4433cb0fb72288f1e1b34851c8637e

Observation f814db51-67bd-4305-8466-5d994260db4c · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Benchmarking the Pedagogical Knowledge of Large Language Models Measuring Massive Multitask Language Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.291469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.291469Z digest=sha256:d6a297d01bdf0ea8cc6fd413728e3547243c1ffd17e4fd601e5af19f3aafce59

Observation a7977192-8b60-4100-8ec4-7b10630d5d89 · outbound

This paper cites Kasenberg, A.

Benchmarking the Pedagogical Knowledge of Large Language Models Kasenberg, A

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.383691Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.383691Z digest=sha256:6f21036c006ebc6130e2d162e682ae2c056918f82712b6b336313c7104091e30

Observation a825aa2b-cc49-4c4d-bcc5-922451871eec · outbound

This paper cites MinorBench: A hand-built benchmark for content-based risks for children.

Benchmarking the Pedagogical Knowledge of Large Language Models MinorBench: A hand-built benchmark for content-based risks for children

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.494890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.494890Z digest=sha256:19580a67aca6571fb7fff796cdf50fe2307f74ae3edc182b0bf63663e19e5020

Observation 5ebea77e-c4d2-444b-ba25-befa5559ed01 · outbound

This paper cites an unresolved cited work.

Benchmarking the Pedagogical Knowledge of Large Language Models Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:20:22.906863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:20:18.550027Z digest=sha256:89f41a41016ccef3c3b9c0fe987ad25ec7bb1fc245c73b4427053588d613588e

Observation c44b9237-e4a4-420a-a847-6a0db5c20830 · outbound

This paper cites Macina, N.

Benchmarking the Pedagogical Knowledge of Large Language Models Macina, N

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.624386Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.624386Z digest=sha256:29a3c541dc8c5da34aaa582b5305195612faf1245018eb87bbc193f967fe4f8c

Observation 2f1c3676-d795-4007-a5ca-e4cf5528221d · outbound

This paper cites Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence.

Benchmarking the Pedagogical Knowledge of Large Language Models Inadequacies of Large Language Model Benchmarks in the Era of Generative Artificial Intelligence

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.697366Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.697366Z digest=sha256:e93b9bb9515727ad707bceb4dc5723ca320f3cdb06dc1c24df714938c1f63ff4

Observation 2ec590f1-abd5-4724-a830-c2b6fc3b6637 · outbound

This paper cites Miller and K.

Benchmarking the Pedagogical Knowledge of Large Language Models Miller and K

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:22.542405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:20:18.788322Z digest=sha256:d2d70de1f7e6310c272982e5fddff9df77e1f051d44f66ecf68901d2f2419f82

Observation 66338182-1251-4ef6-9821-3d7cb07ade13 · outbound

This paper cites Nancy, D.

Benchmarking the Pedagogical Knowledge of Large Language Models Nancy, D

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:22.232468Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:20:18.858576Z digest=sha256:422064b14e965220a4849238968692f43df85ae7cdc5d4ee367f029b38a7d1fc

Observation dcd55ea8-78ac-492d-8336-9986fa188871 · outbound

This paper cites Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions.

Benchmarking the Pedagogical Knowledge of Large Language Models Large Language Models Sensitivity to The Order of Options in Multiple-Choice Questions

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:18.894389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:18.894389Z digest=sha256:2c1646fad308ba92be89acac88542cf2ae12d4540759a8220160948b235447a8

Observation f8c1cade-0438-4809-ab45-d0ef468f386b · outbound

This paper cites an unresolved cited work.

Benchmarking the Pedagogical Knowledge of Large Language Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T23:20:21.989970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:20:18.941933Z digest=sha256:a5d0e8e2335437a3bbdc24b2ffec230aa4dd70dcdbe68df114367ae6c86e1e0a

Observation 79513115-4e9f-4085-9aa8-aa8a4158768d · outbound

This paper cites GPQA: A Graduate-Level Google-Proof Q&A Benchmark.

Benchmarking the Pedagogical Knowledge of Large Language Models GPQA: A Graduate-Level Google-Proof Q&A Benchmark

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:19.015163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:19.015163Z digest=sha256:cc895cc1635ee00f354846bc64c5850a38feabef7b0af692fcb521b97b5d9b5b

Observation 2f14bc13-72fa-4de4-9d86-c1a26c846bcf · outbound

This paper cites an unresolved cited work.

Benchmarking the Pedagogical Knowledge of Large Language Models Unresolved cited work

Reference 21

Resolution
verified exact
doi, observed 2026-08-06T23:20:20.188179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:20:19.083382Z digest=sha256:8893817fb896b91241aeb7cc6cd939be904f2168ca63dceb42a8026a7fc98087

Observation f8914a9b-4d96-4faf-8417-c8c5e8d0b485 · outbound

This paper cites LearnLM: Improving Gemini for Learning.

Benchmarking the Pedagogical Knowledge of Large Language Models LearnLM: Improving Gemini for Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:19.149082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:19.149082Z digest=sha256:48c4014b58c22bbb3165d076a245aebbb961c3ba66c63d19ed7fae62f83694fb

Observation 80662d06-63a6-4f2f-bbd1-0c5299a439b6 · outbound

This paper cites Evaluating Gemini in an arena for learning.

Benchmarking the Pedagogical Knowledge of Large Language Models Evaluating Gemini in an arena for learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:19.262891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:19.262891Z digest=sha256:4fb5edafdf1545fa304f653dc845348884538bdda572e76040e5684e6d717692

Observation 256efb3c-ade7-419c-8204-c485a7f13395 · outbound

This paper cites Large Language Models are not Fair Evaluators.

Benchmarking the Pedagogical Knowledge of Large Language Models Large Language Models are not Fair Evaluators

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:19.346836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:19.346836Z digest=sha256:8e7a4cc5f183a5b828555358a4bbdf9f36986bb5a79235b5b3f7e09133f8e46c

Observation 2210cf71-da4c-40ff-9f41-bf3ab00f527c · outbound

This paper cites Large Language Models for Education: A Survey and Outlook.

Benchmarking the Pedagogical Knowledge of Large Language Models Large Language Models for Education: A Survey and Outlook

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:19.392025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:19.392025Z digest=sha256:18367efce631fa0355ea71cd19ae49bb7caff5c13e5fbbef848687781a2c00c0

Observation c0bf046d-b888-472b-94a8-13ab792af7b1 · outbound

This paper cites "My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models.

Benchmarking the Pedagogical Knowledge of Large Language Models "My Answer is C": First-Token Probabilities Do Not Match Text Answers in Instruction-Tuned Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:19.476553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:19.476553Z digest=sha256:0a4b39ea787374be35bb7f5d88db9da98c47203815794c15004f2c520f875b7f

Observation 77f20ff9-76b7-42cc-85f2-bcabae8501c1 · outbound

This paper cites MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark.

Benchmarking the Pedagogical Knowledge of Large Language Models MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:19.533268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:19.533268Z digest=sha256:6a5f13db728228862ccfc71a121f62cf2beb4bd820ffe0c42fe749107dd9f6ba

Observation d81164cd-622c-4e6f-8ba6-db336f662f38 · outbound

This paper cites A Survey of Large Language Models.

Benchmarking the Pedagogical Knowledge of Large Language Models A Survey of Large Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T23:20:19.618219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:20:19.618219Z digest=sha256:d1ca6647ea8849778ac12ff406fbfc2a473e65261a3f1078335307f481b8c60d

Observation 6d599a70-0eef-4cc6-9f62-cbc8081f806e · outbound

This paper cites Zheng, H.

Benchmarking the Pedagogical Knowledge of Large Language Models Zheng, H

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:21.782219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:20:19.682163Z digest=sha256:e0e49ee47ea7730677a9fbd4560dcc12e9fad372fe85712f8cf82acdd3e9d72b

Observation fe32bd1a-bcbb-4f4d-8740-b4e7ef7995f3 · outbound

This paper cites Wheredoyouthinkthereismoreplasticine?.

Benchmarking the Pedagogical Knowledge of Large Language Models Wheredoyouthinkthereismoreplasticine?

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:21.493176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:20:19.753332Z digest=sha256:49c6497b91049a9249ce0b0bda896f82cd53144790475599af7839ac972f7169

Observation 0846e8bf-2909-4ddd-af7e-d10ca4ffaa55 · outbound

This paper cites Question.

Benchmarking the Pedagogical Knowledge of Large Language Models Question

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:21.120052Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:20:19.816947Z digest=sha256:0694ae9a1d640fc8e73056513731c69979e5da6062eba3cc276204e98c583553

Observation a8232885-ba20-49c7-8ce3-03c1bca6a1c9 · outbound

This paper cites The text is as follows: —– paragraph —–.

Benchmarking the Pedagogical Knowledge of Large Language Models The text is as follows: —– paragraph —–

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:20.945282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:20:19.868065Z digest=sha256:d648d7e0a011481c18cde5e5284a54a23c9f2db8ade84b5984985878aeb379d5

Observation 99315814-de92-45c3-9951-84e64ec29d9f · outbound

This paper cites question.

Benchmarking the Pedagogical Knowledge of Large Language Models question

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:20.827125Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:20:19.914280Z digest=sha256:8e9dce7e6156d2043b3ab0bf6fcf20493a68fafb6d165a96b809140126fd6106

Observation 7278ccc6-e8be-452c-8a1c-6b0307cd8019 · outbound

This paper cites question.

Benchmarking the Pedagogical Knowledge of Large Language Models question

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:20:20.698192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-06T23:20:19.995889Z digest=sha256:b394b79c693fe974b288c877e3267b472cf2ff8ee1831107821093d60775fb70

Pith citing papers

Observation b77a933a-aeb1-469a-ac90-b173740e58cf · inbound

Application-Driven Pedagogical Knowledge Optimization of Open-Source LLMs via Reinforcement Learning and Supervised Fine-Tuning cites this paper.

Application-Driven Pedagogical Knowledge Optimization of Open-Source LLMs via Reinforcement Learning and Supervised Fine-Tuning Benchmarking the Pedagogical Knowledge of Large Language Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:20:49.432408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-10T19:55:17.077281Z digest=sha256:02944b3eac526c1074a3cc1227bb4e48e26c79bae6d2e62018b7a66cbc3a9836

Observation da3fb9dd-9cfb-47b3-acdd-2272bc508f34 · inbound

Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education cites this paper.

Calibrating Trustworthiness: Co-Designing Metrics and Visualizations for Evaluating LLMs in Education Benchmarking the Pedagogical Knowledge of Large Language Models

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-05T04:17:25.648174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T04:17:25.648174Z digest=sha256:8e2c2dbed1bea024d42f66e542edde81de5b87e48b05e588ac0550d910835dce