Pith. sign in

Paper Citation Record · LEDGER

(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges

As of 15 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2501.01588.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.01588 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:27:15.532176Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved10
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d6ad1e48-9f66-42ca-a3fb-5aaf9ef04a85 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:15.489330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:15.489330Z digest=sha256:5456659754039ec083e185ff42de2e398d97b4d84f533cc67e31baac7ca55dca

Observation 84b44dac-f64b-40f5-9aea-4a0366644dc4 · outbound

This paper cites Multiple-Choice Questions are Efficient and Robust LLM Evaluators.

(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges Multiple-Choice Questions are Efficient and Robust LLM Evaluators

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:15.493787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:15.493787Z digest=sha256:e8f32348112ef300d41d1576f7b7bee3f622f951df8d894f0cfe390bbab570ef

Observation d12c1df2-7bcf-45d2-b124-caaef8f28beb · outbound

This paper cites Generating multiple choice questions from a textbook: Llms match human performance on most metrics,.

(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges Generating multiple choice questions from a textbook: Llms match human performance on most metrics,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:27:15.663550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T22:27:15.497723Z digest=sha256:c716217b8b24d8eb5ff69b476a5676cfbb5d814a7b60f216063992375b6bdf89

Observation 90df8baa-0b88-4629-973b-b0de585f97bf · outbound

This paper cites Can Large Language Models Be an Alternative to Human Evaluations?.

(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges Can Large Language Models Be an Alternative to Human Evaluations?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:15.501335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:15.501335Z digest=sha256:fd59f25c5bd133abc4f15a493861ead9bf83e75f852547459b6829b398d1801a

Observation 1d586043-adb9-49e9-b4f0-6e2c49dbf84b · outbound

This paper cites Can multiple-choice questions really be useful in detecting the abilities of llms?.

(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges Can multiple-choice questions really be useful in detecting the abilities of llms?

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:27:15.654071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T22:27:15.504985Z digest=sha256:c3fecdfb5a34e4218532601eaa3eefa094571a8d201d2d9682d074da25660033

Observation 2de37227-0c18-46f5-8f7f-268f4cc607e1 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:15.508868Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:15.508868Z digest=sha256:353eb779b9cdc411e348e705db0d240cd4601363df16c3653f2bc129dec8a728

Observation 890337d4-9d63-444e-a079-8bfae93a0a4e · outbound

This paper cites WinoGrande: An Adversarial Winograd Schema Challenge at Scale.

(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges WinoGrande: An Adversarial Winograd Schema Challenge at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:15.512501Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:15.512501Z digest=sha256:d4cf6bc68e0a09035235ee3c53db9061a5d2e0d0edaf99adddda1d41f8664b12

Observation d8a34c59-aec6-411b-a581-99709ea201e5 · outbound

This paper cites Training verifiers to solve math word problems,.

(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges Training verifiers to solve math word problems,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T22:27:15.643613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T22:27:15.516467Z digest=sha256:35fd162096032e4e4040b1ed59beb42b5768b42fbac54c941f8787eddc9644db

Observation 7bb84f82-5f93-4eb9-aee6-1fa1d99be9a7 · outbound

This paper cites Measuring Massive Multitask Language Understanding.

(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges Measuring Massive Multitask Language Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:15.522694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:15.522694Z digest=sha256:b24931ef2237942c620ccc5ab4f2153690bf9bcdf6216b583eafad19c2ee8e94

Observation 117e356c-b6e1-43d9-90c2-85c131bb3dff · outbound

This paper cites Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena.

(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:15.525599Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:15.525599Z digest=sha256:54d8113cfcfd3401a381e6a2241b5debaa5fa2a32baef1d5e56be851a54f6ca1

Observation 90ee99b9-3907-473d-8c89-1cad6d325aa5 · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:15.528407Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:15.528407Z digest=sha256:87dac932c6367fcc713aecc619618899a5f7c48b24be8bf1df5647776fb7b14b

Observation 47590e5e-a39f-4910-a352-27a662850092 · outbound

This paper cites TruthfulQA: Measuring How Models Mimic Human Falsehoods.

(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges TruthfulQA: Measuring How Models Mimic Human Falsehoods

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:15.532176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:15.532176Z digest=sha256:6039d7d5998f3fda9065064788fd6a08a1d79872199f7cb0c08e29879cc4bdd1

Observation 85e783e0-b169-4857-82a0-eabc38ecdb0b · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges Training Verifiers to Solve Math Word Problems

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-10T22:27:15.519792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:27:15.519792Z digest=sha256:0a151768aa56a90979859bc5448ab162d7673a56aedef3528ba216cbd44dc691

Pith citing papers

No inbound Pith citation observations are available.