Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:27:15.532176Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2501.01588.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:27:15.532176Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
13 of 13 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d6ad1e48-9f66-42ca-a3fb-5aaf9ef04a85 · outbound
(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84b44dac-f64b-40f5-9aea-4a0366644dc4 · outbound
(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges Multiple-Choice Questions are Efficient and Robust LLM Evaluators
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d12c1df2-7bcf-45d2-b124-caaef8f28beb · outbound
(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges Generating multiple choice questions from a textbook: Llms match human performance on most metrics,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 90df8baa-0b88-4629-973b-b0de585f97bf · outbound
(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges Can Large Language Models Be an Alternative to Human Evaluations?
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1d586043-adb9-49e9-b4f0-6e2c49dbf84b · outbound
(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges Can multiple-choice questions really be useful in detecting the abilities of llms?
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 2de37227-0c18-46f5-8f7f-268f4cc607e1 · outbound
(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 890337d4-9d63-444e-a079-8bfae93a0a4e · outbound
(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges WinoGrande: An Adversarial Winograd Schema Challenge at Scale
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d8a34c59-aec6-411b-a581-99709ea201e5 · outbound
(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges Training verifiers to solve math word problems,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7bb84f82-5f93-4eb9-aee6-1fa1d99be9a7 · outbound
(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges Measuring Massive Multitask Language Understanding
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 117e356c-b6e1-43d9-90c2-85c131bb3dff · outbound
(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges Open-LLM-Leaderboard: From Multi-choice to Open-style Questions for LLMs Evaluation, Benchmark, and Arena
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 90ee99b9-3907-473d-8c89-1cad6d325aa5 · outbound
(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 47590e5e-a39f-4910-a352-27a662850092 · outbound
(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges TruthfulQA: Measuring How Models Mimic Human Falsehoods
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 85e783e0-b169-4857-82a0-eabc38ecdb0b · outbound
(WhyPHI) Fine-Tuning PHI-3 for Multiple-Choice Question Answering: Methodology, Results, and Challenges Training Verifiers to Solve Math Word Problems
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.