Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T11:25:14.799196Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 0 inbound Pith citation observations for arXiv:2412.15524.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T11:25:14.799196Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
38 of 38 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 7706cd94-1b05-435c-b286-190ff7bc001b · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c881e4d1-d6f0-4a63-85f5-1a841c3a9640 · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models GPT-4 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf12983d-0afe-48b6-aa6b-e7c7a769ba9d · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Qwen Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 031d8da6-90bb-47b1-bb8a-c01afa940d30 · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Meteor: An automatic metric for mt evaluation with improved correlation with human judgments
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cad45276-edd1-49ee-8e1e-b7ba18e741a8 · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Pythia: A suite for analyzing large language models across training and scaling
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a7184216-70a3-482c-84be-7c9953d09c0e · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Language Models are Few-Shot Learners
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad41664b-7b22-43df-8473-1e4923bf080f · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Gonzalez, Ion Stoica, and Eric P
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f0bc3b7b-27ee-4374-8938-66702cf5f376 · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Gonzalez, and Ion Stoica
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bab4b4d1-647b-4872-b176-3c2fa75077ab · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Free dolly: Introducing the world’s first truly open instruction-tuned llm
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 789a0590-c104-4f09-9d69-a9a27dc3041f · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models The Llama 3 Herd of Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dbb5bae7-83ed-47af-a34e-4ddb375f7903 · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Length-Controlled AlpacaEval: A Simple Way to Debias Automatic Evaluators
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b40a680b-4166-4a2e-973b-ba933a162bf3 · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Prolific first, 2014
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e8f6d6a5-2a32-4ffe-ab85-ad6c5c1dc123 · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Koala: A dialogue model for academic research
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c61b52af-9bf4-40e9-8c4f-45cac6a4323e · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models OLMo: Accelerating the Science of Language Models
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5acb2964-ccae-4b79-b572-bf6f503148cf · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Mistral 7B
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7031387-07ef-4e83-8c48-a1284c4773d8 · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Tulu 3: Pushing Frontiers in Open Language Model Post-Training
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 20da4237-8daa-47d7-b4a6-bda86fc023f9 · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models From Crowdsourced Data to High-Quality Benchmarks: Arena-Hard and BenchBuilder Pipeline
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 572ca1ff-ee8c-4166-bf5b-4459465408c6 · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Hashimoto
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 41f67fbc-9b0f-4496-9aed-7946f7789fc2 · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models WildBench: Benchmarking LLMs with Challenging Tasks from Real Users in the Wild
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed911785-929d-4109-8c39-72eb30362baf · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Rouge: A package for automatic evaluation of summaries
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36f8f95d-5a6e-4b0a-9224-0cc542b41a37 · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models RoBERTa: A Robustly Optimized BERT Pretraining Approach
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e061115e-7682-4609-af9e-a0d808278e08 · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Hybrid Preferences: Learning to Route Instances for Human vs. AI Feedback
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 992f0be7-0a2a-4244-b48a-5a6aa77b3024 · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Training language models to follow instructions with human feedback
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a8d359b-4a7c-4894-b8c0-a6f3b0a20d75 · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Bleu: a method for automatic evaluation of machine translation
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2cfb50f5-8bd9-4199-974f-3f991d7c63ec · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Instruction Tuning with GPT-4
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 338f78d9-b7f2-4e35-b073-b1b18bbb01da · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Rush, and Thomas Wolf
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation dd800184-77d9-4d0a-ae27-631d3eb0e644 · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cae3dbb-297c-4350-bb5e-e42e7a5e23d8 · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Self-Instruct: Aligning Language Models with Self-Generated Instructions
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23fba9b4-7939-44c8-ac85-a20c49a25d39 · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Wizard LM : Empowering large pre-trained language models to follow complex instructions
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f726c6b-5af5-4628-8a61-c223d586a93e · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Yi: Open Foundation Models by 01.AI
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 78d5f30d-0b37-4831-8052-de1308fed6ba · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models BERTScore: Evaluating Text Generation with BERT
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f7bbac03-8951-477a-8718-8faaf4e2e05d · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models WildChat: 1M ChatGPT Interaction Logs in the Wild
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 191909aa-9571-4ddf-a930-03927830be7e · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Judging llm-as-a-judge with mt-bench and chatbot arena
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6f5bd239-60f3-4dc1-a9f7-fec408cad0e7 · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Lima: Less is more for alignment
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8705ba5c-2c97-44c5-8f55-c37a6eab485f · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models write newline
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d59554a-79f9-4417-9fc0-d1133d89f1b1 · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models @esa (Ref
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 776f9b1b-3835-48ef-b18d-8f9d18609a68 · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Unresolved cited work
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e33ad84e-b466-403e-b816-10e04b97dc4a · outbound
HREF: Human Response-Guided Evaluation of Instruction Following in Language Models Hide and Seek
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.