Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:34:24.301890Z
Paper Citation Record · LEDGER
As of 20 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 0 inbound Pith citation observations for arXiv:2506.07731.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:34:24.301890Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 95e80969-1d94-45df-8bf0-4334f80bc952 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models The Llama 3 Herd of Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 91fdd820-2027-4d60-96ff-0f045e4af147 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Qwen2.5 Technical Report
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7951fad2-7143-48e8-a759-05cbf3e6fc03 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Falcon2-11B Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0b345c9a-2afe-4fba-8302-c587dec33ebd · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models A Survey of Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b05157fc-a886-4ca6-8c9f-3703de78beb9 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models HellaSwag: Can a Machine Really Finish Your Sentence?
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1e551fb-b465-41af-92c3-d963a481847b · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Winogrande: An adversarial winograd schema challenge at scale
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fdd2bbe-bd98-4434-aa8f-bb9622febac3 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Measuring massive multitask language understanding
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d1f35d8-1ee2-4ebb-a52a-11869d87a5b0 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Mmlu-pro: A more robust and challenging multi-task language understanding benchmark
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6e8fa017-751c-4f4c-95c6-8e98ed6dc1ef · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Gpqa: A graduate-level google-proof q&a benchmark
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 344ce59a-6fd8-40d4-b415-f2f574edc337 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Training Verifiers to Solve Math Word Problems
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4572fc88-9d5e-4b5d-95cc-74cb531ba2ae · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Measuring Mathematical Problem Solving With the MATH Dataset
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffb23c3c-4554-42f2-9141-b0ec61f7c258 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models LiveCodeBench: Holistic and Contamination Free Evaluation of Large Language Models for Code
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53bbafaf-a6e2-4903-92fa-a0c8564ca637 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models What is the Role of Small Models in the LLM Era: A Survey
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52abb96c-b5eb-434a-a4a2-e9fb7259ebac · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Tensor Programs V: Tuning Large Neural Networks via Zero-Shot Hyperparameter Transfer
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4bfc5bee-5f46-407c-a886-a0599802c537 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models OLMoE: Open Mixture-of-Experts Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eec567c6-c754-4ab4-8baa-31b7459185dc · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1f48485-dae5-4d48-abad-457f907e18e4 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Liu, and Matt Gardner
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 67010805-7da1-477f-b792-654e356efed1 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Toward an Evaluation Science for Generative AI Systems
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation adb9d9d7-c967-4860-a65b-ae619bb15522 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Efficient Large Language Models: A Survey
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3870a5ef-4ab3-4d4e-a946-e1d41744f0c6 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Scaling Laws for Neural Language Models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9151034a-1c74-4880-a784-ddef281cfe68 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Training Compute-Optimal Large Language Models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e23300a7-7e69-4026-b516-7132558fa918 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models A survey on large language models: Applications, challenges, limitations, and practical usage
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a53c502c-186c-4836-a1b7-3b8763c440c6 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Llm merging: Building llms efficiently through merging
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3e6fed67-cb35-49a5-b64d-64ef0cb9b366 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Edge-llms: Edge-device large language model competition
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b96889db-bc37-443d-bd61-e9e03ada860c · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models https://llm-efficiency- challenge.github.io/, 2024
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ed11da27-2811-4ce5-89ed-3a19ce82a889 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Stages and individual differences in cognitive development
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation dd9d1bbc-a117-46dc-ae91-cd14e863b6ed · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models The new taxonomy of educational objectives
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 84d14a56-b06c-4121-b867-d78a700c5d94 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Taxonomy ofeducational objectives the clas- sification ofeducational goals
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 54c502d6-2abc-499a-9973-f1f726de7ec0 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models The fineweb datasets: Decanting the web for the finest text data at scale
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 43967017-ed14-46a0-ad3d-c0459a6db40e · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Fineweb-edu: the finest collection of educational content, 2024
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f3c32860-37fb-444a-804c-f831236b37e9 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models The stack: 3 tb of permissively licensed source code
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 23d2bf85-5658-4171-845f-89efdaab09d4 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Starcoder 2 and the stack v2: The next generation, 2024
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d79a0753-0b9e-4eac-9394-85654265dbe5 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Infimm-webmath-40b: Advancing multimodal pre-training for enhanced mathematical reasoning, 2024
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 997b618d-957d-4123-a691-13e7afbd9a01 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Unresolved cited work
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3da28a5b-baa0-4e81-9e21-65db149cc632 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Physics of Language Models: Part 3.1, Knowledge Storage and Extraction
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a66c55f8-e0fe-4544-8742-5198e6cc34d2 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 19f8f75e-e5fa-4486-aec8-96599f6670a7 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Finetasks: Finding signal in a haystack of 200+ multilingual tasks
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 97c02f2b-3666-44f4-a5c3-9e03237df4ea · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Codabench: Flexible, easy-to-use, and reproducible meta-benchmark platform
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation d8777a78-7700-4dc6-b742-6ad439ae53ea · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models The LAMBADA dataset: Word prediction requiring a broad discourse context
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3f422071-bfc0-44b8-8b57-3febe9296ab5 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Piqa: Reasoning about physical commonsense in natural language
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cde405fe-40ec-434a-93fc-9eb8df6c4de5 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Instruction-following evaluation for large language models, 2023
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6a9663b7-dbfa-41a9-ba14-678b57f0f707 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models RACE: Large-scale ReAding comprehension dataset from examinations
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4dc3f492-4861-45ce-8153-8d1086f0a9b0 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Semantic parsing on Freebase from question-answer pairs
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 63f10966-727f-4796-b79a-122ff02a3f14 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Challenging BIG-Bench Tasks and Whether Chain-of-Thought Can Solve Them
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 54ca9199-39f0-4c6b-8df8-ae5f61b43f84 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7bff9da3-0843-4fb2-a6f5-b26d72e9d769 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Boolq: Exploring the surprising difficulty of natural yes/no questions
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e4f64901-3ebc-44c7-8dfe-19f6722f4711 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Adversarial nli: A new benchmark for natural language understanding
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e3940417-37d4-48e3-b1cd-6776893ee0d6 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models TruthfulQA: Measuring How Models Mimic Human Falsehoods
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b6a31d19-0550-468b-a7ec-bf06c35be34f · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Mask R-CNN
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4cdcbf4b-16f7-489a-b5f5-182b4571eea9 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models 26 Figure 8: Validation loss across different datasets for two model variants: Scientific-DataMix (blue) and webOnly-Data (red)
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 5f8a63bd-5263-46f2-97e0-305c3ee1c572 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Lower loss on one dataset does not necessarily translate to higher capability on corresponding benchmarks
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b7d0f763-0cae-400e-8500-32697d552b51 · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8169890c-602a-4ab6-a869-f4992a17611c · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Unresolved cited work
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 8e7be798-6258-4d8b-aba3-98392e39ca8a · outbound
NeurIPS 2025 E2LM Competition : Early Training Evaluation of Language Models Unresolved cited work
Reference 2013
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
No inbound Pith citation observations are available.