Pith. sign in

Paper Citation Record · LEDGER

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars

As of 18 August 2026, this Paper Citation Record lists 17 of 17 outbound references and 1 inbound Pith citation observation for arXiv:2504.17562.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.17562 v2

Coverage vector

measured 17 of 17 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T10:40:21.640422Z

measured 18 of 18 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-17T04:46:33.641714Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-17T04:49:02.953758Z

Reference resolution

17 of 17 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved11
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5f6a8e5f-4826-437a-bd56-fb07154f952c · outbound

This paper cites HTLM: Hyper-Text Pre-Training and Prompting of Language Models.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars HTLM: Hyper-Text Pre-Training and Prompting of Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T10:40:21.569567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:40:21.569567Z digest=sha256:7990beccfd5b7447415a2f4508a1e892f41d6baa52f3d3e4d56d98f6c44d7f45

Observation 5ca024e4-4918-47e3-bf12-a9e51078cda4 · outbound

This paper cites Plug and Play Language Models: A Simple Approach to Controlled Text Generation.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars Plug and Play Language Models: A Simple Approach to Controlled Text Generation

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T10:40:21.597843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:40:21.597843Z digest=sha256:548d4d517daeda40a29be34919a10b0e5399e70c6c11da98901937ce5b416c20

Observation 29d8bd5e-f1ac-48ef-b137-78c7c3bb4333 · outbound

This paper cites A Theory of Emergent In-Context Learning as Implicit Structure Induction.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars A Theory of Emergent In-Context Learning as Implicit Structure Induction

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T10:40:21.607260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:40:21.607260Z digest=sha256:ebb3041ca155feefc99d792041b45742a3b985e46df87f383c068241fe76b125

Observation 5da20810-687a-4991-a0a4-8bb7687e7b92 · outbound

This paper cites CTRL: A Conditional Transformer Language Model for Controllable Generation.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars CTRL: A Conditional Transformer Language Model for Controllable Generation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-16T10:40:21.611507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:40:21.611507Z digest=sha256:18f165b1a99a49b4b37943f0cba6090569c2b2823d18fac9d86ed8d0836891a8

Observation 5e9c4b23-e569-4400-9b61-08e4f262930b · outbound

This paper cites Source-Aware Training Enables Knowledge Attribution in Language Models.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars Source-Aware Training Enables Knowledge Attribution in Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T10:40:21.615899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:40:21.615899Z digest=sha256:e6bbc49c7d512e81ee8e8807a37201a361578b1ef3214446726cc36b04cb2f4a

Observation 0c8d3eb8-84ea-472b-a12a-d0bfcf67ede1 · outbound

This paper cites Implicit meta-learning may lead language models to trust more reliable sources.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars Implicit meta-learning may lead language models to trust more reliable sources

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T10:40:21.620137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:40:21.620137Z digest=sha256:086f765c502be2fbde27e787d7899af6e7aa95391b8ef4e94521afee12eee041

Observation fad96e55-2756-4869-9d56-10ff2f9de7b7 · outbound

This paper cites Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth´ee Lacroix, Baptiste Rozi `ere, Naman Goyal, Eric Hambro, Faisal Azhar, et al.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars Hugo Touvron, Thibaut Lavril, Gautier Izacard, Xavier Martinet, Marie-Anne Lachaux, Timoth´ee Lacroix, Baptiste Rozi `ere, Naman Goyal, Eric Hambro, Faisal Azhar, et al

Reference 13

Resolution
malformed identifier
no resolver link, observed 2026-08-16T10:40:21.624364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:40:21.624364Z digest=sha256:4b6c5285c84b2c1007261f86675145815dd843c85a0699edf04d0892a5a1785a

Observation 59c8a36f-fe50-424f-a446-ae69a9dcffd0 · outbound

This paper cites Conditional Language Learning with Context.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars Conditional Language Learning with Context

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T10:40:21.628207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:40:21.628207Z digest=sha256:b4224adadab8ebf658a07a9f72fa40b4cc2017724d3fc4f886302c7039870390

Observation 196d73a5-9b48-4591-8ac3-6c4af658c168 · outbound

This paper cites Deciphering the impact of pretraining data on large language models through machine unlearning.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars Deciphering the impact of pretraining data on large language models through machine unlearning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:40:21.919692Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:40:21.632162Z digest=sha256:cb3cd962640b5622601ee1587e72b03511b9a9b1247668ae54d6423c2793c3e1

Observation a849e908-be1c-4485-8680-ec13837d9ccb · outbound

This paper cites 12 Published as a conference paper at COLM 2025 A More Details of the Experimental Setting Training Setting.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars 12 Published as a conference paper at COLM 2025 A More Details of the Experimental Setting Training Setting

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:40:21.907550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:40:21.636242Z digest=sha256:aee68f3404abe415f5810cfe4c4203e01b8eccf7cac9bee48040f0de1541daea

Observation cbefdf3e-f022-4134-8587-64e57b6af8f0 · outbound

This paper cites B The loss of next-token prediction for various token positions.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars B The loss of next-token prediction for various token positions

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T10:40:21.895155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T10:40:21.640422Z digest=sha256:307b2ca115bbe9ec4b27953b5385efb14f51cf352d18da3ce2fd502312f7a077

Observation 88cdee63-919e-4970-a1ce-fae8ab7d020f · outbound

This paper cites URL https: //aclanthology.org/P18-1198/.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars URL https: //aclanthology.org/P18-1198/

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-16T10:40:21.592365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:40:21.592365Z digest=sha256:d9db4419a4566f668175003c40da4cd1fd7829606051613bc332c75bc71e490e

Observation e3d15d29-0f1d-41c5-811b-1d81e8edf486 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-16T10:40:21.602496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:40:21.602496Z digest=sha256:1ec27c351dec2dc9828d808d2a39e6cfe76ba5318d6aa54a1f47104507e31aa8

Observation 804b2c0f-49a5-4f8b-bb7f-bceea1a77a75 · outbound

This paper cites CoCon: A Self-Supervised Approach for Controlled Text Generation.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars CoCon: A Self-Supervised Approach for Controlled Text Generation

Reference 2022

Resolution
malformed identifier
no resolver link, observed 2026-08-16T10:40:21.587527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:40:21.587527Z digest=sha256:45741a20b329dc26183a48877a66e57b9182c08cb9d73428ef987901afad50cb

Observation 705191bd-db36-43d6-966a-6c590f03ffbf · outbound

This paper cites Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars Physics of Language Models: Part 3.3, Knowledge Capacity Scaling Laws

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-16T10:40:21.578676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:40:21.578676Z digest=sha256:f9580da92ca19113569c5a6d68be7831c295e8dc070cdccbaaa4f8da615ecf38

Observation 6c110442-d4c6-442d-8a5d-30f0951b371a · outbound

This paper cites A Theory for Emergence of Complex Skills in Language Models.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars A Theory for Emergence of Complex Skills in Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-16T10:40:21.582670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:40:21.582670Z digest=sha256:a5277e41c392165b31fa1b016d0ca2b34cf1bf552cefd4bda9ff0ac46e3bdf43

Observation 99c4e181-5b57-431c-97e9-36560abea958 · outbound

This paper cites Physics of Language Models: Part 1, Learning Hierarchical Language Structures.

When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars Physics of Language Models: Part 1, Learning Hierarchical Language Structures

Reference 2025

Resolution
malformed identifier
no resolver link, observed 2026-08-16T10:40:21.574336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T10:40:21.574336Z digest=sha256:b7e929303088424747fbcdea865f05cb6b5e027cdc58ddf4ab9d9fd6b858b428

Pith citing papers

Observation 3d3eda9a-55b6-4fcf-8478-9f2759dde147 · inbound

Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining cites this paper.

Beyond URLs: Metadata Diversity and Position for Efficient LLM Pretraining When Does Metadata Conditioning (NOT) Work for Language Model Pre-Training? A Study with Context-Free Grammars

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:49:02.956416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T04:46:33.641714Z digest=sha256:af196f3857b17663752fc55111fd1ed1845fb65acdd551999d912c98c7211bb4