Pith. sign in

Paper Citation Record · LEDGER

Reasoning Bias of Next Token Prediction Training

As of 17 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 2 inbound Pith citation observations for arXiv:2502.02007.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02007 v2

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T13:46:43.619470Z

measured 69 of 69 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:05:57.739148Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-15T21:05:57.939113Z

Reference resolution

67 of 67 outbound references displayed

  • verified exact5
  • verified fuzzy23
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 908f4876-0b39-4468-b68b-25e6da6230ad · outbound

This paper cites Phi-4 Technical Report.

Reasoning Bias of Next Token Prediction Training Phi-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.282422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.282422Z digest=sha256:c6452fd81f2ff401b72b38870860ab10cbfe7a13e301e99fbbe5f46ebc8d4f00

Observation 0bc3db89-9f45-4529-a4eb-512b5c8a79a6 · outbound

This paper cites and Nagarajan, V.

Reasoning Bias of Next Token Prediction Training and Nagarajan, V

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.818716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.288710Z digest=sha256:be92c440c9f12035d1b88f30c77f97bac9d529054774c1e8a1ef34108f8dfbc5

Observation 5a091145-20fd-4987-a9f7-37a6a77efe5d · outbound

This paper cites and Giryes, R.

Reasoning Bias of Next Token Prediction Training and Giryes, R

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.803082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.294502Z digest=sha256:8180003808a9dcc6e48023af204f78bfda518ae7709784aeb6af2eba62e74e38

Observation 54745b23-35fb-4dec-9396-4dc895cbfc58 · outbound

This paper cites Multi-Step Deductive Reasoning Over Natural Language: An Empirical Study on Out-of-Distribution Generalisation.

Reasoning Bias of Next Token Prediction Training Multi-Step Deductive Reasoning Over Natural Language: An Empirical Study on Out-of-Distribution Generalisation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.299577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.299577Z digest=sha256:d2bec40918a436373f3b5ebf363cce6c1047b6b0671dbc1f9f2ae4d68b92f6d1

Observation b0fb012e-f872-44c1-b650-21a26278852c · outbound

This paper cites Understanding robustness of transformers for image classification.

Reasoning Bias of Next Token Prediction Training Understanding robustness of transformers for image classification

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.786239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.305137Z digest=sha256:20be8ab11cd194652637e29809ef3f55c101acdae495c0e562b9372d89a34010

Observation b1e68e2b-b26c-43cb-b507-2af169b07173 · outbound

This paper cites R., Angeli, G., Potts, C., and Manning, C.

Reasoning Bias of Next Token Prediction Training R., Angeli, G., Potts, C., and Manning, C

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.310512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.310512Z digest=sha256:7cb4592ad66cf4321c528edcd3a1d48c18940e96c0574b200dc567396e5ec075

Observation 64acb6d5-312d-4437-a06c-2ecbe4500664 · outbound

This paper cites Language Models are Few-Shot Learners.

Reasoning Bias of Next Token Prediction Training Language Models are Few-Shot Learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.316250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.316250Z digest=sha256:61d82ccc79bae919e7396a79c07f0a99f15adc1a77eacf37d7f1e62bff90fc6f

Observation 26d82bc1-9b13-477a-b139-85f4720360ab · outbound

This paper cites Dropout as a low-rank regularizer for matrix factorization.

Reasoning Bias of Next Token Prediction Training Dropout as a low-rank regularizer for matrix factorization

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.770532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.321652Z digest=sha256:3a71cffe6c0dd08930aa2a8025655f58470fed8f56a8f8c32c331afd22503cc8

Observation 72c42146-f7ca-464b-8eeb-13a91c8e4e1c · outbound

This paper cites Transformers as soft reasoners over language.

Reasoning Bias of Next Token Prediction Training Transformers as soft reasoners over language

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.326651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.326651Z digest=sha256:85d906827d2cdea36b969c2cad1866c5b21c876b35038b4b08bb71246c4e2519

Observation e1d7fe60-f30a-416f-a531-3f5520c0b2ca · outbound

This paper cites From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step.

Reasoning Bias of Next Token Prediction Training From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.331535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.331535Z digest=sha256:eb07e462a659632f911cf02090b52f5a35a2f70a9c9db838129c99a67193c876

Observation ec8fa746-963f-4ee2-a080-7addc029617f · outbound

This paper cites Reducing Transformer Depth on Demand with Structured Dropout.

Reasoning Bias of Next Token Prediction Training Reducing Transformer Depth on Demand with Structured Dropout

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.336880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.336880Z digest=sha256:6073e517b569a607299dc3c2a84d5e97d9dfed8a667cab7a9d93c5cc83c3c43b

Observation b3ed6204-64b0-475a-99a1-9b58e512403f · outbound

This paper cites and Tu, Y.

Reasoning Bias of Next Token Prediction Training and Tu, Y

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.754241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.342409Z digest=sha256:613f8acf775d4d8230fae1d3ca40e4e8be62aad0add10cbf1dcf67b57ad535b3

Observation 004c79a9-e648-4a5a-9026-951e34de47d9 · outbound

This paper cites Y., Roziere, B., Lopez-Paz, D., and Synnaeve, G.

Reasoning Bias of Next Token Prediction Training Y., Roziere, B., Lopez-Paz, D., and Synnaeve, G

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.739085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.347125Z digest=sha256:89b0e032beef198caf7f3aa8286d8157232f73638979ee5c811121e6ca91bf1c

Observation dcf418dc-aec7-462a-a395-17c6feb2b8d6 · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

Reasoning Bias of Next Token Prediction Training Training Large Language Models to Reason in a Continuous Latent Space

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.351846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.351846Z digest=sha256:04746a2ed5d23b8a50b54f93c5fefdcf6e475c0cc64964a2d69b051f0e36e8a9

Observation 0373d6dc-8f9b-411e-b6a5-e400175c010a · outbound

This paper cites A Law of Next-Token Prediction in Large Language Models.

Reasoning Bias of Next Token Prediction Training A Law of Next-Token Prediction in Large Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-09T13:46:44.306519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.357172Z digest=sha256:4e3bbbb68a7f280792285cf0042cac1433a25b4b10d80439b2925134d250ccf3

Observation 5889224b-91c3-4f62-977f-ef673b0eab86 · outbound

This paper cites What Matters in Transformers? Not All Attention is Needed.

Reasoning Bias of Next Token Prediction Training What Matters in Transformers? Not All Attention is Needed

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.362438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.362438Z digest=sha256:ac566f19dd4b2dd9458816d75231b60b0ecbde544ea3bd58f452075099bc5ab4

Observation 42335d35-d8bc-4683-8f70-55e86a492a40 · outbound

This paper cites Pretrained transformers improve out-of-distribution robustness.

Reasoning Bias of Next Token Prediction Training Pretrained transformers improve out-of-distribution robustness

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.367522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.367522Z digest=sha256:0a5f42fcdb03e87308890e836c5597f52d898ec2cbc079daa9b71e5e906d9902

Observation 80767613-99e7-4bfa-9773-0c646eae34da · outbound

This paper cites and Schmidhuber, J.

Reasoning Bias of Next Token Prediction Training and Schmidhuber, J

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.724058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.372397Z digest=sha256:e21eb6a9d8477fb8a51691001ec81ba24cd934a816a9507f415e9be1af97b8b6

Observation d3608fce-dd3f-4a81-94ca-1d129db31b3e · outbound

This paper cites NEFTune: Noisy Embeddings Improve Instruction Finetuning.

Reasoning Bias of Next Token Prediction Training NEFTune: Noisy Embeddings Improve Instruction Finetuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.378061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.378061Z digest=sha256:ea11c00c55426b292fdace736a96691b4b4a39699911b6b4aa0c1944802736c5

Observation fdeedeed-ec45-46a3-90c1-5753574ac00b · outbound

This paper cites S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P.

Reasoning Bias of Next Token Prediction Training S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.383548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.383548Z digest=sha256:c35188235a1056ac4eed56f5deeeec0914c09b3c9dc503a974aa9e0071fe66b0

Observation d34e9dcb-e86e-48a9-8da7-8b7bdf6dbfce · outbound

This paper cites N., Hellmann, S., Morsey, M., Van Kleef, P., Auer, S., et al.

Reasoning Bias of Next Token Prediction Training N., Hellmann, S., Morsey, M., Van Kleef, P., Auer, S., et al

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.697997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.388908Z digest=sha256:da10116a352dc32305b12202e1bd2631aa5d560924ccd0db5bf4ef86adbd9f6a

Observation a0bf94c7-c237-48c6-9fb2-05d1b59d2d50 · outbound

This paper cites J., Xing, E., and Caruana, R.

Reasoning Bias of Next Token Prediction Training J., Xing, E., and Caruana, R

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.681991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.393776Z digest=sha256:44b0039b2c36740b1c018c94d250e3fdbb5dc6c22d68757e3e84d2d5ed88d20d

Observation 599dc58c-32e0-4abc-9f01-e43c660afdad · outbound

This paper cites DropKey.

Reasoning Bias of Next Token Prediction Training DropKey

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.398762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.398762Z digest=sha256:2de3ec116669e3f0da9940a6540ee8f47f7bf2a7eb3336dc9fca6bcb7acd7d8d

Observation 045c0bc8-f000-49af-bbff-6e95ce11acf6 · outbound

This paper cites Challenging large language models with new tasks: A study on their adaptability and robustness.

Reasoning Bias of Next Token Prediction Training Challenging large language models with new tasks: A study on their adaptability and robustness

Reference 24

Resolution
verified exact
doi, observed 2026-08-09T13:46:43.800168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.403768Z digest=sha256:b0185b0cf82ab160b6975d4b3e9528b2af8fb682029e754562f481c60efbd613

Observation ce9ae42e-ea01-4df8-91f7-653a262e3cbd · outbound

This paper cites Visualizing the loss landscape of neural nets.

Reasoning Bias of Next Token Prediction Training Visualizing the loss landscape of neural nets

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.665175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.408608Z digest=sha256:5bf6f6f804bda3d8e332eea822bb1a9de289062a56338b5e7e85179959f49385

Observation 71f357b0-9897-4fb4-9bc3-24524f166e75 · outbound

This paper cites E., Singh Rawat, A., and Oymak, S.

Reasoning Bias of Next Token Prediction Training E., Singh Rawat, A., and Oymak, S

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.648043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.413206Z digest=sha256:2e409db525a1190d010b415ac2e1d20580838216494a5ecf8c8543d93310987c

Observation 638fbb91-2244-4cfb-bbf0-bd7c71d5aa05 · outbound

This paper cites Rho-1: Not All Tokens Are What You Need.

Reasoning Bias of Next Token Prediction Training Rho-1: Not All Tokens Are What You Need

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.418162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.418162Z digest=sha256:f6552854bbda5fc40b9520d453033a586d632738b782dc3c8fffca4e4e3f293c

Observation bb035bc5-c676-4945-be4e-7ed6c8263df2 · outbound

This paper cites M., Li, Z., and Ma, T.

Reasoning Bias of Next Token Prediction Training M., Li, Z., and Ma, T

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.632706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.423209Z digest=sha256:711d62c34e43536479e01eb9c93a248d765512d25c3c782eb8b32df999d5934f

Observation 405bb9c5-4ebb-4351-829a-b8ec1772e82f · outbound

This paper cites and Ying, L.

Reasoning Bias of Next Token Prediction Training and Ying, L

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.616824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.428213Z digest=sha256:19cc8b66df4c6528e157b2481b6931a47dd0326e8962590dc1b2d7e97654f766

Observation 2ea55aa7-4095-49a2-8ac6-d49fba158f4d · outbound

This paper cites Next-token prediction capacity: general upper bounds and a lower bound for transformers, 2024.

Reasoning Bias of Next Token Prediction Training Next-token prediction capacity: general upper bounds and a lower bound for transformers, 2024

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.433208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.433208Z digest=sha256:81111d5ccedd8d28b72430efa81eac5a62e11bd90e553a599e9b6c3539392b46

Observation 3f18b2ec-3d31-4193-9860-5a52fe1ae1bb · outbound

This paper cites On the implicit bias of dropout.

Reasoning Bias of Next Token Prediction Training On the implicit bias of dropout

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.601307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.438140Z digest=sha256:6590a35b461739609d3c8e7a0fc67640c361fae280bce089908324d87b5d7eb0

Observation 53d00bd4-9b33-4657-b951-5a431f8cafe1 · outbound

This paper cites Pretrained Transformers Do not Always Improve Robustness.

Reasoning Bias of Next Token Prediction Training Pretrained Transformers Do not Always Improve Robustness

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-09T13:46:44.127239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.443060Z digest=sha256:9e06782354c7f6e7f0f537808b4ffa135bc7d3bd5159182ce87b0ff771fd574b

Observation 128a33fd-553f-42f0-a89e-e447c3a3bb11 · outbound

This paper cites and Samwald, M.

Reasoning Bias of Next Token Prediction Training and Samwald, M

Reference 33

Resolution
verified exact
doi, observed 2026-08-09T13:46:43.783504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.448286Z digest=sha256:b30ab2c39caa4c8f9f8afb55bc9d891f0dd019a8ce7d867a9156a3527ab905e5

Observation 79db3ce0-125d-490e-9828-df83de123781 · outbound

This paper cites Power-law escape rate of SGD.

Reasoning Bias of Next Token Prediction Training Power-law escape rate of SGD

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.453263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.453263Z digest=sha256:7db46d520d6e7995381a87c074de8241b33aa22bc34a6462f232b27095cdda0f

Observation 36c2b9ac-3e90-46a9-8e47-10777162b5a6 · outbound

This paper cites LogicInference: A New Dataset for Teaching Logical Inference to seq2seq Models.

Reasoning Bias of Next Token Prediction Training LogicInference: A New Dataset for Teaching Logical Inference to seq2seq Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.459652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.459652Z digest=sha256:a44fa6f834afb2ed0688a754feba208dd9e0891dc77c66af60e6dacbbdf9d302

Observation 991996d3-8a7c-4857-b014-46ed8f8781cc · outbound

This paper cites and Narasimhan, K.

Reasoning Bias of Next Token Prediction Training and Narasimhan, K

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.584285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.464879Z digest=sha256:fe22d48a72be64933063ce5cf898875385646c8d3545246d6b17d4fe7a709803

Observation 7cbfafc1-b7aa-4038-b36f-4521b0666e6d · outbound

This paper cites Language models are unsupervised multitask learners.

Reasoning Bias of Next Token Prediction Training Language models are unsupervised multitask learners

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.469944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.469944Z digest=sha256:90636d600a4471043fa78f6052b2ad2a0c49da8a2724abf071c7df30455ace7e

Observation d24bcdc6-d40d-449b-b2bd-fdd6d1898055 · outbound

This paper cites R obust LR : A diagnostic benchmark for evaluating logical robustness of deductive reasoners.

Reasoning Bias of Next Token Prediction Training R obust LR : A diagnostic benchmark for evaluating logical robustness of deductive reasoners

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.474845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.474845Z digest=sha256:3fb8f7a934a774ee6c3132172a50eed104a28b2408e6a9e92f16613c71632bf5

Observation 1242e8b6-90c7-41f1-a5a8-e94f5255b8ed · outbound

This paper cites and He, H.

Reasoning Bias of Next Token Prediction Training and He, H

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.479930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.479930Z digest=sha256:6d57faefc83b976ea650b6dbaacdab623d5d0281f729a51e05b4861e11e6c0d2

Observation e0def47a-ad1f-4840-9711-164b47e0c829 · outbound

This paper cites Stepgame: A new benchmark for robust multi-hop spatial reasoning in texts.

Reasoning Bias of Next Token Prediction Training Stepgame: A new benchmark for robust multi-hop spatial reasoning in texts

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.485090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.485090Z digest=sha256:5d96c4ab5a7dcf58a69d007623ff8bb6bb778cd923478aaf587b4edab3de610a

Observation 03ca9248-61bf-40b2-8d41-56918fad7417 · outbound

This paper cites an unresolved cited work.

Reasoning Bias of Next Token Prediction Training Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.490010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.490010Z digest=sha256:ab2c6d74519b1076da70a3498ac257290206eed97cb4117fa9f77ee848099640

Observation 50ab1b23-a4ab-427b-86b2-c90363ac2245 · outbound

This paper cites P roof W riter: Generating implications, proofs, and abductive statements over natural language.

Reasoning Bias of Next Token Prediction Training P roof W riter: Generating implications, proofs, and abductive statements over natural language

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.495140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.495140Z digest=sha256:a24c0f2ca365ba7ab9f80fa7932a4afb8605c6970251d2c34d32d030a40ac2dc

Observation 2d6c2844-7d4b-412d-b39f-f2f929cb2d8d · outbound

This paper cites Memorisation versus generalisation in pre-trained language models.

Reasoning Bias of Next Token Prediction Training Memorisation versus generalisation in pre-trained language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.500250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.500250Z digest=sha256:820a29ed987ac3900ac7ff0ddef644d81bd191cd5a47649c08c5f0a1db0f2ee1

Observation 4d0b59b2-752d-417e-a8e7-4396418c8023 · outbound

This paper cites Implicit optimization bias of next-token prediction in linear models.

Reasoning Bias of Next Token Prediction Training Implicit optimization bias of next-token prediction in linear models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.547038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.505179Z digest=sha256:ad0c1c18d6f572eca70ecbb929fd6cba5f660314fe59224e4817759f134c8b90

Observation 9104245d-a585-445a-9e04-f9e73cdc3f99 · outbound

This paper cites An empirical study on robustness to spurious correlations using pre-trained language models.

Reasoning Bias of Next Token Prediction Training An empirical study on robustness to spurious correlations using pre-trained language models

Reference 45

Resolution
verified exact
doi, observed 2026-08-09T13:46:43.713990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.509922Z digest=sha256:21589100528e1d68c6427d20843b0a6e2608fa2810c97a58f5820ba12213d089

Observation 4d3139d1-8c38-47b7-a762-b88b682e82df · outbound

This paper cites L ogic A sker: Evaluating and improving the logical reasoning ability of large language models.

Reasoning Bias of Next Token Prediction Training L ogic A sker: Evaluating and improving the logical reasoning ability of large language models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.514920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.514920Z digest=sha256:302f3e5b5d0d3ff7a9d5b20e2aa1a03d1b8b6cd9d626f7972136978d4afc2e0b

Observation 9d3daf19-287c-4f20-aa3b-8b685bcc4b2b · outbound

This paper cites Are large language models really robust to word-level perturbations? In Socially Responsible Language Modelling Research, 2023.

Reasoning Bias of Next Token Prediction Training Are large language models really robust to word-level perturbations? In Socially Responsible Language Modelling Research, 2023

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.529595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.519722Z digest=sha256:77ffddb0b66751a49c1bcf29b84a0a716326b58621fde4f60a38495f2e776827

Observation cc99b161-f150-445d-a9f1-c18593e4f879 · outbound

This paper cites The implicit and explicit regularization effects of dropout.

Reasoning Bias of Next Token Prediction Training The implicit and explicit regularization effects of dropout

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.513937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.524447Z digest=sha256:c3d213bf4f803f8509f6970c13be36df2336578f9e1ee4d0899abac9e118479a

Observation 0448aee8-c29c-46bb-825d-985cba3b0ada · outbound

This paper cites Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks.

Reasoning Bias of Next Token Prediction Training Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.529135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.529135Z digest=sha256:27c39b70baa6a07fcefe28c86cf01d308e53c3b66794e8a0b890da29b8dfc4ee

Observation 7fcd35b3-2ada-4779-b3dc-6ee73103d0c0 · outbound

This paper cites On the noisy gradient descent that generalizes as sgd.

Reasoning Bias of Next Token Prediction Training On the noisy gradient descent that generalizes as sgd

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.496363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.534025Z digest=sha256:d2027660f2238245aac40d9c9acac5f957d6594ff8a8731c062f9235cc82b210

Observation 21ebed3d-a2d9-431a-a33c-299eba61061f · outbound

This paper cites How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective.

Reasoning Bias of Next Token Prediction Training How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.480659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.539227Z digest=sha256:326bcdaf74ec711d44d95b4a820fde25cd274eb6881a626b31b07ca68466ce8a

Observation e45537e8-832f-4ceb-9a8b-bfad031b41e3 · outbound

This paper cites UniDrop: A Simple yet Effective Technique to Improve Transformer without Extra Cost.

Reasoning Bias of Next Token Prediction Training UniDrop: A Simple yet Effective Technique to Improve Transformer without Extra Cost

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.544075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.544075Z digest=sha256:f9402c56d38601009ebfc1216da2f9134fb99dc4b293776528bc3fc0f0512bd8

Observation 87d8e884-4966-49fa-aae4-9ba4027f9c3c · outbound

This paper cites D., and Potts, C.

Reasoning Bias of Next Token Prediction Training D., and Potts, C

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.549000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.549000Z digest=sha256:2f2c539ff9c6f2f4c0f31b221c1bf8271bba46876492a9d79cb551c32e497993

Observation 93be4192-acd6-4e66-9346-1ba25f5a8a5f · outbound

This paper cites A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima.

Reasoning Bias of Next Token Prediction Training A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.554031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.554031Z digest=sha256:6ae1f1e7cc41e606d7627a49f94d5bb75f7a110dcaff4a7d7c748e6e308f70ee

Observation ef9ae870-53ef-4629-90b6-895e2c03ed87 · outbound

This paper cites URL http://www.yelp.com/ dataset_challenge.

Reasoning Bias of Next Token Prediction Training URL http://www.yelp.com/ dataset_challenge

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.465135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.559095Z digest=sha256:6cbf0be85b3eb13824f4d0575333cc7dd1bc13e4bf030798f7c8fe1bce48c4e5

Observation ffb8e00d-fa33-4783-b3d8-65f358a901f5 · outbound

This paper cites InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning.

Reasoning Bias of Next Token Prediction Training InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.563724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.563724Z digest=sha256:a4488c251848d0a02b75baf4cd2f5ce89a9a9c7751b9f9d02c249cf0eab811e8

Observation c837e3f2-6754-489a-a334-e4e3b937e06f · outbound

This paper cites Natural language reasoning, a survey.

Reasoning Bias of Next Token Prediction Training Natural language reasoning, a survey

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.568890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.568890Z digest=sha256:fe78e31e190e113cc3f838012abd73f79ce8f8a1461cb5d7fdbe2ddf85fa0ba0

Observation 7752d722-275c-47e2-a23a-c363ea5c1b53 · outbound

This paper cites DropAttention: A Regularization Method for Fully-Connected Self-Attention Networks.

Reasoning Bias of Next Token Prediction Training DropAttention: A Regularization Method for Fully-Connected Self-Attention Networks

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.573794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.573794Z digest=sha256:7f950dadf8d63a43c813941f64aa51a59744b220f3c8e15e6a9fa8196e39d044

Observation 3b913036-4a56-4b15-87c9-92fa484ac6e2 · outbound

This paper cites Dropdim: A regularization method for transformer networks.

Reasoning Bias of Next Token Prediction Training Dropdim: A regularization method for transformer networks

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.579115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.579115Z digest=sha256:a17f24646a0644f8dc398ba15a0613115735a87629dd9f41f81fc6bb92b164a2

Observation 74d55690-5c9c-4951-b325-91bfac772083 · outbound

This paper cites H., Meng, T., Chang, K.-W., and Van den Broeck, G.

Reasoning Bias of Next Token Prediction Training H., Meng, T., Chang, K.-W., and Van den Broeck, G

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.583817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.583817Z digest=sha256:464a880235e84874b67cd3ffe7aae409684fa9c520025ca2bcabe4664f431ce5

Observation 33ebdccc-6c5d-42b1-8b4f-0d681e7bd3f0 · outbound

This paper cites and Xu, Z.-Q.

Reasoning Bias of Next Token Prediction Training and Xu, Z.-Q

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.588944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.588944Z digest=sha256:93003fa1033c42e67dcf3b74ac1465c0f6ceea5049a88978f30c3113a37a3ea8

Observation f61a50ba-e91a-4684-979c-b1dcc1ae298b · outbound

This paper cites Stochastic Modified Equations and Dynamics of Dropout Algorithm.

Reasoning Bias of Next Token Prediction Training Stochastic Modified Equations and Dynamics of Dropout Algorithm

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.594201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.594201Z digest=sha256:a19d550986cb9a7e3d182899efd417f0bdfaf5c7545819f9ac131a9f975a98df

Observation 98a53005-a589-4721-9d72-b65c5d036bc9 · outbound

This paper cites Initialization is Critical to Whether Transformers Fit Composite Functions by Reasoning or Memorizing.

Reasoning Bias of Next Token Prediction Training Initialization is Critical to Whether Transformers Fit Composite Functions by Reasoning or Memorizing

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.599362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.599362Z digest=sha256:a809a6137babb60f5f1b6551149729f6a71593394e30f4323e4a1fe411cdb375

Observation bc20519d-1dad-4b65-ae7c-c7d1f8a0e807 · outbound

This paper cites Anchor function: a type of benchmark functions for studying language models.

Reasoning Bias of Next Token Prediction Training Anchor function: a type of benchmark functions for studying language models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.604752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.604752Z digest=sha256:0af517116d7e863a50497858267557a182bfd06581f422a12783833236bca0d0

Observation ddbdd9be-60e5-4cbb-8e00-b098c371f2df · outbound

This paper cites Implicit geometry of next-token prediction: From language sparsity patterns to model representations.

Reasoning Bias of Next Token Prediction Training Implicit geometry of next-token prediction: From language sparsity patterns to model representations

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.440791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.609837Z digest=sha256:15d1bd54e8f7d9756bd35253e24300a6e58df8aead90e69d2e09bd548eb9dd94

Observation d9675f5f-907b-4fcc-814f-96896e1c0ded · outbound

This paper cites Scheduled D rop H ead: A regularization method for transformer models.

Reasoning Bias of Next Token Prediction Training Scheduled D rop H ead: A regularization method for transformer models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.614551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.614551Z digest=sha256:c83b708d4960ffe93b9423228f598b86e78e704dbb8599bec0f887079cc5b66b

Observation 30af6bad-2d3a-4fcf-9b1f-dc57d4457767 · outbound

This paper cites The anisotropic noise in stochastic gradient descent: Its behavior of escaping from sharp minima and regularization effects.

Reasoning Bias of Next Token Prediction Training The anisotropic noise in stochastic gradient descent: Its behavior of escaping from sharp minima and regularization effects

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.424735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.619470Z digest=sha256:6a460f5aba2335aca70ee8080bcb3e1c3335881b112b20a585d196bead6ef65b

Pith citing papers

Observation 7fa8fe6c-a3ce-45af-9b93-81c736ac7eb3 · inbound

On Next-Token Prediction in LLMs: How End Goals Determine the Consistency of Decoding Algorithms cites this paper.

On Next-Token Prediction in LLMs: How End Goals Determine the Consistency of Decoding Algorithms Reasoning Bias of Next Token Prediction Training

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:05:57.943965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-15T21:05:57.739148Z digest=sha256:3d26923dea672b784254526ae2f8a147e06a375eec4044f311a513a0d701f253

Observation 291098ac-6303-4e3b-9c1c-f6473047bec8 · inbound

Weight-norm Criticality: A Mechanism for Loss Spikes Induced by the Normalization and Weight Decay cites this paper.

Weight-norm Criticality: A Mechanism for Loss Spikes Induced by the Normalization and Weight Decay Reasoning Bias of Next Token Prediction Training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T08:50:43.217179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T08:50:43.217179Z digest=sha256:135bad1918a7e04205b0c223781c4d11b9aabc04c805f9f778c6ae37de6a3253