Pith. sign in

Paper Citation Record · LEDGER

Reasoning Bias of Next Token Prediction Training

As of 10 August 2026, this Paper Citation Record lists 67 of 67 outbound references and 1 inbound Pith citation observation for arXiv:2502.02007.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.02007 v2

Coverage vector

measured 67 of 67 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T13:46:43.619470Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T08:50:43.217179Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

67 of 67 outbound references displayed

  • verified exact5
  • verified fuzzy23
  • unresolved39
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 908f4876-0b39-4468-b68b-25e6da6230ad · outbound

This paper cites Phi-4 Technical Report.

Reasoning Bias of Next Token Prediction Training Phi-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.282422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.282422Z digest=sha256:47ed4deae1fcf54cad9a2baf840d61e9466130ce1a9dd122a877244c0062bab4

Observation 0bc3db89-9f45-4529-a4eb-512b5c8a79a6 · outbound

This paper cites and Nagarajan, V.

Reasoning Bias of Next Token Prediction Training and Nagarajan, V

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.818716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.288710Z digest=sha256:4ec32695d256cb7bc3cae5a322f5bd1357cbb095eb84c4da454095ee9f00bb15

Observation 5a091145-20fd-4987-a9f7-37a6a77efe5d · outbound

This paper cites and Giryes, R.

Reasoning Bias of Next Token Prediction Training and Giryes, R

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.803082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.294502Z digest=sha256:201317ae19d1f528d005e3207511fa19525164099d827384f5ca4638e81f96c8

Observation 54745b23-35fb-4dec-9396-4dc895cbfc58 · outbound

This paper cites Multi-Step Deductive Reasoning Over Natural Language: An Empirical Study on Out-of-Distribution Generalisation.

Reasoning Bias of Next Token Prediction Training Multi-Step Deductive Reasoning Over Natural Language: An Empirical Study on Out-of-Distribution Generalisation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.299577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.299577Z digest=sha256:dbeef8b0d5e306719d88abc586d1db383d13286d5c326f3aa41d528ddef8671d

Observation b0fb012e-f872-44c1-b650-21a26278852c · outbound

This paper cites Understanding robustness of transformers for image classification.

Reasoning Bias of Next Token Prediction Training Understanding robustness of transformers for image classification

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.786239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.305137Z digest=sha256:23fbe9d27d16080c4e7bb5939042a80e3ce38b29cdc021675132ae54412d69b1

Observation b1e68e2b-b26c-43cb-b507-2af169b07173 · outbound

This paper cites R., Angeli, G., Potts, C., and Manning, C.

Reasoning Bias of Next Token Prediction Training R., Angeli, G., Potts, C., and Manning, C

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.310512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.310512Z digest=sha256:f825e72894e93da578c974fe797ce79f98e9d8d6743554e79eb83babd1b40bcf

Observation 64acb6d5-312d-4437-a06c-2ecbe4500664 · outbound

This paper cites Language Models are Few-Shot Learners.

Reasoning Bias of Next Token Prediction Training Language Models are Few-Shot Learners

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.316250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.316250Z digest=sha256:52c1cde37158aa5bde63f49053c2f7536226cf45fdc43c0623e5c511f0f74af3

Observation 26d82bc1-9b13-477a-b139-85f4720360ab · outbound

This paper cites Dropout as a low-rank regularizer for matrix factorization.

Reasoning Bias of Next Token Prediction Training Dropout as a low-rank regularizer for matrix factorization

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.770532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.321652Z digest=sha256:c9ce08563945894651caf69be1a27b11cdd97aa699bfdac723dc942f37e3c2db

Observation 72c42146-f7ca-464b-8eeb-13a91c8e4e1c · outbound

This paper cites Transformers as soft reasoners over language.

Reasoning Bias of Next Token Prediction Training Transformers as soft reasoners over language

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.326651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.326651Z digest=sha256:7b3f18a979352f885e5d421cb05094473632fa9d497cef6d02ff1609206cf449

Observation e1d7fe60-f30a-416f-a531-3f5520c0b2ca · outbound

This paper cites From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step.

Reasoning Bias of Next Token Prediction Training From Explicit CoT to Implicit CoT: Learning to Internalize CoT Step by Step

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.331535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.331535Z digest=sha256:3cf9dfaeaf0774f41ea7aeb5c7c3ae0cb6fa339d47616a99a6e086fc68739e4b

Observation ec8fa746-963f-4ee2-a080-7addc029617f · outbound

This paper cites Reducing Transformer Depth on Demand with Structured Dropout.

Reasoning Bias of Next Token Prediction Training Reducing Transformer Depth on Demand with Structured Dropout

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.336880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.336880Z digest=sha256:03c3bcaab1f6d28428ccf1c2075f8088e283e0744e97b56e99086cd5291f647a

Observation b3ed6204-64b0-475a-99a1-9b58e512403f · outbound

This paper cites and Tu, Y.

Reasoning Bias of Next Token Prediction Training and Tu, Y

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.754241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.342409Z digest=sha256:1cb9fb24deb53f70b7f6b7bbb29e12479e01664ff6b4bcb455cf289758eb1655

Observation 004c79a9-e648-4a5a-9026-951e34de47d9 · outbound

This paper cites Y., Roziere, B., Lopez-Paz, D., and Synnaeve, G.

Reasoning Bias of Next Token Prediction Training Y., Roziere, B., Lopez-Paz, D., and Synnaeve, G

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.739085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.347125Z digest=sha256:36d343611efa47a001f9e03b66d2f7ec5f3c6a62cde6ee3bbb43c3702057e0fc

Observation dcf418dc-aec7-462a-a395-17c6feb2b8d6 · outbound

This paper cites Training Large Language Models to Reason in a Continuous Latent Space.

Reasoning Bias of Next Token Prediction Training Training Large Language Models to Reason in a Continuous Latent Space

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.351846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.351846Z digest=sha256:a65dce0e090ff5e5272f75375f012ea28d7a15dcb45bfb3ddbea54cc64b7e997

Observation 0373d6dc-8f9b-411e-b6a5-e400175c010a · outbound

This paper cites A Law of Next-Token Prediction in Large Language Models.

Reasoning Bias of Next Token Prediction Training A Law of Next-Token Prediction in Large Language Models

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-09T13:46:44.306519Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.357172Z digest=sha256:616269c8f823813e5daa45170b78824c57c6694059714c9587d34be4bde447b6

Observation 5889224b-91c3-4f62-977f-ef673b0eab86 · outbound

This paper cites What Matters in Transformers? Not All Attention is Needed.

Reasoning Bias of Next Token Prediction Training What Matters in Transformers? Not All Attention is Needed

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.362438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.362438Z digest=sha256:f818f132f701487c0b70dde70391cf7bc3c844960d976755a2b355c8b96136e0

Observation 42335d35-d8bc-4683-8f70-55e86a492a40 · outbound

This paper cites Pretrained transformers improve out-of-distribution robustness.

Reasoning Bias of Next Token Prediction Training Pretrained transformers improve out-of-distribution robustness

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.367522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.367522Z digest=sha256:ba251d925cae60c62c66eba4ff9febe9530d42c966144bd940e9eb7e1608344f

Observation 80767613-99e7-4bfa-9773-0c646eae34da · outbound

This paper cites and Schmidhuber, J.

Reasoning Bias of Next Token Prediction Training and Schmidhuber, J

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.724058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.372397Z digest=sha256:c19093eb8da5e170a2ce1d6e3b51cf5eace882ddba8322a6f4cc036eb299d82a

Observation d3608fce-dd3f-4a81-94ca-1d129db31b3e · outbound

This paper cites NEFTune: Noisy Embeddings Improve Instruction Finetuning.

Reasoning Bias of Next Token Prediction Training NEFTune: Noisy Embeddings Improve Instruction Finetuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.378061Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.378061Z digest=sha256:64af371bb9b3794371e5cb6f9c5e2f2637473c0cc8ebad5a31be7a2219844f9e

Observation fdeedeed-ec45-46a3-90c1-5753574ac00b · outbound

This paper cites S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P.

Reasoning Bias of Next Token Prediction Training S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.383548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.383548Z digest=sha256:a1a4e88dd6cb6b73bc4c111d5cc6faf5a7c974d31967cd9fc221b52e5fd52c86

Observation d34e9dcb-e86e-48a9-8da7-8b7bdf6dbfce · outbound

This paper cites N., Hellmann, S., Morsey, M., Van Kleef, P., Auer, S., et al.

Reasoning Bias of Next Token Prediction Training N., Hellmann, S., Morsey, M., Van Kleef, P., Auer, S., et al

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.697997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.388908Z digest=sha256:664c8258b8383b7d74ad77947848f814dfc89bd12bfbc9c0b30285013c67cc9d

Observation a0bf94c7-c237-48c6-9fb2-05d1b59d2d50 · outbound

This paper cites J., Xing, E., and Caruana, R.

Reasoning Bias of Next Token Prediction Training J., Xing, E., and Caruana, R

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.681991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.393776Z digest=sha256:11406c6356a03c9ea3718d5008851bd69d956756c240c35f8d37f616b09c860d

Observation 599dc58c-32e0-4abc-9f01-e43c660afdad · outbound

This paper cites DropKey.

Reasoning Bias of Next Token Prediction Training DropKey

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.398762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.398762Z digest=sha256:b97e317f5c7d5c0cd982f962fd1d8d7602da546669a930885e596215c4adfdcf

Observation 045c0bc8-f000-49af-bbff-6e95ce11acf6 · outbound

This paper cites Challenging large language models with new tasks: A study on their adaptability and robustness.

Reasoning Bias of Next Token Prediction Training Challenging large language models with new tasks: A study on their adaptability and robustness

Reference 24

Resolution
verified exact
doi, observed 2026-08-09T13:46:43.800168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.403768Z digest=sha256:5c7d37c0f6a6dd2ad4047215b1e9a6c1f5ce28615b6c079716f490685cfdf20d

Observation ce9ae42e-ea01-4df8-91f7-653a262e3cbd · outbound

This paper cites Visualizing the loss landscape of neural nets.

Reasoning Bias of Next Token Prediction Training Visualizing the loss landscape of neural nets

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.665175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.408608Z digest=sha256:737b2e20ecf592ee33d1df7ed3d62546d364e23eb920558e36ae40e6eb007c6b

Observation 71f357b0-9897-4fb4-9bc3-24524f166e75 · outbound

This paper cites E., Singh Rawat, A., and Oymak, S.

Reasoning Bias of Next Token Prediction Training E., Singh Rawat, A., and Oymak, S

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.648043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.413206Z digest=sha256:b7969ac074966fd80162de5450c1f49acb2b7b2e0bf1ad3f20b93ceae1d1da92

Observation 638fbb91-2244-4cfb-bbf0-bd7c71d5aa05 · outbound

This paper cites Rho-1: Not All Tokens Are What You Need.

Reasoning Bias of Next Token Prediction Training Rho-1: Not All Tokens Are What You Need

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.418162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.418162Z digest=sha256:661257f225aa4c4b621152120a7e53ddb687dd456f6b39bbcc18cc4fb0a339ec

Observation bb035bc5-c676-4945-be4e-7ed6c8263df2 · outbound

This paper cites M., Li, Z., and Ma, T.

Reasoning Bias of Next Token Prediction Training M., Li, Z., and Ma, T

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.632706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.423209Z digest=sha256:ee251117106c8d0cd165fe8a02e11b4f36c11a606f2d24a0a009672b126360f4

Observation 405bb9c5-4ebb-4351-829a-b8ec1772e82f · outbound

This paper cites and Ying, L.

Reasoning Bias of Next Token Prediction Training and Ying, L

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.616824Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.428213Z digest=sha256:7cfed0fe7c139a82061a6688555e48393d55035797c3d15896baa5db473ad5db

Observation 2ea55aa7-4095-49a2-8ac6-d49fba158f4d · outbound

This paper cites Next-token prediction capacity: general upper bounds and a lower bound for transformers, 2024.

Reasoning Bias of Next Token Prediction Training Next-token prediction capacity: general upper bounds and a lower bound for transformers, 2024

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.433208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.433208Z digest=sha256:99f3f8175dd7650445daac70ae496919b24368757511b311009ea6056e12c4e8

Observation 3f18b2ec-3d31-4193-9860-5a52fe1ae1bb · outbound

This paper cites On the implicit bias of dropout.

Reasoning Bias of Next Token Prediction Training On the implicit bias of dropout

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.601307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.438140Z digest=sha256:32bc6e6355cbbe751295a92600e13fa393c900acbe2b880836522efba5b88b76

Observation 53d00bd4-9b33-4657-b951-5a431f8cafe1 · outbound

This paper cites Pretrained Transformers Do not Always Improve Robustness.

Reasoning Bias of Next Token Prediction Training Pretrained Transformers Do not Always Improve Robustness

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-09T13:46:44.127239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.443060Z digest=sha256:989f00ed00c120322e864544a6ef8b9bee3ff1d92624a15ef06948c5f9ef7642

Observation 128a33fd-553f-42f0-a89e-e447c3a3bb11 · outbound

This paper cites and Samwald, M.

Reasoning Bias of Next Token Prediction Training and Samwald, M

Reference 33

Resolution
verified exact
doi, observed 2026-08-09T13:46:43.783504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.448286Z digest=sha256:f628b812aef5d013953fd72eee3d3431de89957a4f979815c2bf95129bbef768

Observation 79db3ce0-125d-490e-9828-df83de123781 · outbound

This paper cites Power-law escape rate of SGD.

Reasoning Bias of Next Token Prediction Training Power-law escape rate of SGD

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.453263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.453263Z digest=sha256:a6f24b1b78387123f78a457f01498e673027b5ca5f95e23657063e4a506225c0

Observation 36c2b9ac-3e90-46a9-8e47-10777162b5a6 · outbound

This paper cites LogicInference: A New Dataset for Teaching Logical Inference to seq2seq Models.

Reasoning Bias of Next Token Prediction Training LogicInference: A New Dataset for Teaching Logical Inference to seq2seq Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.459652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.459652Z digest=sha256:eece06752b97672d8954f6ac711c34e14aaa2f198654e34709a21aa8abaf7901

Observation 991996d3-8a7c-4857-b014-46ed8f8781cc · outbound

This paper cites and Narasimhan, K.

Reasoning Bias of Next Token Prediction Training and Narasimhan, K

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.584285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.464879Z digest=sha256:8415474dca650e8a0d2d9bd32c4fc93ef7ac5f0d9d907b2a594cc73ff0fcc072

Observation 7cbfafc1-b7aa-4038-b36f-4521b0666e6d · outbound

This paper cites Language models are unsupervised multitask learners.

Reasoning Bias of Next Token Prediction Training Language models are unsupervised multitask learners

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.469944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.469944Z digest=sha256:95e96e7d98cf4762bc2ccb96883261a7d066b44a6a33437e84105d7db6e76a83

Observation d24bcdc6-d40d-449b-b2bd-fdd6d1898055 · outbound

This paper cites R obust LR : A diagnostic benchmark for evaluating logical robustness of deductive reasoners.

Reasoning Bias of Next Token Prediction Training R obust LR : A diagnostic benchmark for evaluating logical robustness of deductive reasoners

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.474845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.474845Z digest=sha256:5c07ae802a15970ee10b8ccd9d292d6df9f87aee3305287bce1ea99e9e6045de

Observation 1242e8b6-90c7-41f1-a5a8-e94f5255b8ed · outbound

This paper cites and He, H.

Reasoning Bias of Next Token Prediction Training and He, H

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.479930Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.479930Z digest=sha256:d86fa26aab43cbb3fbfca6b92718d4631e2ec8847645dc6489f4c0afad5637fd

Observation e0def47a-ad1f-4840-9711-164b47e0c829 · outbound

This paper cites Stepgame: A new benchmark for robust multi-hop spatial reasoning in texts.

Reasoning Bias of Next Token Prediction Training Stepgame: A new benchmark for robust multi-hop spatial reasoning in texts

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.485090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.485090Z digest=sha256:f7152119eb47cdaf6f697b3ef70173171c97d87538255eb60a9f241adbe7acf2

Observation 03ca9248-61bf-40b2-8d41-56918fad7417 · outbound

This paper cites an unresolved cited work.

Reasoning Bias of Next Token Prediction Training Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.490010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.490010Z digest=sha256:9cb0aae1bf33dd5b5215bdf8845c907f5b847e79b65824ad99bfc588b987fbbe

Observation 50ab1b23-a4ab-427b-86b2-c90363ac2245 · outbound

This paper cites P roof W riter: Generating implications, proofs, and abductive statements over natural language.

Reasoning Bias of Next Token Prediction Training P roof W riter: Generating implications, proofs, and abductive statements over natural language

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.495140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.495140Z digest=sha256:8713d6ed7addd3120efc90da3751fa982e0c6c97ce29de32144c8e801fa19241

Observation 2d6c2844-7d4b-412d-b39f-f2f929cb2d8d · outbound

This paper cites Memorisation versus generalisation in pre-trained language models.

Reasoning Bias of Next Token Prediction Training Memorisation versus generalisation in pre-trained language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.500250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.500250Z digest=sha256:d7844e3f6aa7611fe806873f3ec668fee96b71a92707229381b5b022981c12ec

Observation 4d0b59b2-752d-417e-a8e7-4396418c8023 · outbound

This paper cites Implicit optimization bias of next-token prediction in linear models.

Reasoning Bias of Next Token Prediction Training Implicit optimization bias of next-token prediction in linear models

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.547038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.505179Z digest=sha256:085d5c45ced85da8880a384f7d9ba70775be1753e7880769e3d5ede08ea9559a

Observation 9104245d-a585-445a-9e04-f9e73cdc3f99 · outbound

This paper cites An empirical study on robustness to spurious correlations using pre-trained language models.

Reasoning Bias of Next Token Prediction Training An empirical study on robustness to spurious correlations using pre-trained language models

Reference 45

Resolution
verified exact
doi, observed 2026-08-09T13:46:43.713990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.509922Z digest=sha256:f7fb746e69f61c87f4fb8c727bd0146241f90ec9cda544a8c60be21d32c73456

Observation 4d3139d1-8c38-47b7-a762-b88b682e82df · outbound

This paper cites L ogic A sker: Evaluating and improving the logical reasoning ability of large language models.

Reasoning Bias of Next Token Prediction Training L ogic A sker: Evaluating and improving the logical reasoning ability of large language models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.514920Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.514920Z digest=sha256:725a643412267c83efa96888444399073f2117f83f2e04c68b41739fda47c044

Observation 9d3daf19-287c-4f20-aa3b-8b685bcc4b2b · outbound

This paper cites Are large language models really robust to word-level perturbations? In Socially Responsible Language Modelling Research, 2023.

Reasoning Bias of Next Token Prediction Training Are large language models really robust to word-level perturbations? In Socially Responsible Language Modelling Research, 2023

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.529595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.519722Z digest=sha256:48f98f39946635ae7834a1f446986c3935b6807d5cafaef83fac31b2805fbce2

Observation cc99b161-f150-445d-a9f1-c18593e4f879 · outbound

This paper cites The implicit and explicit regularization effects of dropout.

Reasoning Bias of Next Token Prediction Training The implicit and explicit regularization effects of dropout

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.513937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.524447Z digest=sha256:5678afd81e8c6666ad39a4d24eefe3e01a2696ac6c98e27ced844e931404e1bd

Observation 0448aee8-c29c-46bb-825d-985cba3b0ada · outbound

This paper cites Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks.

Reasoning Bias of Next Token Prediction Training Towards AI-Complete Question Answering: A Set of Prerequisite Toy Tasks

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.529135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.529135Z digest=sha256:b6e94e6739a1d3f661d5bb594aaac8b770e48ffce67afb9c57a035d89f7893ad

Observation 7fcd35b3-2ada-4779-b3dc-6ee73103d0c0 · outbound

This paper cites On the noisy gradient descent that generalizes as sgd.

Reasoning Bias of Next Token Prediction Training On the noisy gradient descent that generalizes as sgd

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.496363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.534025Z digest=sha256:4b9a42c6d1dff59056fb68e75c11454997c27b9aea65980f15773e4dc3518886

Observation 21ebed3d-a2d9-431a-a33c-299eba61061f · outbound

This paper cites How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective.

Reasoning Bias of Next Token Prediction Training How sgd selects the global minima in over-parameterized learning: A dynamical stability perspective

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.480659Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.539227Z digest=sha256:acf76c8633fdeb7997df689a86d0b33fa61e2bced61ed19a43d8219f33c130a0

Observation e45537e8-832f-4ceb-9a8b-bfad031b41e3 · outbound

This paper cites UniDrop: A Simple yet Effective Technique to Improve Transformer without Extra Cost.

Reasoning Bias of Next Token Prediction Training UniDrop: A Simple yet Effective Technique to Improve Transformer without Extra Cost

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.544075Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.544075Z digest=sha256:45bf4568dbc261c2f630b51b844f4be8625e9ec0ec7b0ac13884a8c082b67c44

Observation 87d8e884-4966-49fa-aae4-9ba4027f9c3c · outbound

This paper cites D., and Potts, C.

Reasoning Bias of Next Token Prediction Training D., and Potts, C

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.549000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.549000Z digest=sha256:a9e6bcbce1bbc5ddda48ed43654ed8b5b25b79e516d65294da396794b5959d85

Observation 93be4192-acd6-4e66-9346-1ba25f5a8a5f · outbound

This paper cites A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima.

Reasoning Bias of Next Token Prediction Training A Diffusion Theory For Deep Learning Dynamics: Stochastic Gradient Descent Exponentially Favors Flat Minima

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.554031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.554031Z digest=sha256:65c26bbb55ac5e8f4f887ce2fe26ab7605dee9d6827f4e83e3107fcdf84c1a07

Observation ef9ae870-53ef-4629-90b6-895e2c03ed87 · outbound

This paper cites URL http://www.yelp.com/ dataset_challenge.

Reasoning Bias of Next Token Prediction Training URL http://www.yelp.com/ dataset_challenge

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.465135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.559095Z digest=sha256:56cea96fbd0ad06724d0a3c0aed0f02995d7d86e3a2f7cd5aff659b9dbb79065

Observation ffb8e00d-fa33-4783-b3d8-65f358a901f5 · outbound

This paper cites InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning.

Reasoning Bias of Next Token Prediction Training InternLM-Math: Open Math Large Language Models Toward Verifiable Reasoning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.563724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.563724Z digest=sha256:6d8ee19557f3faca3887d10130971363826d78cc6636364bf174e6b802939ca8

Observation c837e3f2-6754-489a-a334-e4e3b937e06f · outbound

This paper cites Natural language reasoning, a survey.

Reasoning Bias of Next Token Prediction Training Natural language reasoning, a survey

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.568890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.568890Z digest=sha256:d36f1fb78dcc278973c0786a94c59700e5643cbae1f8bc71e79c6cda49663d00

Observation 7752d722-275c-47e2-a23a-c363ea5c1b53 · outbound

This paper cites DropAttention: A Regularization Method for Fully-Connected Self-Attention Networks.

Reasoning Bias of Next Token Prediction Training DropAttention: A Regularization Method for Fully-Connected Self-Attention Networks

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.573794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.573794Z digest=sha256:65907cfa3e946b3dce213da511dad38babd308bd91c1a30625279c45ad98a13b

Observation 3b913036-4a56-4b15-87c9-92fa484ac6e2 · outbound

This paper cites Dropdim: A regularization method for transformer networks.

Reasoning Bias of Next Token Prediction Training Dropdim: A regularization method for transformer networks

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.579115Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.579115Z digest=sha256:952a4698874c8925b76a5336554242b7c1d88f1e311eb9431ae4d10dc87552e6

Observation 74d55690-5c9c-4951-b325-91bfac772083 · outbound

This paper cites H., Meng, T., Chang, K.-W., and Van den Broeck, G.

Reasoning Bias of Next Token Prediction Training H., Meng, T., Chang, K.-W., and Van den Broeck, G

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.583817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.583817Z digest=sha256:94a809bb27c48a16613db0c52cf0716c9234b5e7b298fde0880c5425094ebc77

Observation 33ebdccc-6c5d-42b1-8b4f-0d681e7bd3f0 · outbound

This paper cites and Xu, Z.-Q.

Reasoning Bias of Next Token Prediction Training and Xu, Z.-Q

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.588944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.588944Z digest=sha256:7438dee5fa96231f570da3c25291dd661fcd32824e4f60c5e33f4cb44fc89d4f

Observation f61a50ba-e91a-4684-979c-b1dcc1ae298b · outbound

This paper cites Stochastic Modified Equations and Dynamics of Dropout Algorithm.

Reasoning Bias of Next Token Prediction Training Stochastic Modified Equations and Dynamics of Dropout Algorithm

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.594201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.594201Z digest=sha256:4525ca1441d14269cf4ae3bb088b4a31278edc4cb9bfc5df9183705a70abab4e

Observation 98a53005-a589-4721-9d72-b65c5d036bc9 · outbound

This paper cites Initialization is Critical to Whether Transformers Fit Composite Functions by Reasoning or Memorizing.

Reasoning Bias of Next Token Prediction Training Initialization is Critical to Whether Transformers Fit Composite Functions by Reasoning or Memorizing

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.599362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.599362Z digest=sha256:5354b5c794d4929bb228e796e395219c5421bff67e3dfb97be12431a8a889d1e

Observation bc20519d-1dad-4b65-ae7c-c7d1f8a0e807 · outbound

This paper cites Anchor function: a type of benchmark functions for studying language models.

Reasoning Bias of Next Token Prediction Training Anchor function: a type of benchmark functions for studying language models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.604752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.604752Z digest=sha256:a67eb88bf5a590c917668d06b88178280b08efbc640b5dbe8a2d0d653438fe80

Observation ddbdd9be-60e5-4cbb-8e00-b098c371f2df · outbound

This paper cites Implicit geometry of next-token prediction: From language sparsity patterns to model representations.

Reasoning Bias of Next Token Prediction Training Implicit geometry of next-token prediction: From language sparsity patterns to model representations

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.440791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.609837Z digest=sha256:239c6fce71b490fdd9ebea34461c1948e4fcf3e5f9628e2d0fc147faf901b287

Observation d9675f5f-907b-4fcc-814f-96896e1c0ded · outbound

This paper cites Scheduled D rop H ead: A regularization method for transformer models.

Reasoning Bias of Next Token Prediction Training Scheduled D rop H ead: A regularization method for transformer models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-09T13:46:43.614551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T13:46:43.614551Z digest=sha256:2ef2fac840663fffc805ea2453506e8fefe687b47c2b5129d370060b74345a22

Observation 30af6bad-2d3a-4fcf-9b1f-dc57d4457767 · outbound

This paper cites The anisotropic noise in stochastic gradient descent: Its behavior of escaping from sharp minima and regularization effects.

Reasoning Bias of Next Token Prediction Training The anisotropic noise in stochastic gradient descent: Its behavior of escaping from sharp minima and regularization effects

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T13:46:44.424735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-09T13:46:43.619470Z digest=sha256:389f85664d13d423a01c2943337bd7bb2e73312f82f9d744041534bded9b3f68

Pith citing papers

Observation 291098ac-6303-4e3b-9c1c-f6473047bec8 · inbound

Weight-norm Criticality: A Mechanism for Loss Spikes Induced by the Normalization and Weight Decay cites this paper.

Weight-norm Criticality: A Mechanism for Loss Spikes Induced by the Normalization and Weight Decay Reasoning Bias of Next Token Prediction Training

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T08:50:43.217179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-01T08:50:43.217179Z digest=sha256:273e466607f89bd4f0b3018c015dd8e4c640747fe62b151eea356799c6da8cea