Pith. sign in

Paper Citation Record · LEDGER

An Analysis for Reasoning Bias of Language Models with Small Initialization

As of 10 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 6 inbound Pith citation observations for arXiv:2502.04375.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.04375 v2

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-09T05:28:52.279862Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:01:15.054309Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-22T20:05:04.830513Z

Reference resolution

60 of 60 outbound references displayed

  • verified exact3
  • verified fuzzy17
  • unresolved40
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e863a94a-9e15-4fd7-a0af-a7c4fd290b6a · outbound

This paper cites Phi-4 Technical Report.

An Analysis for Reasoning Bias of Language Models with Small Initialization Phi-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.220951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.220951Z digest=sha256:4a95289bca6e0d0fa5ae5f4371d317b5776f81ee4d9d3ba835a27bf8a48a4cf0

Observation 572a90c5-ba55-436e-9ae3-e26e705ce555 · outbound

This paper cites GPT-4 Technical Report.

An Analysis for Reasoning Bias of Language Models with Small Initialization GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.227174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.227174Z digest=sha256:3766a87ec44840c9bd4cd26af1aa3e22caa36605f440d99d6c026be1172437a1

Observation 7612f387-27eb-47fa-9cf5-e1ead065d700 · outbound

This paper cites S., Hu, W., Li, Z., Salakhutdinov, R.

An Analysis for Reasoning Bias of Language Models with Small Initialization S., Hu, W., Li, Z., Salakhutdinov, R

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.666924Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.232212Z digest=sha256:12de735014e6f124e82d534b75d7db1777981a11abf787be5f32da71bd3a6fa4

Observation bb436b54-89ab-415d-bcfc-f53260046733 · outbound

This paper cites Reflections after refereeing papers for nips.

An Analysis for Reasoning Bias of Language Models with Small Initialization Reflections after refereeing papers for nips

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.616478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.236969Z digest=sha256:b12d690100cd7e3f7116e2dddc5d1640064edbb2bc6e7e5165b795fe38077432

Observation 01915d98-d04e-4ebe-9319-62ae92630134 · outbound

This paper cites and Rathie, P.

An Analysis for Reasoning Bias of Language Models with Small Initialization and Rathie, P

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.525810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.241569Z digest=sha256:0bc62e3352fe759117bdcdd665bc580867b5d8be3e8a7bd79bf4eeacd08e8506

Observation 7280dc0a-53e8-4036-8f73-99b8c833acf1 · outbound

This paper cites an unresolved cited work.

An Analysis for Reasoning Bias of Language Models with Small Initialization Unresolved cited work

Reference 6

Resolution
verified exact
doi, observed 2026-08-09T05:28:52.349960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.246643Z digest=sha256:da6c3aa628970d31d3a40bfb7777b4ad5427cc5f2e69c72ea1080590cc5bbd08

Observation 0d67c68a-3e61-4241-829f-9a3692893d1b · outbound

This paper cites and Bach, F.

An Analysis for Reasoning Bias of Language Models with Small Initialization and Bach, F

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.464356Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.252040Z digest=sha256:187eed90eadb0ac71e4e5e8286dbc5b81bbb33f1d01a213fc283e13c9c72b9dd

Observation 2ed755b8-7569-4ff0-89d1-601a80a9f9b5 · outbound

This paper cites Faithful Reasoning Using Large Language Models.

An Analysis for Reasoning Bias of Language Models with Small Initialization Faithful Reasoning Using Large Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.257146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.257146Z digest=sha256:dfe05474c432228336da8f32420989a1422dbcabff52ffe4f5691593349ad60f

Observation 54507407-694a-4d15-9c72-0177863cc7de · outbound

This paper cites Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning.

An Analysis for Reasoning Bias of Language Models with Small Initialization Selection-Inference: Exploiting Large Language Models for Interpretable Logical Reasoning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.262280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.262280Z digest=sha256:cd02cf030fedef378a5ed159a591ebc95a3f5644958c70dce5d98cd0bdd52f9e

Observation ece0bfbf-5f29-4af2-8523-40b8fc116d48 · outbound

This paper cites The Neural Data Router: Adaptive Control Flow in Transformers Improves Systematic Generalization.

An Analysis for Reasoning Bias of Language Models with Small Initialization The Neural Data Router: Adaptive Control Flow in Transformers Improves Systematic Generalization

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.278903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.278903Z digest=sha256:5d8a4c6a53ce1c626e6547be2d1057ee900f82a404decce17d7026cd87a87e2b

Observation 36299e03-8cb4-4498-8d23-30005147fb43 · outbound

This paper cites CTL++: Evaluating Generalization on Never-Seen Compositional Patterns of Known Functions, and Compatibility of Neural Representations.

An Analysis for Reasoning Bias of Language Models with Small Initialization CTL++: Evaluating Generalization on Never-Seen Compositional Patterns of Known Functions, and Compatibility of Neural Representations

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.310278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.310278Z digest=sha256:e3dee582f5846554668958b5cbd5e2757b9443453ad1d89bc14e201f256e4d39

Observation 8344b049-1e7f-4a95-9e9d-78aca8964860 · outbound

This paper cites L., Jiang, L., Lin, B.

An Analysis for Reasoning Bias of Language Models with Small Initialization L., Jiang, L., Lin, B

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.343200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.343200Z digest=sha256:8d07f0b58af6753a8939207cab386184cdc877bb3cb0b3d82a06a9db65bf0b65

Observation 5c83dc4b-a778-4551-9439-c7bc29679656 · outbound

This paper cites A comparative analysis of optimization and generalization properties of two-layer neural network and random feature models under gradient descent dynamics.

An Analysis for Reasoning Bias of Language Models with Small Initialization A comparative analysis of optimization and generalization properties of two-layer neural network and random feature models under gradient descent dynamics

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.422368Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.375582Z digest=sha256:fc7c18fcc5502480aa8355084a0d032629c7a7d6754157eacbdd2d0876fcc33f

Observation 73bb8b33-8326-404f-a4e5-7a52f025a538 · outbound

This paper cites TinyStories: How Small Can Language Models Be and Still Speak Coherent English?.

An Analysis for Reasoning Bias of Language Models with Small Initialization TinyStories: How Small Can Language Models Be and Still Speak Coherent English?

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.403494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.403494Z digest=sha256:95bc166e37a4866931da2a8310b105b49612d6dbaca9010daf3ab361f7bfdf13

Observation f571c1fe-67d4-478d-84f4-639840cf1a11 · outbound

This paper cites How does gpt obtain its ability? tracing emergent abilities of language models to their sources.

An Analysis for Reasoning Bias of Language Models with Small Initialization How does gpt obtain its ability? tracing emergent abilities of language models to their sources

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.406308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.442629Z digest=sha256:b593e6b36dee340a71e6ff34001081762af5ba41e7df5a3d24d7b1d7f1f0140b

Observation 671aca84-d7a9-4b9b-b25c-0c7a3538cda2 · outbound

This paper cites Delving deep into rectifiers: Surpassing human-level performance on imagenet classification.

An Analysis for Reasoning Bias of Language Models with Small Initialization Delving deep into rectifiers: Surpassing human-level performance on imagenet classification

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.501648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.501648Z digest=sha256:8e20917985c486a8c0fe6422acdfc8811f27acd67029991321ccc00b297f4777

Observation bce18684-9761-4eff-b1b7-d8fc826d7f09 · outbound

This paper cites S., Perez, F., Ba, J., and Volkovs, M.

An Analysis for Reasoning Bias of Language Models with Small Initialization S., Perez, F., Ba, J., and Volkovs, M

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.379227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.547458Z digest=sha256:de9f8542e93bc929ee8405bd5f2b78a00727fd623a8f6c5866d70afeb684ac45

Observation eefeb4ba-6f10-4d26-85c3-dcebc66f833d · outbound

This paper cites Learning compositionally through attentive guidance.

An Analysis for Reasoning Bias of Language Models with Small Initialization Learning compositionally through attentive guidance

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-09T05:28:52.745846Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.552290Z digest=sha256:a0a2aad4672b57caacb1fe85faf2292d1b0e73e126f45792d4c38ae44ff0dd7d

Observation 0fa9d1e1-9024-4aa2-bee3-4c2504ddd66c · outbound

This paper cites Neural Tangent Kernel : Convergence and Generalization in Neural Networks.

An Analysis for Reasoning Bias of Language Models with Small Initialization Neural Tangent Kernel : Convergence and Generalization in Neural Networks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.363554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.557359Z digest=sha256:1ea8f5564e429c05e0745399b4945f83ab7fff9d74a7776c8cd08d138c9fe82b

Observation 718e5914-bfe6-4e6b-88e0-52a57fd03c0d · outbound

This paper cites B., and M \"u ller, K.

An Analysis for Reasoning Bias of Language Models with Small Initialization B., and M \"u ller, K

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.562035Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.562035Z digest=sha256:8a6d9f34b08bae5259f5adbbc02f070e12041d46438682f21d29d62de18334ce

Observation d68ae65b-4b0e-4cc9-9b77-c09de6140f3c · outbound

This paper cites Break It Down: Evidence for Structural Compositionality in Neural Networks.

An Analysis for Reasoning Bias of Language Models with Small Initialization Break It Down: Evidence for Structural Compositionality in Neural Networks

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.566749Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.566749Z digest=sha256:487f5c250d2d34621b5ce6f8f0be41a69a2ce1a7f57fdcdcbb752458b71c67db

Observation f9f4ffcc-a086-4b30-9e69-316ee80546cf · outbound

This paper cites Not all tokens are what you need for pretraining.

An Analysis for Reasoning Bias of Language Models with Small Initialization Not all tokens are what you need for pretraining

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.571389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.571389Z digest=sha256:27a65784640042bc28d18b67483891dedd17a9f973f5255526043939b35562ab

Observation 7d20637e-1d00-4009-a74c-95afaac33a48 · outbound

This paper cites DeepSeek-V3 Technical Report.

An Analysis for Reasoning Bias of Language Models with Small Initialization DeepSeek-V3 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.575637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.575637Z digest=sha256:73b2576af914c89ad8e9eb5295b9bda5803be526af935cc73148dfc14991123e

Observation 276bfca0-e664-4254-95a6-d80ad59018ca · outbound

This paper cites Transformers Learn Shortcuts to Automata.

An Analysis for Reasoning Bias of Language Models with Small Initialization Transformers Learn Shortcuts to Automata

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.580568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.580568Z digest=sha256:9486795e303b0460b95403c1daa4208a0ec185cf5b773a564705c2f83282469b

Observation a35bad4b-12e6-4bb5-9f02-49f411c4d22c · outbound

This paper cites Understanding the Difficulty of Training Transformers.

An Analysis for Reasoning Bias of Language Models with Small Initialization Understanding the Difficulty of Training Transformers

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.585595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.585595Z digest=sha256:d3ddecc3f0cddbffa823a5daaf50907f7f4f4bc4e4c5e671a7c678259cf7bdf1

Observation 17c8b97d-acaf-48b7-b6b2-195ea4becdf4 · outbound

This paper cites J., Ma, Z., and Zhang, Y.

An Analysis for Reasoning Bias of Language Models with Small Initialization J., Ma, Z., and Zhang, Y

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.319975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.589800Z digest=sha256:465753bd2ee9cd27c36b0f9b6d8c823bb12ef55ab52bf083bc506d9da9b08b43

Observation c396f774-5388-4fc7-85ab-05341015456f · outbound

This paper cites an unresolved cited work.

An Analysis for Reasoning Bias of Language Models with Small Initialization Unresolved cited work

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.594177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.594177Z digest=sha256:985cfa96e6eb7e16da7c4235145877fef0c5bece1b5f1e7ae5bad5b5418175f9

Observation 318fdf09-8ad9-422b-8499-08325e10cf91 · outbound

This paper cites A mean field view of the landscape of two-layer neural networks.

An Analysis for Reasoning Bias of Language Models with Small Initialization A mean field view of the landscape of two-layer neural networks

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.598613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.598613Z digest=sha256:83d71e35a4d1e93f664b9e7570a74bf891f158cc5937e733e281a0061895da5b

Observation e96eff98-44a6-4c7b-af36-3131d9171e0c · outbound

This paper cites Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic Task.

An Analysis for Reasoning Bias of Language Models with Small Initialization Compositional Abilities Emerge Multiplicatively: Exploring Diffusion Models on a Synthetic Task

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.603410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.603410Z digest=sha256:42b2a79dd138e4dcb9d2a4db1375074f136257e22ce94ba0f326fefa9f8244ef

Observation 18875ebd-5a96-486f-883c-61a8d0949742 · outbound

This paper cites Language models are unsupervised multitask learners.

An Analysis for Reasoning Bias of Language Models with Small Initialization Language models are unsupervised multitask learners

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.620847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.620847Z digest=sha256:1757cb89698010cae74b73e01f2ff64892d7f37187db761bf9b2c0eb6b81edfe

Observation 857d318c-56eb-4190-b762-acc542225018 · outbound

This paper cites Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks.

An Analysis for Reasoning Bias of Language Models with Small Initialization Compositional Capabilities of Autoregressive Transformers: A Study on Synthetic, Interpretable Tasks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.641065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.641065Z digest=sha256:1686be201bc7ed4020ec803de901c7c0a40f3cbd2493e11039d49fee1cb0fb90

Observation e81b832a-d782-4d54-97f1-9355dea9d994 · outbound

This paper cites and Vanden-Eijnden, E.

An Analysis for Reasoning Bias of Language Models with Small Initialization and Vanden-Eijnden, E

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.285195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.700823Z digest=sha256:bbdfe74b19a3126218f6f271c9856d69c8352f6f6d81805a47d927e7350c6055

Observation 82d661f1-f854-4d06-8bfb-bd376c14dbac · outbound

This paper cites and He, H.

An Analysis for Reasoning Bias of Language Models with Small Initialization and He, H

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.269539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.738043Z digest=sha256:0e6d1219e5076584660a7913033ce79cc8e5bb42d1563fbbb49fb08eb829b53b

Observation 485a5d23-6b5d-4d98-8247-abf04b726b2d · outbound

This paper cites and Spiliopoulos, K.

An Analysis for Reasoning Bias of Language Models with Small Initialization and Spiliopoulos, K

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.757701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.757701Z digest=sha256:5aa5a22045802f11b7944255f7329f1d74fd52bf6c4071deacf12e32486a6ae6

Observation 6756689a-d32d-49bd-9937-a4f787d68eb8 · outbound

This paper cites Neurocompositional computing: From the central paradox of cognition to a new generation of ai systems.

An Analysis for Reasoning Bias of Language Models with Small Initialization Neurocompositional computing: From the central paradox of cognition to a new generation of ai systems

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.255005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.784618Z digest=sha256:5aef3a22811d0797163609fdd5742eadc1862810502b4ee53a060cad8c1531a0

Observation 42c48d34-487b-4c4c-a673-849b591f0d9d · outbound

This paper cites Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models.

An Analysis for Reasoning Bias of Language Models with Small Initialization Beyond the Imitation Game: Quantifying and extrapolating the capabilities of language models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.788616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.788616Z digest=sha256:ab0ce9b1bd6f1112994a5674b943ad976285564b55499621e9bb211273d6b81b

Observation 58695419-ce1f-4005-9824-f486b8936afa · outbound

This paper cites and Kolter, J.

An Analysis for Reasoning Bias of Language Models with Small Initialization and Kolter, J

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.239723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.793071Z digest=sha256:1fc74da46917e0b9230fcd7188eccfbd4f32b3d75b2263c69142821a09763f05

Observation 4569c100-5d1e-4abf-bf1c-3329d336f44d · outbound

This paper cites Deepnet: Scaling transformers to 1,000 layers.

An Analysis for Reasoning Bias of Language Models with Small Initialization Deepnet: Scaling transformers to 1,000 layers

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.224112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.797195Z digest=sha256:5bb9c1cf3d492fcd514bfe3307e2ff44ae8f89eb59ea9a26af1e761d410dc16f

Observation bd5896f8-8a41-4384-8067-67e48f6e9198 · outbound

This paper cites Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning.

An Analysis for Reasoning Bias of Language Models with Small Initialization Label Words are Anchors: An Information Flow Perspective for Understanding In-Context Learning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.801401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.801401Z digest=sha256:a88bc661ac6251d6a69f18d6c8aeb4cf3251b303675588689367dfe3e922a455

Observation be13f3e7-9ac0-4b58-9233-54524425ebdd · outbound

This paper cites Improving Generalization and Convergence by Enhancing Implicit Regularization.

An Analysis for Reasoning Bias of Language Models with Small Initialization Improving Generalization and Convergence by Enhancing Implicit Regularization

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-09T05:28:52.592722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.806335Z digest=sha256:8e73092c32f77574c0fdcf601709add05644efdedb2b2453d48b2b46519b6cce

Observation ed352f33-948f-4aad-a02e-453e9df0eec5 · outbound

This paper cites Understanding the Expressive Power and Mechanisms of Transformer for Sequence Modeling.

An Analysis for Reasoning Bias of Language Models with Small Initialization Understanding the Expressive Power and Mechanisms of Transformer for Sequence Modeling

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.810922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.810922Z digest=sha256:241f497a5c8f33f76d2044db7b371899ffd0eaf2c3041c62679b7d3b23677b23

Observation 39617b73-84ee-4467-93fb-da96bb31208f · outbound

This paper cites Understanding the Language Model to Solve the Symbolic Multi-Step Reasoning Problem from the Perspective of Buffer Mechanism.

An Analysis for Reasoning Bias of Language Models with Small Initialization Understanding the Language Model to Solve the Symbolic Multi-Step Reasoning Problem from the Perspective of Buffer Mechanism

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.815997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.815997Z digest=sha256:a0af858ed20f7b0152842c188bec8bfd2e54c1ebddce4a289827974a7f8e283b

Observation cd478da5-6f8d-4d05-b085-6e5dad8133da · outbound

This paper cites H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W.

An Analysis for Reasoning Bias of Language Models with Small Initialization H., Hashimoto, T., Vinyals, O., Liang, P., Dean, J., and Fedus, W

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.199486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.820681Z digest=sha256:99e42ae3b61b879294724800f6aa004ec060e967082b5887378b7f4b995b0edd

Observation b2f1df8b-9f3c-4c1d-b4b7-40cc19ef6a2f · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

An Analysis for Reasoning Bias of Language Models with Small Initialization Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.824742Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.824742Z digest=sha256:604602f7c52155733529b7c018a6b6fdee15024018984548e755e2194dd84138

Observation fb5e145f-64c4-4f27-bd06-d99b826300bc · outbound

This paper cites Gradient Dynamics of Shallow Univariate ReLU Networks.

An Analysis for Reasoning Bias of Language Models with Small Initialization Gradient Dynamics of Shallow Univariate ReLU Networks

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.829313Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.829313Z digest=sha256:52a01b8f5d2e455750941d08d57c1ab2394faf8c9f92adc00b3308c656b3ea2c

Observation b8ea9d09-8e59-40d0-912b-10b067bfa388 · outbound

This paper cites An overview of condensation phenomenon in deep learning.

An Analysis for Reasoning Bias of Language Models with Small Initialization An overview of condensation phenomenon in deep learning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.834240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.834240Z digest=sha256:1d14bce7b10f0c9c0696c60aaa786daa6a0ba50680173067b0470afcd0d13263

Observation 54c580f6-6c77-4e44-9636-09a75b1d5d17 · outbound

This paper cites Do Vision-Language Pretrained Models Learn Composable Primitive Concepts?.

An Analysis for Reasoning Bias of Language Models with Small Initialization Do Vision-Language Pretrained Models Learn Composable Primitive Concepts?

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.838929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.838929Z digest=sha256:051ca9f1cd4237c4f64f5913b1afe7c7d4262c797fc6b1556fb95b53af60424d

Observation 172990d3-9470-4d19-b1b3-d190604287c2 · outbound

This paper cites Improving Deep Transformer with Depth-Scaled Initialization and Merged Attention.

An Analysis for Reasoning Bias of Language Models with Small Initialization Improving Deep Transformer with Depth-Scaled Initialization and Merged Attention

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.843650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.843650Z digest=sha256:c035360c02efe23e8db17adde26f7af4dfc05847c5ec1b9f491e0eaa51a1e2c2

Observation fe9d802e-28cb-41b2-a937-ba971030ca36 · outbound

This paper cites Understanding deep learning requires rethinking generalization.

An Analysis for Reasoning Bias of Language Models with Small Initialization Understanding deep learning requires rethinking generalization

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.848843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.848843Z digest=sha256:b2e7582a787f1628f6da5c4ff540e94d989a8fca1152f84c3bb9fb5858e7c38e

Observation 38bcef4a-6788-4061-a647-f77c37aae9d4 · outbound

This paper cites A type of generalization error induced by initialization in deep neural networks.

An Analysis for Reasoning Bias of Language Models with Small Initialization A type of generalization error induced by initialization in deep neural networks

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.853798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.853798Z digest=sha256:bb9432677189b9ea16d0ec1df20a02d76e449d9436d54346f4363337e4177666

Observation add21f00-2939-48fa-90f9-e7ad69b09681 · outbound

This paper cites Linear Stability Hypothesis and Rank Stratification for Nonlinear Models.

An Analysis for Reasoning Bias of Language Models with Small Initialization Linear Stability Hypothesis and Rank Stratification for Nonlinear Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.858900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.858900Z digest=sha256:5ec2039a3f2f9b556f59dc93f8d10e88efe4cc780a7245dc134136ff84ef6cf9

Observation f822a794-272e-45e8-9b20-e0794a811df3 · outbound

This paper cites Loss Spike in Training Neural Networks.

An Analysis for Reasoning Bias of Language Models with Small Initialization Loss Spike in Training Neural Networks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:51.863442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:51.863442Z digest=sha256:8206666dc131d68e85e54a5586f272220836430ffbca6c1779155d0a962a52e3

Observation 1c412bb4-6094-4512-b027-703e74fbdc37 · outbound

This paper cites and Xu, Z.-Q.

An Analysis for Reasoning Bias of Language Models with Small Initialization and Xu, Z.-Q

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:53.141956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:28:51.928943Z digest=sha256:fc818dcb1d6307fc4cc41309cfdc34ea7dd3e13fe720436098f0bcc4e79682ce

Observation a52381c9-da47-499e-8db9-e5007416eeac · outbound

This paper cites Stochastic Modified Equations and Dynamics of Dropout Algorithm.

An Analysis for Reasoning Bias of Language Models with Small Initialization Stochastic Modified Equations and Dynamics of Dropout Algorithm

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:52.046608Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:52.046608Z digest=sha256:c29e72515947ed7d542282182cc0e45ad33dbfcaf888a78b0f6f138f1bbc57d4

Observation ec5d6f10-a553-4fa8-b59c-f0dcbcacc58d · outbound

This paper cites an unresolved cited work.

An Analysis for Reasoning Bias of Language Models with Small Initialization Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-09T05:28:53.029678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:28:52.153482Z digest=sha256:46a219ba8fde406d29aa5ec9754b3e1c8ec061a1f8a958fb1456d27a77825617

Observation ef590fc0-5244-4723-aa78-0af3468e01e1 · outbound

This paper cites Anchor function: a type of benchmark functions for studying language models.

An Analysis for Reasoning Bias of Language Models with Small Initialization Anchor function: a type of benchmark functions for studying language models

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:52.234294Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:52.234294Z digest=sha256:4d62f518bea7f7e41cca0dcd84134bd2a62f2aaf5f03ecccab66f391465cf352

Observation f2c555c6-f9b4-4532-8140-f292ba9d8258 · outbound

This paper cites Complexity Control Facilitates Reasoning-Based Compositional Generalization in Transformers.

An Analysis for Reasoning Bias of Language Models with Small Initialization Complexity Control Facilitates Reasoning-Based Compositional Generalization in Transformers

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:52.265267Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:52.265267Z digest=sha256:24ba41f3bb16cc74853ef333d272bc51c1d52eda288a7ed23e7d57bb1b80f3e2

Observation d9937e27-fdfb-49b8-bc92-a1b7cf3b48f5 · outbound

This paper cites an unresolved cited work.

An Analysis for Reasoning Bias of Language Models with Small Initialization Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-09T05:28:52.931262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:28:52.270541Z digest=sha256:303eabb7f99ac3883a1fe6a845d008a7bfe1f34b8c2aabb63f7320220e0e3a72

Observation 0c30e867-4b53-4098-b86a-c8789600bfe7 · outbound

This paper cites R., and Goldstein, T.

An Analysis for Reasoning Bias of Language Models with Small Initialization R., and Goldstein, T

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-09T05:28:52.897390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=arxiv_source observed=2026-08-09T05:28:52.275817Z digest=sha256:518ad239afa84ce395c245dd512aa2129d1d20841d3f5fda89b6a862de5f65bf

Observation 2bf11d20-18e9-4d4b-9d59-2321367ceb6a · outbound

This paper cites write newline.

An Analysis for Reasoning Bias of Language Models with Small Initialization write newline

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-09T05:28:52.279862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T05:28:52.279862Z digest=sha256:0e472444c0327bdb36377c6120bd395482879258dcded06957e1d1c3351fc3ef

Pith citing papers

Observation 66fefe0f-4f93-47e8-9afe-9276c46157ab · inbound

An overview of condensation phenomenon in deep learning cites this paper.

An overview of condensation phenomenon in deep learning An Analysis for Reasoning Bias of Language Models with Small Initialization

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-22T20:05:04.833173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-22T20:02:41.105786Z digest=sha256:1857ce61d8fc5eaa9e9439a5b95536c1365eb3bf9fd91415e9e2e2be542aef3d

Observation b8fc0767-8d84-497a-abfa-5756056ee175 · inbound

Scalable Complexity Control Facilitates Reasoning Ability of LLMs cites this paper.

Scalable Complexity Control Facilitates Reasoning Ability of LLMs An Analysis for Reasoning Bias of Language Models with Small Initialization

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T13:01:15.054309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:01:15.054309Z digest=sha256:38d704497e224bb03a517a9460cbf2262b646cc36b5dc6f67856ec212bc7c07f

Observation 6ec5b63b-4bbc-4e57-864d-1447d941c7f6 · inbound

Unveiling the Mechanisms of Multi-Hop Reasoning in Transformers via Identity Bridge cites this paper.

Unveiling the Mechanisms of Multi-Hop Reasoning in Transformers via Identity Bridge An Analysis for Reasoning Bias of Language Models with Small Initialization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T13:54:18.723775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T13:54:18.723775Z digest=sha256:bc272b202b79a5d3f870561e2d1f4d867ee9cf756be7ad1a24fb481080cb40d6

Observation a616f226-79c2-4b59-a397-8de0f7fbb763 · inbound

How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability cites this paper.

How Do Transformers Learn to Associate Tokens: Gradient Leading Terms Bring Mechanistic Interpretability An Analysis for Reasoning Bias of Language Models with Small Initialization

Reference 20

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:20:52.679420Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-16T11:20:33.400885Z digest=sha256:514a230c083f95827d6285d007d101ae93b634b97b96b644f50965c9ae447adc

Observation 1c17a3da-46b0-4515-aaf6-d2123e7139f9 · inbound

Understanding LoRA as Knowledge Memory: An Empirical Analysis cites this paper.

Understanding LoRA as Knowledge Memory: An Empirical Analysis An Analysis for Reasoning Bias of Language Models with Small Initialization

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-15T18:10:13.250276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-15T18:09:54.899039Z digest=sha256:b8c2edf510af4782abea7fb43a53ea2b848768c037927286d3dab8e71dd416c3

Observation 978df83f-d197-48a9-84d2-e4b6899c320e · inbound

Understanding LoRA as Knowledge Memory: An Empirical Analysis cites this paper.

Understanding LoRA as Knowledge Memory: An Empirical Analysis An Analysis for Reasoning Bias of Language Models with Small Initialization

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T19:49:44.634914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T19:49:44.634914Z digest=sha256:df23f0e8221ba9110f3939b2db7bedbdee337ec01e92c8418ce4ab57dd279364