Pith. sign in

Paper Citation Record · LEDGER

Understanding Emergent Abilities of Language Models from the Loss Perspective

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 13 inbound Pith citation observations for arXiv:2403.15796.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.15796 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 13 of 13 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-09T14:27:57.423803Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c31e2163-919e-4684-958e-f1391d816486 · inbound

MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies cites this paper.

MiniCPM: Unveiling the Potential of Small Language Models with Scalable Training Strategies Understanding Emergent Abilities of Language Models from the Loss Perspective

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T18:00:53.457262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T18:00:53.389420Z digest=sha256:66a9df5efe8d4f3fb8c556bef62f7fdb17f53ba067b407cc3ffd42bc3294d371

Observation e84ceb55-0f61-4ce1-b195-7ba48fdcab56 · inbound

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging cites this paper.

Soup-of-Experts: Pretraining Specialist Models via Parameters Averaging Understanding Emergent Abilities of Language Models from the Loss Perspective

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-09T14:27:57.423803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T14:27:57.423803Z digest=sha256:bae383d8976e9a1084e091552bcb746f0c71d5f751f09b68f4c758a24f1c3dec

Observation 3a6eefdf-0aad-4e04-a859-e9f4de283a6b · inbound

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model cites this paper.

SmolLM2: When Smol Goes Big -- Data-Centric Training of a Small Language Model Understanding Emergent Abilities of Language Models from the Loss Perspective

Reference 168

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T17:30:02.937523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-13T17:30:02.803757Z digest=sha256:8b727ec8326e62ff94d8a69b044b986f85ce0e134267dca9d9632b2930c44e15

Observation 1b5153ba-a9e3-4f99-8924-8726896eb432 · inbound

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws cites this paper.

LLMs on the Line: Data Determines Loss-to-Loss Scaling Laws Understanding Emergent Abilities of Language Models from the Loss Perspective

Reference 10

Resolution
metadata mismatch
arxiv_id, observed 2026-05-23T02:52:27.109203Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-23T02:47:37.492619Z digest=sha256:697374d65a197692ae560f3fdf055ff3ef881e1a882f92a7592a54fe3b20dd2b

Observation 8dae9bd0-9a5c-419a-ba21-37e4a47687ba · inbound

Capability Salience Vector: Fine-grained Alignment of Loss and Capabilities for Downstream Task Scaling Law cites this paper.

Capability Salience Vector: Fine-grained Alignment of Loss and Capabilities for Downstream Task Scaling Law Understanding Emergent Abilities of Language Models from the Loss Perspective

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T00:42:00.296033Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T00:42:00.296033Z digest=sha256:3b5303af3d6bc626f058689a2cacd958e7b26c7da380437ff056c95d440ce66f

Observation d7fbd434-b4e6-48a5-9caf-54111cd0313b · inbound

AMix-1: A Pathway to Test-Time Scalable Protein Foundation Model cites this paper.

AMix-1: A Pathway to Test-Time Scalable Protein Foundation Model Understanding Emergent Abilities of Language Models from the Loss Perspective

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T18:16:47.823217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:16:47.823217Z digest=sha256:d439435d4bf64bfbcf58fae7ea51df470cc69004f93ec935f4228b6a31380672

Observation 743a70bc-e96e-4fe1-9ae3-8f5d5265b47f · inbound

Feature Learning in Linear-Width Two-Layer Networks: Two vs. One Step of Gradient Descent cites this paper.

Feature Learning in Linear-Width Two-Layer Networks: Two vs. One Step of Gradient Descent Understanding Emergent Abilities of Language Models from the Loss Perspective

Reference 213

Resolution
verified exact
arxiv_id, observed 2026-05-20T01:32:56.072235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-20T01:29:14.555216Z digest=sha256:66e5d7d5570ca436a8e30f8ff048e019138e8492cc70f05c0bc4f1ec23204f5f

Observation baebb411-75cc-44f2-bddd-99bcf941aa9b · inbound

Feature Learning in Linear-Width Two-Layer Networks: Two vs. One Step of Gradient Descent cites this paper.

Feature Learning in Linear-Width Two-Layer Networks: Two vs. One Step of Gradient Descent Understanding Emergent Abilities of Language Models from the Loss Perspective

Reference 213

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:40:24.962878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-25T06:39:16.246591Z digest=sha256:c7956d96655a715b1f3ebaa8fc4872f84b32e3c2a7aa2569f4f48de7dfccc1fd

Observation f42199a4-1035-428f-9c13-99fc0995c091 · inbound

The Thermodynamic Costs of Simple Linear Regression cites this paper.

The Thermodynamic Costs of Simple Linear Regression Understanding Emergent Abilities of Language Models from the Loss Perspective

Reference 69

Resolution
metadata mismatch
arxiv_id, observed 2026-05-20T07:03:23.237151Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T07:03:21.987679Z digest=sha256:cc6fd98fc4a0339771169bac7ff65314f617602d22b6d9953f0addd97420af4d

Observation 69de36b4-5800-460a-810d-8f6305cf4edb · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning Understanding Emergent Abilities of Language Models from the Loss Perspective

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-04T09:59:44.688100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T09:19:50.623741Z digest=sha256:429a22809bd2b64bd29cf6101d2eb6245d80573e0d483a2b4deeda8bd1247274

Observation 527d4801-de94-4f6e-bbff-b3961bfd114c · inbound

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning cites this paper.

SingGuard: A Policy-Adaptive Multimodal LLM Guardrail with Dynamic Reasoning Understanding Emergent Abilities of Language Models from the Loss Perspective

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-01T18:55:59.667241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-29T01:18:19.195007Z digest=sha256:3ccca6a584b4ee1774becb0db6f00d14cb9450e946606da8957aa0cd45e4b750

Observation 586b36fa-c9c2-4a24-a13e-89909a3a01da · inbound

When transformers learn "impossible" languages, what do they learn? cites this paper.

When transformers learn "impossible" languages, what do they learn? Understanding Emergent Abilities of Language Models from the Loss Perspective

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-01T02:15:14.469922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-07-01T02:13:58.839175Z digest=sha256:b425f7cd53bb5857f2e2bbf08989ad225d36b6d207ed4012c4c6e86b7a9173bc

Observation 2e698ea0-b7b7-4b00-8086-b84d28f5ab9f · inbound

Foundation Models for Astrophysics cites this paper.

Foundation Models for Astrophysics Understanding Emergent Abilities of Language Models from the Loss Perspective

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-04T04:31:48.703352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T04:31:48.703352Z digest=sha256:cd5bac9a18c8159442008fe75c50d8f8d3959f1452382a9845e7fcc1a062a467