Pith. sign in

Paper Citation Record · LEDGER

Model Merging in Pre-training of Large Language Models

As of 19 August 2026, this Paper Citation Record lists 57 of 57 outbound references and 11 inbound Pith citation observations for arXiv:2505.12082.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.12082 v3

Coverage vector

measured 57 of 57 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T20:46:28.238597Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 11 of 11 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:30:10.163726Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:49:46.245243Z

Reference resolution

57 of 57 outbound references displayed

  • verified exact0
  • verified fuzzy24
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dc5cf18c-451a-4f70-9c22-712ccb546981 · outbound

This paper cites GPT-4 Technical Report.

Model Merging in Pre-training of Large Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:27.985936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:27.985936Z digest=sha256:434eeee59107a23316621d64de97b9a0cf511b894e3f712fff692bb0cca10adb

Observation 113a2c22-078a-44da-9bb5-3ac6372e80b4 · outbound

This paper cites Evolutionary optimization of model merging recipes.

Model Merging in Pre-training of Large Language Models Evolutionary optimization of model merging recipes

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:27.991379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:27.991379Z digest=sha256:e45d4749e60df96a1b6c0c88c1b4dbd596376a79e27cf114465631b51cbbf293

Observation a06fe965-8853-4ab4-9c45-54f33b7cc8c9 · outbound

This paper cites Program Synthesis with Large Language Models.

Model Merging in Pre-training of Large Language Models Program Synthesis with Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:27.996068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:27.996068Z digest=sha256:5cb145773a6c9fc2bf3f56beb368e0355abc37ba008e83ed83c786836c7819f8

Observation 5bcbde28-ffdc-4097-9834-52c36772cf39 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

Model Merging in Pre-training of Large Language Models DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:28.001379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:28.001379Z digest=sha256:117144479cac4bf947dba085360eb291bdf2b06747871b8a0a29f3ebd3058b39

Observation abf61e49-c49b-4a58-970b-e6ea3f5c7287 · outbound

This paper cites Evaluating Large Language Models Trained on Code.

Model Merging in Pre-training of Large Language Models Evaluating Large Language Models Trained on Code

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:28.006262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:28.006262Z digest=sha256:cf04f7917c4c5f3b2d76aab23cabf9ac2ef1c329298f834fadd243f967ffa317

Observation c91f5bf3-8d9e-48fc-94aa-93bffa4db82e · outbound

This paper cites Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge.

Model Merging in Pre-training of Large Language Models Think you have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:28.011465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:28.011465Z digest=sha256:ee0dbc6e99827bb1f20140d0ac5793484c6922cb645e6bdc3e0b3c8aea94fe38

Observation e1657b69-436f-4b9a-b7ab-d23823604a78 · outbound

This paper cites Training Verifiers to Solve Math Word Problems.

Model Merging in Pre-training of Large Language Models Training Verifiers to Solve Math Word Problems

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:28.016606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:28.016606Z digest=sha256:386c9eef68c0216f6d2ce2d1f9238a183dbc804354be760c59c99a02aa7217be

Observation b8d5fa3a-2383-49ef-ae99-9e16c020f42b · outbound

This paper cites Gradient descent on neural networks typically occurs at the edge of stability.

Model Merging in Pre-training of Large Language Models Gradient descent on neural networks typically occurs at the edge of stability

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:28.021188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:28.021188Z digest=sha256:75b3251ae4240cdc93059b230bf5902079085cd53316ea68cf2a941b149b1851

Observation 7f6af6af-c17b-452e-9cb9-7cdb04977025 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Model Merging in Pre-training of Large Language Models An image is worth 16x16 words: Transformers for image recognition at scale

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:28.025569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:28.025569Z digest=sha256:18b70adea9b2f0b9a80137074f0bbbdc890031c26c33ec9bf32faae77726c298

Observation 694e03df-c1b9-420e-9715-8601d83ca4e1 · outbound

This paper cites DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs.

Model Merging in Pre-training of Large Language Models DROP: A reading comprehension benchmark requiring discrete reasoning over paragraphs

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:29.088974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:46:28.029866Z digest=sha256:c5c58087774e6f936a68c3e5889374f34dd2bfdaa02047bd7777f350f19c6e83

Observation afdfec0b-13eb-410b-90a2-86a92e6d70e3 · outbound

This paper cites The llama 3 herd of models.arXiv e-prints, pages arXiv–2407, 2024.

Model Merging in Pre-training of Large Language Models The llama 3 herd of models.arXiv e-prints, pages arXiv–2407, 2024

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:28.034400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:28.034400Z digest=sha256:9d7aa85dbe04031df797b751251cd73bcd4537447ec27704820040b5a7e3fc7f

Observation d02f92f1-ae8a-4335-b87b-c16443f94f05 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

Model Merging in Pre-training of Large Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:28.038785Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:28.038785Z digest=sha256:711d9e5f3942d2bf61ca5dce1d265065a17c822e2c4ebdacebe2d3cf0fefe16a

Observation 0fdbbd56-4a27-4c7a-a7f2-d1535cbc2ff4 · outbound

This paper cites Scaling laws and compute-optimal training beyond fixed training durations.Advances in Neural Information Processing Systems, 37:76232–76264, 2024.

Model Merging in Pre-training of Large Language Models Scaling laws and compute-optimal training beyond fixed training durations.Advances in Neural Information Processing Systems, 37:76232–76264, 2024

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:29.065656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:46:28.043088Z digest=sha256:f2c720ff871373d9a7511552377c374225ad4580a32d24338e4a3f806dbc67a0

Observation 161b767d-48e6-4e61-939a-a37620cbbca0 · outbound

This paper cites Measuring massive multitask language understanding.

Model Merging in Pre-training of Large Language Models Measuring massive multitask language understanding

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:28.047542Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:28.047542Z digest=sha256:da996b6808c8487637bbfa4693c18ab78ff8bd80beea0920131febf78479aba2

Observation a7168105-42fe-4f7b-be6e-5506274842cf · outbound

This paper cites MiniCPM: Unveiling the potential of small language models with scalable training strategies.

Model Merging in Pre-training of Large Language Models MiniCPM: Unveiling the potential of small language models with scalable training strategies

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:29.042041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:46:28.051774Z digest=sha256:5f1ecdd8cdbe016a3b0b75c7f73172b6cb64ec1c372bc9667b51461d56fa3067

Observation 4370adf0-5a6c-4cb9-937f-865e0524054d · outbound

This paper cites C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models.

Model Merging in Pre-training of Large Language Models C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:29.027766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:46:28.056087Z digest=sha256:06edb0d7f62a82ec70a45ee5cae35a5e23d40008b5e1b360a664fbebcb5958c2

Observation d96def94-1097-4ab3-90da-66a2037988e4 · outbound

This paper cites The exponentially weighted moving average.Journal of quality technology, 18(4):203–210, 1986.

Model Merging in Pre-training of Large Language Models The exponentially weighted moving average.Journal of quality technology, 18(4):203–210, 1986

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:28.060418Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:28.060418Z digest=sha256:737d1484141681051e27ba4b9e6f6849e8b885d7242d256bd62e4b8263fbe649

Observation caade15a-3dbd-4e92-934c-7eeadf223440 · outbound

This paper cites Editing models with task arithmetic.

Model Merging in Pre-training of Large Language Models Editing models with task arithmetic

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:29.003251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:46:28.064763Z digest=sha256:908e6e4415a005b174b81c3fdbcdd63f971061b8a42434cb115bc00c20c19178

Observation 15e89984-c6ad-4484-935f-64099d86899e · outbound

This paper cites Dataless knowledge fusion by merging weights of language models.

Model Merging in Pre-training of Large Language Models Dataless knowledge fusion by merging weights of language models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:28.987743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:46:28.069243Z digest=sha256:7795f53e1d0b8f79f3c629ea0d6f6025fc8eea8db7d45aa625ffc308cd2032e3

Observation 2429eb69-8451-44db-801b-b86f1afe2439 · outbound

This paper cites Some properties of a simple moving average when applied to forecasting a time series.Journal of the Operational Research Society, 50(12):1267–1271, 1999.

Model Merging in Pre-training of Large Language Models Some properties of a simple moving average when applied to forecasting a time series.Journal of the Operational Research Society, 50(12):1267–1271, 1999

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:28.973248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:46:28.073511Z digest=sha256:dfe76332c5fc83352c9e6196e9bbfd5831d7a791525eda12d55fecb3604021e4

Observation a7fe01eb-919e-48e0-a51d-2a8c413791fa · outbound

This paper cites PopulAtion Parameter Averaging (PAPA).

Model Merging in Pre-training of Large Language Models PopulAtion Parameter Averaging (PAPA)

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:28.077730Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:28.077730Z digest=sha256:2f9a82b92f6814d21d1728c47745b7b4fefdffef8803d86b1a09737ed8e8e318

Observation ac51e743-4037-4e1c-b1bf-25e121691780 · outbound

This paper cites TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension.

Model Merging in Pre-training of Large Language Models TriviaQA: A large scale distantly supervised challenge dataset for reading comprehension

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:28.959066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:46:28.082242Z digest=sha256:3320d0d53990d1ce3415d5bf12513f963ecd28db01d95a146fa2482442d29e08

Observation 2b0de40c-ddfa-430a-8e69-f771b2f930d0 · outbound

This paper cites Stop wasting my time! saving days of imagenet and BERT training with latest weight averaging.

Model Merging in Pre-training of Large Language Models Stop wasting my time! saving days of imagenet and BERT training with latest weight averaging

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:28.944286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:46:28.086596Z digest=sha256:a989c8ff7acabd1678fb1685e266f30da15dc871f68cab9ea9046735bda5ad64

Observation 6e4a93c5-356a-4446-9a3a-9e46f31c46e4 · outbound

This paper cites Scaling Laws for Neural Language Models.

Model Merging in Pre-training of Large Language Models Scaling Laws for Neural Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:28.091049Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:28.091049Z digest=sha256:6be672ff10f1e6ccb2cb0e437be6b2259d39011df625547f03d93b2e2ad9ef0e

Observation e54e9882-bdcf-4eaa-9728-ef474e5c1122 · outbound

This paper cites Trainable weight averaging: Efficient training by optimizing historical solutions.

Model Merging in Pre-training of Large Language Models Trainable weight averaging: Efficient training by optimizing historical solutions

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:28.930075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:46:28.095470Z digest=sha256:be9081d17b84e40511e7d07397bbe3de7c58c4063dc515c55a50fd2000f11b70

Observation 08235f68-b1ff-48d8-961b-a118c409527e · outbound

This paper cites DeepSeek-V3 Technical Report.

Model Merging in Pre-training of Large Language Models DeepSeek-V3 Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:28.099762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:28.099762Z digest=sha256:534184c6cb1010668b24b5af8f464813b5f95bc4d9b1d3bf97f84e6ead1604f8

Observation ceb07be9-756a-4a49-a66e-d7565e759fcd · outbound

This paper cites Checkpoint Merging via Bayesian Optimization in LLM Pretraining.

Model Merging in Pre-training of Large Language Models Checkpoint Merging via Bayesian Optimization in LLM Pretraining

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:28.104948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:28.104948Z digest=sha256:b02b1738120973da77d1796f92abfcbefe8a762e8c067fc44568a2ba2aa258f0

Observation 3309126d-ca29-4e54-a4e6-7c4fee9cfeb0 · outbound

This paper cites SGDR: Stochastic gradient descent with warm restarts.

Model Merging in Pre-training of Large Language Models SGDR: Stochastic gradient descent with warm restarts

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:28.109336Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:28.109336Z digest=sha256:05934ecd74e0b8b9f0d414c51b35c100beaa4560522d5faaf6dccdb00f5007bc

Observation fcf1b735-1106-4bb8-a3d5-e7a6358126f7 · outbound

This paper cites Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct.

Model Merging in Pre-training of Large Language Models Wizardmath: Empowering mathematical reasoning for large language models via reinforced evol-instruct

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:28.906903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:46:28.113454Z digest=sha256:b877e4c6056342dd0be5d253b1d5c9db63f275af79a9583da266e123edcf13d9

Observation 13cb1706-2de0-43c7-a4af-c32ebba4822d · outbound

This paper cites Merging models with fisher-weighted averaging.Advances in Neural Information Processing Systems, 35:17703–17716, 2022.

Model Merging in Pre-training of Large Language Models Merging models with fisher-weighted averaging.Advances in Neural Information Processing Systems, 35:17703–17716, 2022

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:28.117878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:28.117878Z digest=sha256:ef164e213898c01fab2136dfeb802fe3a740241f2dab8a329d9bd79772fc5de3

Observation fcb3d693-7a59-41d3-937b-9a222a560c6c · outbound

This paper cites An Empirical Model of Large-Batch Training.

Model Merging in Pre-training of Large Language Models An Empirical Model of Large-Batch Training

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:28.122787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:28.122787Z digest=sha256:f6e4d0a8be6ab4ba5933510c10936b271d18475825f5485b01e7ad8e0c6bd2ce

Observation e0cfc1d1-ea07-4974-a2ab-daedd6058278 · outbound

This paper cites The weighted moving average technique.Wiley Encyclopedia of Operations Research and Management Science, 2010.

Model Merging in Pre-training of Large Language Models The weighted moving average technique.Wiley Encyclopedia of Operations Research and Management Science, 2010

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:28.883090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:46:28.127526Z digest=sha256:a4f999a86a47f5f68ff0f7a4ee83f749f4f1d287714a6db059227778f476c118

Observation 15e7104d-207f-4cdc-af77-ce5ef72d3731 · outbound

This paper cites Gpqa: A graduate-level google-proof q&a benchmark.

Model Merging in Pre-training of Large Language Models Gpqa: A graduate-level google-proof q&a benchmark

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:28.131557Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:28.131557Z digest=sha256:5cfcb08a609bcca2f2a4424fd9d137d93f3fe027e9b30b1c2246caf2f15bf85d

Observation 6017fe27-3ec2-46c4-acce-6ac0f84d3535 · outbound

This paper cites Winogrande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106, 2021.

Model Merging in Pre-training of Large Language Models Winogrande: An adversarial winograd schema challenge at scale.Communications of the ACM, 64(9):99–106, 2021

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:28.136077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:28.136077Z digest=sha256:4fd0d49fd9450847502e6ef08ff7b4f202633cff3060e85d4e9c739e742bc5c5

Observation f733f241-f864-4aa3-b47b-e00bc7be65b6 · outbound

This paper cites Early weight averaging meets high learning rates for LLM pre-training.

Model Merging in Pre-training of Large Language Models Early weight averaging meets high learning rates for LLM pre-training

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:28.850780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:46:28.140589Z digest=sha256:d463f27ff0a3543ac297436088386839e2824af15204ef71d1801c8ad4063f0f

Observation ab4f75e6-a192-43b5-982f-1f1bf5f31dd0 · outbound

This paper cites Seed-thinking-v1.

Model Merging in Pre-training of Large Language Models Seed-thinking-v1

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:28.145028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:28.145028Z digest=sha256:46b2efe61d7b522c8d638bde87642f66c1eb465b8eb1ee089cf7b32965b71188

Observation 3e384067-98dc-4227-a9bd-81e473be7b7e · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

Model Merging in Pre-training of Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:28.149485Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:28.149485Z digest=sha256:3d619c34f509f2d64bf9813f447dd3d101b6eb927c2afb9a307766ded3fb2499

Observation 7a370094-ffcc-4ee8-ae56-bd33208f8ad3 · outbound

This paper cites Outrageously large neural networks: The sparsely-gated mixture-of-experts layer.

Model Merging in Pre-training of Large Language Models Outrageously large neural networks: The sparsely-gated mixture-of-experts layer

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:28.154477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:28.154477Z digest=sha256:dfc125552a78274f1780b59b4cc23d2f14dc6da3e7dbc53094eaad42e31b7d3c

Observation 1205ee6d-ee6c-4766-bad1-f09d51b5f3ee · outbound

This paper cites Challenging BIG-bench tasks and whether chain-of- thought can solve them.

Model Merging in Pre-training of Large Language Models Challenging BIG-bench tasks and whether chain-of- thought can solve them

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:28.828274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:46:28.159389Z digest=sha256:b4cf2a05a1d06e94d429b033627946c250078f98cea411cced4d87d1e4871a36

Observation 73f8b1d7-a9a4-479b-a23e-544e4c0951be · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Model Merging in Pre-training of Large Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:28.163645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:28.163645Z digest=sha256:f4f9677abb3353d23fc2eacae346c3a2a44afac5c7e5a419b87fa06a467f06b6

Observation 6c6fd893-5ccb-4866-8714-d9865e9d7e1a · outbound

This paper cites Mmlu-pro: A more robust and challenging multi-task language understanding benchmark.

Model Merging in Pre-training of Large Language Models Mmlu-pro: A more robust and challenging multi-task language understanding benchmark

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:28.814108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:46:28.168009Z digest=sha256:47a9836c182734d284ba7cd52490b5ed65d4377e47536cd562601f5284908763

Observation 3f52f625-1ed8-4e1d-b7c6-b32076afabdb · outbound

This paper cites Dai, and Quoc V Le.

Model Merging in Pre-training of Large Language Models Dai, and Quoc V Le

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:28.800404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:46:28.172072Z digest=sha256:a84333ab99edbe5d6c9378a555a6f64dd3c98cdc81046bd697495e472a844cef

Observation 8b17a5cd-4a8d-4f45-bece-510fea2ae4d2 · outbound

This paper cites Livebench: A challenging, contamination-limited LLM benchmark.

Model Merging in Pre-training of Large Language Models Livebench: A challenging, contamination-limited LLM benchmark

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:28.786598Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:46:28.176555Z digest=sha256:b61939a89fa9a40c6c27b4b56f85b7f3e4fc6118bd3958dfde2902948d21c9c0

Observation 5ed6c65a-7216-4061-91ea-c214d88deb82 · outbound

This paper cites Small-scale proxies for large-scale transformer training instabilities.

Model Merging in Pre-training of Large Language Models Small-scale proxies for large-scale transformer training instabilities

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:28.771728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:46:28.181037Z digest=sha256:40c2096e3d5284315ce2d3fbc569ce83f317414d273571e0db5720f23360e59c

Observation 07f1acda-f168-41bb-8595-4c7eb99b71b7 · outbound

This paper cites WizardLM: Empowering large pre-trained language models to follow complex instructions.

Model Merging in Pre-training of Large Language Models WizardLM: Empowering large pre-trained language models to follow complex instructions

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:28.756060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:46:28.185394Z digest=sha256:20a5ea768ed23597d31d51d186b1d35668f92e788ab62694799ebfa5e6622790

Observation 3f9a73ed-d049-4584-b97b-64c716ee8d7e · outbound

This paper cites TIES-merging: Resolving interference when merging models.

Model Merging in Pre-training of Large Language Models TIES-merging: Resolving interference when merging models

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:28.739730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:46:28.189876Z digest=sha256:302792cb32ea430ea204d819141955051e6c1d35e0da8809742cbb0c88eb59b7

Observation 975fafc3-1f35-4e9e-acd6-997afcb40bc5 · outbound

This paper cites Baichuan 2: Open Large-scale Language Models.

Model Merging in Pre-training of Large Language Models Baichuan 2: Open Large-scale Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:28.194200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:28.194200Z digest=sha256:14206d8e3a8939d5637418e2caa0a0926f409aa997cc3df09063b31744a73fe3

Observation 1419f775-e006-4a89-8060-f18f0bf6a6ae · outbound

This paper cites Qwen2.5 Technical Report.

Model Merging in Pre-training of Large Language Models Qwen2.5 Technical Report

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:28.198793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:28.198793Z digest=sha256:76ce5b0893356b983c68157e86d0bdac2f79d5ccfa9f0d3dce3dc688a69e0f10

Observation 20d6aab7-72a0-45d2-a645-5d0e0a4bd855 · outbound

This paper cites Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities.

Model Merging in Pre-training of Large Language Models Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:28.203139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:28.203139Z digest=sha256:4521f64b6eb675f86f4788e50a1a2574b4ed82408be5f20e9d29515a134b19dc

Observation aa8c7e5b-22e5-4888-959e-d68ffeec81a2 · outbound

This paper cites Adamerging: Adaptive model merging for multi-task learning.

Model Merging in Pre-training of Large Language Models Adamerging: Adaptive model merging for multi-task learning

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:28.725421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:46:28.207520Z digest=sha256:f2858c01bca96f167377437ce3ff05084600bc0732c4f52086e03eb0ee444ed4

Observation bc2a334e-9ee0-4af7-82c9-269437af24c5 · outbound

This paper cites Language models are super mario: Absorbing abilities from homologous models as a free lunch.

Model Merging in Pre-training of Large Language Models Language models are super mario: Absorbing abilities from homologous models as a free lunch

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:28.711011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:46:28.211878Z digest=sha256:175b20dfa43a3cfc7bfc2507d6773ee94047687f93b49adf7958e8bed05b964b

Observation 3e4cb3b1-6dcc-4ef7-9ac8-a149d50e25fe · outbound

This paper cites DAPO: An Open-Source LLM Reinforcement Learning System at Scale.

Model Merging in Pre-training of Large Language Models DAPO: An Open-Source LLM Reinforcement Learning System at Scale

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:28.216068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:28.216068Z digest=sha256:f9d9fad4d3ba60ddc9e5ea1bdbbcf0c4958de3f6185b172d8a725e97551413ca

Observation fa64e4dd-c4bf-4bef-8fa8-94024c6853c2 · outbound

This paper cites VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks.

Model Merging in Pre-training of Large Language Models VAPO: Efficient and Reliable Reinforcement Learning for Advanced Reasoning Tasks

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:28.220562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:28.220562Z digest=sha256:b3dc4455a55eea4ff601fdd5f0705725abb0c9135b99b4393938a2c426aa7759

Observation 4d7d8925-ac6d-4eca-9cb6-31086e773e10 · outbound

This paper cites HellaSwag: Can a Machine Really Finish Your Sentence?.

Model Merging in Pre-training of Large Language Models HellaSwag: Can a Machine Really Finish Your Sentence?

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:28.225152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:28.225152Z digest=sha256:4aae133f78c866ce285d48e4bcaaaba074ac8b7572fd7eaed04f02287f531778

Observation 9e386bee-32b1-4cea-863e-91acaf16e72a · outbound

This paper cites Ape210K: A Large-Scale and Template-Rich Dataset of Math Word Problems.

Model Merging in Pre-training of Large Language Models Ape210K: A Large-Scale and Template-Rich Dataset of Math Word Problems

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T20:46:28.229714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T20:46:28.229714Z digest=sha256:5b735f055873d13112fa8d37841ba060750c2e7de0b6d8212b4bc1b25db5232f

Observation 6694c25b-3f3b-41a3-bd03-50540bfc15eb · outbound

This paper cites AGIEval: A human-centric benchmark for evaluating foundation models.

Model Merging in Pre-training of Large Language Models AGIEval: A human-centric benchmark for evaluating foundation models

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:28.695843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:46:28.234367Z digest=sha256:c4bbc3e2a183596938c78840a61e1605b09e23a9eb574ee980c6d925aeef6ab9

Observation bf9d46d5-0779-438b-aab2-fd7d94bc92b3 · outbound

This paper cites MetaGPT: Merging large language models using model exclusive task arithmetic.

Model Merging in Pre-training of Large Language Models MetaGPT: Merging large language models using model exclusive task arithmetic

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T20:46:28.681134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T20:46:28.238597Z digest=sha256:497d574b0d7f0aa74ac0b5bfffcb5298f5dd28997380d08fba7843e3fcd41c2a

Pith citing papers

Observation 11f46116-417b-46e2-9a8f-da9c9ba887cd · inbound

Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities cites this paper.

Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities Model Merging in Pre-training of Large Language Models

Reference 129

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:16:04.524049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-17T22:16:04.386706Z digest=sha256:516e53fef0d737d37b66b017983b55fe0b25b9fcf88d9f1bf8dbf21a67f2b06d

Observation 04dde608-7a4e-4db6-9a65-e5d8367a23ee · inbound

Kwai Keye-VL Technical Report cites this paper.

Kwai Keye-VL Technical Report Model Merging in Pre-training of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T20:45:08.181516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:45:08.181516Z digest=sha256:f6177f266303dbc05dcb5ecce1812712a0a3398bfbc3b4ba9d2ebadf1d6c3e96

Observation b01d7e69-effd-4359-869c-aaa6507bee03 · inbound

WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training cites this paper.

WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training Model Merging in Pre-training of Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T14:49:40.270904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:49:40.270904Z digest=sha256:a50cb116dc4e1a92d443c7de4bca9fbf237b6f23fc4976bf4860a028d14e60ad

Observation d532f56f-1d36-4e14-a307-9d752f503eed · inbound

ReaLM: Reflection-Enhanced Autonomous Reasoning with Small Language Models cites this paper.

ReaLM: Reflection-Enhanced Autonomous Reasoning with Small Language Models Model Merging in Pre-training of Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T17:30:10.163726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T17:30:10.163726Z digest=sha256:37883cb45df128738049d6630f07ad1d6dab2a8afa9a44e523638b2b9b5f2bf5

Observation a7a559c1-04ee-4264-b5dd-27d1d9222cd3 · inbound

Waver: Wave Your Way to Lifelike Video Generation cites this paper.

Waver: Wave Your Way to Lifelike Video Generation Model Merging in Pre-training of Large Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-05T17:45:16.870761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T17:45:16.870761Z digest=sha256:dc95b0a978d5b5a3de94a14e116da19ee797c0ddf07c24246d477c2220ecc880

Observation cdfded52-de3c-4810-8fbc-99db82938247 · inbound

ScheduleFree+: Scaling Learning-Rate-Free & Schedule-Free Learning to Large Language Models cites this paper.

ScheduleFree+: Scaling Learning-Rate-Free & Schedule-Free Learning to Large Language Models Model Merging in Pre-training of Large Language Models

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-20T12:23:16.787163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-20T12:22:30.263086Z digest=sha256:31e72c7aa24623069c0eb0b899b135695ba44af3edbdc44f94904528d29f71ff

Observation e7bd2b19-81d5-452c-b607-0ede6828ea68 · inbound

Anytime Training with Schedule-Free Spectral Optimization cites this paper.

Anytime Training with Schedule-Free Spectral Optimization Model Merging in Pre-training of Large Language Models

Reference 27

Resolution
verified exact
arxiv_id, observed 2026-05-25T05:40:24.455637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-05-25T05:38:16.958574Z digest=sha256:80049ca55d1857311ada581cefaac4413d99546ab4107b93f2560102b30198e5

Observation 6cac034e-95b0-4736-a464-d13520d85f90 · inbound

Kwai Keye-VL-2.0 Technical Report cites this paper.

Kwai Keye-VL-2.0 Technical Report Model Merging in Pre-training of Large Language Models

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-07-03T04:27:37.028942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-27T13:53:10.352603Z digest=sha256:b0b5442be4b86fa0579934c9bb7bef97e40c037d5e61e35f0d4439fee5eafbd4

Observation a161417f-12aa-443a-9431-55420b5ce9fc · inbound

Humanoid-OmniOcc: Stereo-Based Full-View Occupancy Dataset for Embodied AI cites this paper.

Humanoid-OmniOcc: Stereo-Based Full-View Occupancy Dataset for Embodied AI Model Merging in Pre-training of Large Language Models

Reference 65

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:49:46.246632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-06-26T08:29:13.389779Z digest=sha256:fef8c4681d9aadcea297d8b4bd9eaf8f6be5399ace182c02bd46be50891ee9e7

Observation 556c49c3-1768-4df0-bda2-1621d33479bb · inbound

LLM Evolution as an Industry-Scale Ecosystem: A Lifecycle Perspective on Continual Learning cites this paper.

LLM Evolution as an Industry-Scale Ecosystem: A Lifecycle Perspective on Continual Learning Model Merging in Pre-training of Large Language Models

Reference 59

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:48:39.896841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-06-27T05:02:18.347642Z digest=sha256:bc41e8c927c070d601e86073f6545305b0a9d6cfe008e73e58de1979194ab3a4

Observation 8a40e590-512f-447e-9ff7-8ec128222178 · inbound

Optimizing Visual Generative Models via Distribution-wise Rewards cites this paper.

Optimizing Visual Generative Models via Distribution-wise Rewards Model Merging in Pre-training of Large Language Models

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-07-03T16:48:39.787654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-07-03T16:39:12.711424Z digest=sha256:bbc4e0e03a1aadfeecfde23380b68a0ee39ed3e582974e63952a8409c04e1774