Pith. sign in

Paper Citation Record · LEDGER

Why Do More Experts Fail? A Theoretical Analysis of Model Merging

As of 9 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 5 inbound Pith citation observations for arXiv:2505.21226.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.21226 v2

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:44:36.927598Z

measured 43 of 43 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T02:42:45.039852Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T11:24:38.144959Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact1
  • verified fuzzy7
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e29d4997-1560-4b85-9b1a-a27a5b968a5e · outbound

This paper cites Evolutionary optimization of model merging recipes.Nature Machine Intelligence, pages 1–10, 2025.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Evolutionary optimization of model merging recipes.Nature Machine Intelligence, pages 1–10, 2025

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:32.608750Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:32.608750Z digest=sha256:18387bcb2e89ee20b9c5913046e1afc6d3cfeee72bd96bfef208143a247e5a1c

Observation 9a59c3fa-eed0-4c63-975b-6ee3caa393bc · outbound

This paper cites Living on the edge: Phase transitions in convex programs with random data.Information and Inference: A Journal of the IMA, 3(3):224–294, 2014.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Living on the edge: Phase transitions in convex programs with random data.Information and Inference: A Journal of the IMA, 3(3):224–294, 2014

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:44:38.665865Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:44:32.677745Z digest=sha256:9a6e80ece6f5b3889de99efc2da1a0990cdd1951441b54d2a2649771c8137bd0

Observation 5b7ccb7a-ad41-4bfe-bc71-3eea40db3453 · outbound

This paper cites Program Synthesis with Large Language Models.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Program Synthesis with Large Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:32.790393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:32.790393Z digest=sha256:970f8dce3fc09bd39dd6b36a904e56c24083ddf97fda757196e42cd7ccdb61d4

Observation d265a83f-900c-46d6-ba8d-a30de0a616bc · outbound

This paper cites Revisiting Weight Averaging for Model Merging.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Revisiting Weight Averaging for Model Merging

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:32.908553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:32.908553Z digest=sha256:036b58eef17015924778466d276a70599a67486a85ab8dc0ff80d2306b76093e

Observation 733253d7-b3d9-40ee-b755-673c64c1964c · outbound

This paper cites Model Swarms: Collaborative Search to Adapt LLM Experts via Swarm Intelligence.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Model Swarms: Collaborative Search to Adapt LLM Experts via Swarm Intelligence

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:33.011741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:33.011741Z digest=sha256:dc931571d85498dcfc09b24218c31d4006e88b5175316746e40186b3c04e2a47

Observation ffdab5c8-5420-4a83-8fa6-a02dcf3303ea · outbound

This paper cites Measuring Massive Multitask Language Understanding.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Measuring Massive Multitask Language Understanding

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:33.131628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:33.131628Z digest=sha256:044a80ae1e0091b309105ed3024957fd0da127d77f7673e429294bb66aae47c6

Observation d5213ade-951a-4a9c-aa43-dddaed611e6b · outbound

This paper cites Measuring Mathematical Problem Solving With the MATH Dataset.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Measuring Mathematical Problem Solving With the MATH Dataset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:33.275322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:33.275322Z digest=sha256:a0e3287ba3deb3fd9ff401ad82d68d5becc063488e9a3703be6265f1d2fcdd43

Observation 7de861e7-b0a6-4fa2-b594-db4a575157f0 · outbound

This paper cites Lora: Low-rank adaptation of large language models.ICLR, 1 (2):3, 2022.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Lora: Low-rank adaptation of large language models.ICLR, 1 (2):3, 2022

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:33.398276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:33.398276Z digest=sha256:d126064a2570d7f9c37e10d875d87c65a54b17e530242ee7d44ef18951512f68

Observation f0422f54-d67a-40e6-af3e-8cd41f517f1e · outbound

This paper cites LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging LoraHub: Efficient Cross-Task Generalization via Dynamic LoRA Composition

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:33.469848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:33.469848Z digest=sha256:8772919b2f38a4726c8c2b2fa1a326d6c26fe5d514e0a63cbfe227ba9af2f909

Observation 9ca7eda8-fd42-4db0-af8b-0cf7fd2f754f · outbound

This paper cites Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Camels in a Changing Climate: Enhancing LM Adaptation with Tulu 2

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:33.633023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:33.633023Z digest=sha256:7d1b4575ada2cd2245165c3b4863c794e975ffeb35ac259727d971d0368239fc

Observation 2088241f-d960-4be2-a8f4-5cdd46cdb00d · outbound

This paper cites The singular value decomposition: Its computation and some applications.IEEE Transactions on automatic control, 25(2):164–176, 1980.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging The singular value decomposition: Its computation and some applications.IEEE Transactions on automatic control, 25(2):164–176, 1980

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:33.713425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:33.713425Z digest=sha256:ecb0ef12be9fab4ddb6562f7b7fedfbd603d2efed759cb632445c5681361a01a

Observation 45931957-3294-4814-9ff7-98b9280b7feb · outbound

This paper cites How many degrees of freedom do we need to train deep networks: a loss landscape perspective.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging How many degrees of freedom do we need to train deep networks: a loss landscape perspective

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-08-07T13:44:37.342378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:44:33.847753Z digest=sha256:27468aa76b63dda0d12eda4791b5695603d10172e2cdc6154e51bccb1bce3c62

Observation 44f2716b-30b3-49e4-b820-8e7757dcc472 · outbound

This paper cites The Power of Scale for Parameter-Efficient Prompt Tuning.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging The Power of Scale for Parameter-Efficient Prompt Tuning

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:33.957436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:33.957436Z digest=sha256:473dfe65d598823f00f8af63dce2ab97a30b75a947cbde633c142fc919790567

Observation 7eb2afd5-079c-4563-99bd-828a43b0c9fd · outbound

This paper cites Prefix-Tuning: Optimizing Continuous Prompts for Generation.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Prefix-Tuning: Optimizing Continuous Prompts for Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:34.061461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:34.061461Z digest=sha256:a52d1958ce80eb35e0d676a4396fb7da6254bffb22ceb20bbf04303474ea3a57

Observation 7c15703c-d038-4a43-a856-03f26bf24f44 · outbound

This paper cites Gpt understands, too.AI Open, 5:208–215, 2024.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Gpt understands, too.AI Open, 5:208–215, 2024

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:34.182206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:34.182206Z digest=sha256:3d0f35bc3eb0193972d0665b8c4fdcd907717705aa0420864b5425938008eafb

Observation 1c215081-916c-4983-9e2e-78cf062b1f82 · outbound

This paper cites HiFT: A Hierarchical Full Parameter Fine-Tuning Strategy.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging HiFT: A Hierarchical Full Parameter Fine-Tuning Strategy

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:34.329800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:34.329800Z digest=sha256:73a85cfc369cc5bbfe4c88e00f77d08749712a47bb73c6ce76475d23a7cfaf67

Observation 7a8a99c8-b520-4f7e-bd11-4cfa72266b36 · outbound

This paper cites MoELoRA: Contrastive Learning Guided Mixture of Experts on Parameter-Efficient Fine-Tuning for Large Language Models.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging MoELoRA: Contrastive Learning Guided Mixture of Experts on Parameter-Efficient Fine-Tuning for Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:34.475689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:34.475689Z digest=sha256:ddda64b80ae1bc858fa6423407fc583ae5f5e6cb02b16e0cf3648a46ff09f9f5

Observation 965601e2-00c2-4dbe-b1b5-8a9b2d8b731d · outbound

This paper cites Pack of LLMs: Model Fusion at Test-Time via Perplexity Optimization.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Pack of LLMs: Model Fusion at Test-Time via Perplexity Optimization

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:34.591129Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:34.591129Z digest=sha256:5a67b7d8b2129969bbb0bf57f888c2c1087dd7f2c8654e3037a905395b25ee64

Observation fdbba087-fced-4277-9e75-5854c8769165 · outbound

This paper cites Orthogonal adaptation for modular customization of diffusion models.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Orthogonal adaptation for modular customization of diffusion models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:34.686559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:34.686559Z digest=sha256:a298cd0e97ad44bbbcd171f97f335041bff6e45670153bdcd3ea6ac3d02d76a7

Observation 3ed1964a-6fa9-44ec-8605-31cba3b67616 · outbound

This paper cites Ensemble learning.Ensemble machine learning: Methods and applications, pages 1–34, 2012.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Ensemble learning.Ensemble machine learning: Methods and applications, pages 1–34, 2012

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:44:38.461282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:44:34.813147Z digest=sha256:ca8fd02b03dc8688ad4a5b422a7a243647fd48648f5c04973b8305971ea1541c

Observation ae21ea70-f297-4238-ad6c-fc313818264b · outbound

This paper cites Acceleration of stochastic approximation by averaging.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Acceleration of stochastic approximation by averaging

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:34.956683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:34.956683Z digest=sha256:9b56745139b3848bc79eca25a5f4912c6c184ff0b521680c28970ca180c66a84

Observation 28bdaed8-72f2-4824-90a6-750e342590e2 · outbound

This paper cites LoRA Soups: Merging LoRAs for Practical Skill Composition Tasks.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging LoRA Soups: Merging LoRAs for Practical Skill Composition Tasks

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:35.077153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:35.077153Z digest=sha256:1b8720c036c1a71f4d0baee811702ebfec3c1427d2a4702c4b12a57cd26f4a96

Observation 0130fcf2-28b3-47cb-a2ed-04fdd78c52a3 · outbound

This paper cites Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards.Advances in Neural Information Processing Systems, 36:71095–71134, 2023.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Rewarded soups: towards pareto-optimal alignment by interpolating weights fine-tuned on diverse rewards.Advances in Neural Information Processing Systems, 36:71095–71134, 2023

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:35.176600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:35.176600Z digest=sha256:645165ae6d1fe7674b28dc88157b59c1181c1f9c2c73111f5f39f6b80f928dca

Observation 54698c76-c14f-44bb-9fa6-b99da878fa47 · outbound

This paper cites Language Models are Multilingual Chain-of-Thought Reasoners.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Language Models are Multilingual Chain-of-Thought Reasoners

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:35.245526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:35.245526Z digest=sha256:ca46f40cdd1053602a197b46220a0343b5fd0e00930fb0bd18a07fec8859dfe3

Observation b91491c9-ebee-4694-bcdc-98a6d2ef8eea · outbound

This paper cites CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging CommonsenseQA: A Question Answering Challenge Targeting Commonsense Knowledge

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:35.393197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:35.393197Z digest=sha256:81291b8278972eae29ec9388c26f2fd057dd27f75e1df864a0f8baf828a1695d

Observation edba27e7-7af2-41c2-8b25-c472e8ae6012 · outbound

This paper cites Unlocking the potential of model merging for low-resource languages.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Unlocking the potential of model merging for low-resource languages

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:44:38.282317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:44:35.501545Z digest=sha256:4e78e9e35e8ae06c92e6b5abc30c0ddbb9906c5a17b4a8852f8cf565826f2322

Observation 2ee1fe58-fe4d-47ae-994f-d2886fcc51d7 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Gemma 2: Improving Open Language Models at a Practical Size

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:35.591163Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:35.591163Z digest=sha256:7d7e098d434d4b28247bc13b90882ff4add89714e06334685703e25617252a14

Observation a78f64a5-2263-436b-97be-29d03aaa3b5a · outbound

This paper cites Estimation in high dimensions: a geometric perspective.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Estimation in high dimensions: a geometric perspective

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:44:38.115696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:44:35.688556Z digest=sha256:86689f601270890a19fc2c1e4be638c7e34d64653054deadd67a0d54b73e1f19

Observation 6038e34c-c053-442d-8403-119bcdca394e · outbound

This paper cites Principal component analysis.Chemometrics and intelligent laboratory systems, 2(1-3):37–52, 1987.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Principal component analysis.Chemometrics and intelligent laboratory systems, 2(1-3):37–52, 1987

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:44:37.954937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:44:35.866143Z digest=sha256:b17eb0e78248dcaa9313939d6b55ac3f29f5c71c202f5950c261ca113279916d

Observation 713dba1e-8030-4202-a7d1-c82ead5ede4c · outbound

This paper cites Ties-merging: Resolving interference when merging models.Advances in Neural Information Processing Systems, 36:7093–7115, 2023.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Ties-merging: Resolving interference when merging models.Advances in Neural Information Processing Systems, 36:7093–7115, 2023

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:35.949478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:35.949478Z digest=sha256:e2ca0b0c772e902bfe3f71c63d1808aae264bdce31a570a35489dd93818e097f

Observation d75f9e00-ade1-4bfa-87a9-8e15c2c7303c · outbound

This paper cites What Matters for Model Merging at Scale?.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging What Matters for Model Merging at Scale?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:36.048280Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:36.048280Z digest=sha256:b3c87a7a6be159012fff5385ee4ce8c97dc6dedcd496d8f3833e98d68e9a1269

Observation b7b36f27-715f-4f4d-8a43-5462faa7c56e · outbound

This paper cites AdaMerging: Adaptive Model Merging for Multi-Task Learning.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging AdaMerging: Adaptive Model Merging for Multi-Task Learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:36.149510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:36.149510Z digest=sha256:3bcb48507713a30b71ec96e426c7a3fe61f1e2166b73fce3e1039b51612199fa

Observation e76b9d63-4e76-4c68-8ac7-2615fb08f8dd · outbound

This paper cites Language models are super mario: Absorbing abilities from homologous models as a free lunch.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Language models are super mario: Absorbing abilities from homologous models as a free lunch

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:36.280683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:36.280683Z digest=sha256:750fe72cde6f0bfdd2ae3e80d2bf4b5a8b81735927f8914e6c3b11389eab942e

Observation 40992453-269c-43b9-915f-861568f81fa5 · outbound

This paper cites Emotion detection on tv show transcripts with sequence- based convolutional neural networks.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Emotion detection on tv show transcripts with sequence- based convolutional neural networks

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:44:37.766416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:44:36.447501Z digest=sha256:15216175887bc274e21f37179939a2a908841b36ac023e028373953d1051a98f

Observation 3ecc682e-1cf1-4fd6-bd56-cee4ab56e020 · outbound

This paper cites Composing parameter-efficient modules with arithmetic operation.Advances in Neural Information Processing Systems, 36:12589–12610, 2023.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Composing parameter-efficient modules with arithmetic operation.Advances in Neural Information Processing Systems, 36:12589–12610, 2023

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:36.543531Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:36.543531Z digest=sha256:e13f84107425d9471bb6e23388c87f827df019b476981e841ba46fb085b9287d

Observation e151f6e4-f5a9-48ce-8208-125b4414a467 · outbound

This paper cites Nature-Inspired Population-Based Evolution of Large Language Models.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging Nature-Inspired Population-Based Evolution of Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:36.640223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:36.640223Z digest=sha256:43c4def2819b8567fe5b6e1bc2e94fe96ce11fc106e77222469f08a5c2fb125d

Observation ae8bc4e2-e34d-47ee-aacc-744cc08cdf91 · outbound

This paper cites LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging LlamaFactory: Unified Efficient Fine-Tuning of 100+ Language Models

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T13:44:36.776858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:44:36.776858Z digest=sha256:07ec3ef3cc348db4db4b68a5e2c4ca2ef41d4d09ecd1cf29861cae49e900b52f

Observation b5c8c531-4c32-486b-8dce-fc8bdfb86d45 · outbound

This paper cites particle.

Why Do More Experts Fail? A Theoretical Analysis of Model Merging particle

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T13:44:37.554033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T13:44:36.927598Z digest=sha256:315541170835d751bee951fd698bc8ba44e568a836d49afaee0059b36063873f

Pith citing papers

Observation 6d716525-a76d-431b-822c-a9ccb32d56a8 · inbound

Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities cites this paper.

Model Merging in LLMs, MLLMs, and Beyond: Methods, Theories, Applications and Opportunities Why Do More Experts Fail? A Theoretical Analysis of Model Merging

Reference 245

Resolution
verified exact
arxiv_id, observed 2026-05-17T22:16:04.730252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T22:16:04.386706Z digest=sha256:7f62dbdc182661c20ada7b972a5daa21180142ee18cfb8d0685b240a8be636e5

Observation 26ade7fe-5909-4180-85f5-c4b39f0e4243 · inbound

DiM\textsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging cites this paper.

DiM\textsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging Why Do More Experts Fail? A Theoretical Analysis of Model Merging

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:42:58.884086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-14T20:24:30.679219Z digest=sha256:b77f37e08cf68d9d89a640de5906e312e4bf5e66a63d1ae702933528baf13ec0

Observation ad5d758d-41b7-48d4-a036-7e2ace2d781f · inbound

DiM\textsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging cites this paper.

DiM\textsuperscript{3}: Bridging Multilingual and Multimodal Models via Direction- and Magnitude-Aware Merging Why Do More Experts Fail? A Theoretical Analysis of Model Merging

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-21T09:14:05.733226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T09:12:22.712240Z digest=sha256:d1d63e4f34d3c4053596b7153dc719b28b46ef47b5cfb32034ecdf0af159f85c

Observation 2d3685b4-69b6-40b0-bf5d-a0e5f3951412 · inbound

Model Merging to Evolution: Parameter Space Exploration for Expert Models cites this paper.

Model Merging to Evolution: Parameter Space Exploration for Expert Models Why Do More Experts Fail? A Theoretical Analysis of Model Merging

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T11:24:38.146287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T11:20:04.579617Z digest=sha256:bec9592150e9ee94c7e43978af376394bb09aec23334946f59f04fd371ae94be

Observation dc62f227-3112-43cc-a5f8-3d2b3c7ff4d5 · inbound

Sharpness-aware Model Merging with Salience Recovery for LLM-based Cross-Domain Sequential Recommendation cites this paper.

Sharpness-aware Model Merging with Salience Recovery for LLM-based Cross-Domain Sequential Recommendation Why Do More Experts Fail? A Theoretical Analysis of Model Merging

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T02:42:45.039852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T02:42:45.039852Z digest=sha256:354f5e33052d6d95b935e7b2398dee79548785ae3dd14912a1e4262394e2a811