Pith. sign in

Paper Citation Record · LEDGER

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning

As of 15 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 1 inbound Pith citation observation for arXiv:2412.08147.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.08147 v2

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T18:15:11.185463Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:38:48.719839Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T22:38:49.070684Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact4
  • verified fuzzy41
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 395755ce-a8b7-4cf9-9dfb-08a3d245f1ba · outbound

This paper cites write newline.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:10.730523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:10.730523Z digest=sha256:4b1d0fb381c77c754b1b5b51dead72c569926284c85a9c5abac61418bdf29935

Observation 964b5bda-5aeb-4d6e-8f45-272fa8877e9f · outbound

This paper cites Muppet: Massive multi-task representations with pre-finetuning.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Muppet: Massive multi-task representations with pre-finetuning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.917923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.737695Z digest=sha256:ce21ad597acb8121c0e8d57dee60ebcbf123d91e0d51f89cd32b8dc82fb524ac

Observation 0e954d8c-e2b0-4e09-bfb8-138d5095de3f · outbound

This paper cites Safety-tuned LL a MA s: Lessons from improving the safety of large language models that follow instructions.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Safety-tuned LL a MA s: Lessons from improving the safety of large language models that follow instructions

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.877649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.743431Z digest=sha256:18b9cd11c8c7f0a031710f04f07cb46cfdb18c541efc49613756a221fa9c6fde

Observation 0b35ded3-ba47-40aa-b1b3-a6976ba82b26 · outbound

This paper cites Neural networks for pattern recognition.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Neural networks for pattern recognition

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.839602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.750856Z digest=sha256:eab88a7dfdd1c2626c8bde5837b6d49c04b5e6ae1b676d82f6e2db2cae8b1f2a

Observation de994e55-adbb-4314-b309-acbb17e2bb7b · outbound

This paper cites an unresolved cited work.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:10.758533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:10.758533Z digest=sha256:77bc44edd5bda17379340c86b4d9c0aa879e19cfa5d44d0a6b85f94e126ed384

Observation 3d774df1-e3d0-4e8c-9950-48a0ca900174 · outbound

This paper cites Mode-finding for mixtures of G aussian distributions.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Mode-finding for mixtures of G aussian distributions

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.809183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.765228Z digest=sha256:095fdaa6cb7ca5ef2ed2bd8022035407e11c87c11d5700cdf3054d1c75b90d5b

Observation 18c311c5-d512-4450-b645-bc3a55f4abe4 · outbound

This paper cites Multitask learning.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Multitask learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:10.771952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:10.771952Z digest=sha256:2e1d9f913abc5f54ec0980d99b6476b0875a324829a4afb77821da2d48fd4786

Observation 84cd8a65-06e0-402c-855d-eb11a7205484 · outbound

This paper cites PAC-B ayesian supervised classification: The thermodynamics of statistical learning.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning PAC-B ayesian supervised classification: The thermodynamics of statistical learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.781424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.780121Z digest=sha256:02cad6413c6a130a5620f5bac3a0965dfc5d9ce5f54f3e9265449937c2fd592d

Observation 6fa95463-b68c-4400-805a-98fa617405d3 · outbound

This paper cites Overview of the IWSLT 2017 evaluation campaign.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Overview of the IWSLT 2017 evaluation campaign

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.739210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.787548Z digest=sha256:96760626066c15b22ef5d16e689f5cb71a3a07ac8c70e6e8c054a8e9ba0be40b

Observation 12efde92-4baf-4f4d-9937-f47678ccfd7f · outbound

This paper cites GradNorm : Gradient normalization for adaptive loss balancing in deep multitask networks.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning GradNorm : Gradient normalization for adaptive loss balancing in deep multitask networks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.704671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.795884Z digest=sha256:05ba8a82ffacc266e3df659caccd369b757c8645b5da910dff4e494db9c34eec

Observation 020759e6-28e3-4368-8781-18109fdb74a8 · outbound

This paper cites Remote sensing image scene classification: Benchmark and state of the art.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Remote sensing image scene classification: Benchmark and state of the art

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:10.805911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:10.805911Z digest=sha256:ace136c139324c83c8625f99fa902a13a4bb9af6cd612304a2516ca7b940cff1

Observation b746629e-d543-4e30-a57a-d7ff5ec90582 · outbound

This paper cites Zhao, Yanping Huang, Andrew M.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Zhao, Yanping Huang, Andrew M

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.679730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.812381Z digest=sha256:8e06e0792a4e2aca38c7febfc66c658da3b34fc0ea7d411ddb7cdacc4e97b4b4

Observation 67ca8381-9653-4ef4-a298-33ec930f966a · outbound

This paper cites Model merging by uncertainty-based gradient matching.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Model merging by uncertainty-based gradient matching

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.655628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.820287Z digest=sha256:814ef91b2a635f41c858cef261bbb6cdcb1ebe2acdaacb9fc038214b0e4f112b

Observation 7179bd36-8525-43e5-a84e-3d89372e9c13 · outbound

This paper cites ColD fusion: Collaborative descent for distributed multitask finetuning.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning ColD fusion: Collaborative descent for distributed multitask finetuning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.623970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.829496Z digest=sha256:2c7c32c5ca4ccfcddde0382a6ebec82800947c2ff06d5e68df6d9ac196ffec8a

Observation c2553b3d-71bd-40d3-950a-2e55aaecec7c · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning An image is worth 16x16 words: Transformers for image recognition at scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:10.840168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:10.840168Z digest=sha256:7ef4bcc372708aa6fe001cc7575c5d3492a36aa019a3f95efbc324bd7b4ed038

Observation a46855af-b8df-4a75-b6c4-9bbb0d9283d2 · outbound

This paper cites GLaM: efficient scaling of language models with mixture-of-experts.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning GLaM: efficient scaling of language models with mixture-of-experts

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.576240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.850405Z digest=sha256:db146f87c7e01712f45897308b5c02cfdd298e0584be2d6e4c0fc2b386a6b190

Observation 96aab1cc-4881-41ab-a5b4-8666dc84a3f4 · outbound

This paper cites Durrant-Whyte and Mike Stevens.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Durrant-Whyte and Mike Stevens

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.544903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.858209Z digest=sha256:37aea998a4bd2f69f4101d049104afcabb259ba420489759efe23d4df32077f0

Observation 7db04345-154f-4cc2-9c3b-8e329aeeceaf · outbound

This paper cites Knowledge card: Filling LLM s' knowledge gaps with plug-in specialized language models.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Knowledge card: Filling LLM s' knowledge gaps with plug-in specialized language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.519976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.865222Z digest=sha256:5a484a7be9170aaa9afd93b1acc03cfe5f73de2e20e2e36b163838476be93e2a

Observation 82690596-fd30-4d45-8362-8dae36fbafe3 · outbound

This paper cites Continual pre-training for cross-lingual LLM adaptation: Enhancing japanese language capabilities.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Continual pre-training for cross-lingual LLM adaptation: Enhancing japanese language capabilities

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.488184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.873089Z digest=sha256:4f24e8e672f0ce07b764b6ebeca1cebc7f7df8662f97e665105f95e2fcaeb24f

Observation 557a0e26-fc38-49c2-9823-140db3640a15 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Gemma 2: Improving Open Language Models at a Practical Size

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:10.879522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:10.879522Z digest=sha256:c3a91ceea554b7dec7987fbfe484aa8c04445dd86beaac363bda6dd7de2ec5ef

Observation a6795780-5525-4f77-89a1-277feca12c6a · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Gemma: Open Models Based on Gemini Research and Technology

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:10.886248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:10.886248Z digest=sha256:fad48738101387e5e64186a7ce54d51b9319bba1d8a8fff000053fbb612ab974

Observation 81bc45e0-b04c-4bba-8969-4468630855b0 · outbound

This paper cites Multi-loss weighting with coefficient of variations.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Multi-loss weighting with coefficient of variations

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.461074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.894295Z digest=sha256:8aa5f0400a90671310074ff5e9f9f7c15dd3e481d9f12a6e31c2b150f187b21d

Observation 1183bd3c-63b1-4443-9971-171e25a6ffe4 · outbound

This paper cites Deep residual learning for image recognition.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Deep residual learning for image recognition

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:10.900181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:10.900181Z digest=sha256:b84753efc1bfda2169af1d3099b76c6f5e95727eae0c148bbd09a13a8a54c632

Observation c1a94f2a-443b-4dcd-944e-13ad64eb79ec · outbound

This paper cites Euro SAT : A novel dataset and deep learning benchmark for land use and land cover classification.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Euro SAT : A novel dataset and deep learning benchmark for land use and land cover classification

Reference 24

Resolution
verified exact
raw_fallback, observed 2026-08-11T18:15:12.306623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.907746Z digest=sha256:6d577dfaab5f1a0aaf3cdd43a0b9975abf98c8f56107a2c0d15245a32e64b03b

Observation c0d203fc-989a-4dd8-86ef-e7c0ef3ce1d2 · outbound

This paper cites Detection of traffic signs in real-world images: The G erman T raffic S ign D etection B enchmark.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Detection of traffic signs in real-world images: The G erman T raffic S ign D etection B enchmark

Reference 25

Resolution
verified exact
raw_fallback, observed 2026-08-11T18:15:12.120608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.913625Z digest=sha256:a2e49e29dbc6a3654785d897d00e8eca4da89f442c4ca29226b714e25607fc2c

Observation 572c2ceb-c933-4c12-ab1f-bd6df9ee96c8 · outbound

This paper cites Lo RA : Low-rank adaptation of large language models.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Lo RA : Low-rank adaptation of large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.434148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.919559Z digest=sha256:fad09fb09a5b29db0a80c91fb40344e5d5cd9f717ca16fbbcb9810bb93d7b956

Observation f5eff7f6-483d-418f-afb6-437de8b88dc5 · outbound

This paper cites Editing models with task arithmetic.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Editing models with task arithmetic

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.413070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.926694Z digest=sha256:9ab6f055275b82921de0fbbfb80c38e99d4b53327dd03923ab2845ff5135bdae

Observation 5c23cc78-7eea-49ca-9fe6-e439d8f52599 · outbound

This paper cites Multitask Mayhem: Unveiling and Mitigating Safety Gaps in LLMs Fine-tuning.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Multitask Mayhem: Unveiling and Mitigating Safety Gaps in LLMs Fine-tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:10.933257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:10.933257Z digest=sha256:486df96ca17ea66beb5dad22d3fd10748eba43b714c1f0b5533dca1baf6bad01

Observation 7159d9cb-459c-40fc-9ebd-39ce6e83583b · outbound

This paper cites ForkMerge : Mitigating negative transfer in auxiliary-task learning.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning ForkMerge : Mitigating negative transfer in auxiliary-task learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.380992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.940131Z digest=sha256:8f02bf851e9317f41701580f0eb98142a1198d5e4ab72559922dfd55ad28e27a

Observation 710b5c2c-3aca-4774-b597-a91a5cea8dca · outbound

This paper cites Dataless knowledge fusion by merging weights of language models.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Dataless knowledge fusion by merging weights of language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.342746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.947644Z digest=sha256:71e5bb6b0563d10a271c364acbcb27b2471007011135ac40eadac2bc26fb2894

Observation 52f24e98-7fbd-4fdc-9b91-53599770f4bb · outbound

This paper cites The B ayesian learning rule.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning The B ayesian learning rule

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.301085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.953903Z digest=sha256:0972f8bac9e22e6b501abc77597b67edd80b9437d698f9e693023fbb8bc2b8e3

Observation 8fda7099-3acd-4845-bf22-4f159185ae92 · outbound

This paper cites Fast and scalable bayesian deep learning by weight-perturbation in Adam.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Fast and scalable bayesian deep learning by weight-perturbation in Adam

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.272196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.962011Z digest=sha256:e0e22a709d880cfcac3ceee5a84d96238e9b98648b924f4f231c51e388142e18

Observation 485c367c-903d-41ef-a06c-9ff3e7bb601e · outbound

This paper cites 3D object representations for fine-grained categorization.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning 3D object representations for fine-grained categorization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:10.970832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:10.970832Z digest=sha256:a58f1e8135a1ae0dd400e74351b3718528e7fda2834964ab21e11426e7600116

Observation 9b844140-3071-4f67-aa83-a82d2194c4c7 · outbound

This paper cites Fast and simple natural-gradient variational inference with mixture of exponential-family approximations.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Fast and simple natural-gradient variational inference with mixture of exponential-family approximations

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.242617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.976303Z digest=sha256:d0a84df46bf5f15a68f9e14193eeaf0c20224e6378a5b9dc0f2679e9733d896a

Observation f98c6fc2-88a7-4804-ace3-584a4bf49310 · outbound

This paper cites MFTCoder: Boosting Code LLMs with Multitask Fine-Tuning.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning MFTCoder: Boosting Code LLMs with Multitask Fine-Tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:10.982181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:10.982181Z digest=sha256:aa7278f5ede3539728285a66d3c58e85a7a41f6b8cf4e6052a33349678d2352c

Observation 74a37287-ab42-4f70-94a1-ed5ce9090c0a · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:10.989779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:10.989779Z digest=sha256:41bc795fa282ed3a16af64daedb1b3625711c7b8dfc5b889895f76fca7c21e28

Observation 8c533a58-f88e-41f7-9cf2-fbe53f4d06d8 · outbound

This paper cites Decoupled weight decay regularization.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Decoupled weight decay regularization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:10.996890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:10.996890Z digest=sha256:b21432b8b57f470e2ddea6792066b50d5fecfdde2f329d0e42284a9760fa8cd0

Observation fea2f85d-a94f-4872-ae6f-11c520ffe25c · outbound

This paper cites Maas, Raymond E.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Maas, Raymond E

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.182130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.004139Z digest=sha256:91e3a335e6537795253826f4962e6d13d9621ce25d8da02234ee4e3555710cd8

Observation 0bcd4f09-389d-4ca9-a158-77b305fec2c1 · outbound

This paper cites EuroLLM: Multilingual Language Models for Europe.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning EuroLLM: Multilingual Language Models for Europe

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:11.010961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:11.010961Z digest=sha256:2d830afb25ee3184a33b1c28837bf8f81a60aa2ac60356772584b347e4b6087f

Observation 3fa8acdd-d773-454f-862a-c982f0af8528 · outbound

This paper cites Merging models with F isher-weighted averaging.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Merging models with F isher-weighted averaging

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.145525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.018531Z digest=sha256:7b718eb0a88b9d273b5d6a23dbf286f6adcc4c0f97c78b69360d8a0946e9399b

Observation f98b26df-4d7c-4dab-bb20-a2440de62d48 · outbound

This paper cites an unresolved cited work.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:11.024309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:11.024309Z digest=sha256:9230b5b307eba05b9f96d9a8ae046c44c674997c5120b27c146d6e292b2f2a71

Observation e469d320-642d-4df6-a0ad-027c149e448c · outbound

This paper cites an unresolved cited work.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Unresolved cited work

Reference 42

Resolution
verified exact
raw_fallback, observed 2026-08-11T18:15:11.708518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.030817Z digest=sha256:015a4975b99e1e126459028ff30d6ef438fc58fcdf7cec337acbca19dcbab566

Observation ad718209-7434-42be-a598-4f2e7c83f624 · outbound

This paper cites Reading digits in natural images with unsupervised feature learning.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Reading digits in natural images with unsupervised feature learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.115920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.038823Z digest=sha256:104ba9b19a5e30633c653fb98f7cf172fc5815f729abd2fc18f40db1a94ef9b0

Observation ac059a11-33d3-430f-8586-0df75f5a2bfb · outbound

This paper cites The variational gaussian approximation revisited.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning The variational gaussian approximation revisited

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:11.045381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:11.045381Z digest=sha256:a953d5bbbd0a184c3c43cb6bddf7aac40cb24052c1d5b6f613a38e600c23c5a0

Observation 5647c61b-0e2a-477b-8edc-4ffd8649f13d · outbound

This paper cites Task arithmetic in the tangent space: Improved editing of pre-trained models.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Task arithmetic in the tangent space: Improved editing of pre-trained models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.070218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.050804Z digest=sha256:812ed812146e47d66171dd4b9a478ba80a28660d0cd961d910e0c9ece5c35fed

Observation 9f8a06f5-d6a4-4b8b-9474-c4fde8e006ee · outbound

This paper cites Training language models to follow instructions with human feedback.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Training language models to follow instructions with human feedback

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.042210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.059099Z digest=sha256:30dc61ab7ae97adadd0e7bde420537837e297f83f0ac4b782a17d0e8ae47048d

Observation b8e9ae59-116f-4e3a-95ae-c74eb59673b0 · outbound

This paper cites Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.006404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.065388Z digest=sha256:febed5bba56bbd31c3912aed91bbce931910edf9aa3c0d66406834d452871111

Observation c5436ae4-1b90-417c-ba7d-a69db06a26d5 · outbound

This paper cites Fine-tuning aligned language models compromises safety, even when users do not intend to! In International Conference on Learning Representations (ICLR), 2024.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Fine-tuning aligned language models compromises safety, even when users do not intend to! In International Conference on Learning Representations (ICLR), 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:12.968269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.072473Z digest=sha256:acab17d007f0c146ef4746c423d26a65d5c26acf9ae48016731c27d728bb836e

Observation 6baa1d97-31aa-4c8b-a2f8-46c822a4b443 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Learning transferable visual models from natural language supervision

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:12.943881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.078934Z digest=sha256:b625cee778140c4fbbce7da1078f37b2947099b3cac21b7cdcc157a51048a0ed

Observation 9eec02c0-f56e-4f67-abef-976c2257cc97 · outbound

This paper cites an unresolved cited work.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:15:12.907624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.085942Z digest=sha256:b4f6fd9041f18a28b101d012bcfeea3a90237db68dbe036d1029ca2c66bdd752

Observation b0c52ace-ab0d-4a77-bd1a-98b00360b063 · outbound

This paper cites Learning to reweight examples for robust deep learning.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Learning to reweight examples for robust deep learning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:12.877091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.093757Z digest=sha256:4b19192eb21826d238ebb9a01a5c14442c1041f1ec9cf12ed77255b33086495f

Observation bd4f4ced-f41a-442a-b3a5-228868419e38 · outbound

This paper cites An Overview of Multi-Task Learning in Deep Neural Networks.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning An Overview of Multi-Task Learning in Deep Neural Networks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:11.101238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:11.101238Z digest=sha256:971bef46c576e27c116302bef6a0a8503f1e920b879a7b5b7850af4611dee02d

Observation 19f589c9-c697-428a-a2d7-deafa1fc82c7 · outbound

This paper cites Variational learning is effective for large deep networks.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Variational learning is effective for large deep networks

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:12.841471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.107479Z digest=sha256:35ef7b302d54c36ac825bc72450ad572aa25eab64d3c4dacfbbb00034fedb2b0

Observation 5450d75e-ee04-4d28-bdd3-7c9bff97333b · outbound

This paper cites Manning, Andrew Ng, and Christopher Potts.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Manning, Andrew Ng, and Christopher Potts

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:12.801207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.112906Z digest=sha256:c7e613fd86af17e7aedd85c8fddccf8b216546972df153b04fb3837ff6ecac2d

Observation 8e03159c-696c-4e1c-b959-13f944dabfaf · outbound

This paper cites ZipIt! M erging models from different tasks without training.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning ZipIt! M erging models from different tasks without training

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:12.771096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.119063Z digest=sha256:6af89d8abecd8a32e9b8fb0e293c0913ea43c6d6136cc429151d21bb998688fe

Observation 9a1db191-e886-4f29-84d9-efceb07a0918 · outbound

This paper cites Self-influence guided data reweighting for language model pre-training.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Self-influence guided data reweighting for language model pre-training

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:12.743621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.124939Z digest=sha256:6b772232844e373285649b63edcae1938a621632962f503a0703af4ab3fab1ac

Observation b02077f1-2c83-4e16-9927-6af9930972b8 · outbound

This paper cites A B ayesian committee machine.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning A B ayesian committee machine

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:12.706744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.131511Z digest=sha256:16a9a0f3a799baff8e230c16d1d9a1c1483c702ffc4cf58751a69d75c95cb588

Observation c2478687-5d2f-4d51-b170-201bdb0ea5cf · outbound

This paper cites Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:12.674500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.139994Z digest=sha256:fc26717e7ad757345126ef0eb311e8aa4ce882d2d4fd49d7794708a5b791369d

Observation cfd7f9dc-8af2-40c3-b6f9-9428fea2ec56 · outbound

This paper cites Ehinger, Aude Oliva, and Antonio Torralba.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Ehinger, Aude Oliva, and Antonio Torralba

Reference 59

Resolution
verified exact
raw_fallback, observed 2026-08-11T18:15:11.511223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.146569Z digest=sha256:266dcec3c7c5f0726b1bbd8346d6624aba96f644b6e15b7c879e18e03f61d285

Observation 2585dbbb-1bae-4423-9ed7-200b9bfb7d54 · outbound

This paper cites Data selection for language models via importance resampling.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Data selection for language models via importance resampling

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:12.645460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.152454Z digest=sha256:d064c5f915c7a4cc290f4e9357ead5ec61e5936f0b27c991838f3b34af9ca432

Observation 97340a3b-6d14-49f4-b3a1-e7611ffe84e8 · outbound

This paper cites Towards few-shot adaptation of foundation models via multitask finetuning.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Towards few-shot adaptation of foundation models via multitask finetuning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:12.612745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.158372Z digest=sha256:5dd028fa89a9fc06a9606b779ce96fcf74a5d15d6f161dac4b9c8a24105d0d8f

Observation 602dd2c2-18fe-439b-97ff-1b0e5729ae93 · outbound

This paper cites FORML: Learning to Reweight Data for Fairness.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning FORML: Learning to Reweight Data for Fairness

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:11.167948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:11.167948Z digest=sha256:ebb18ff0efa8862c11b8a8f1b8cf912f22efee14ec5925d7b25fe791b4c38207

Observation 223edb0f-e7c1-4a25-abae-67512cab4bec · outbound

This paper cites Adamerging: Adaptive model merging for multi-task learning.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Adamerging: Adaptive model merging for multi-task learning

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:12.582686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.174240Z digest=sha256:508fc4e78ffa8c8d5986cd1e1e84b8e5b84a102cd72502e555884fbf576d9627

Observation d3ace85e-fbb7-45a0-94b8-f3fafe13ce02 · outbound

This paper cites an unresolved cited work.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:11.179853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:11.179853Z digest=sha256:637633d24d0a308a39f8bd51d170f0bba58b20e36165dfa7feb8684d533abb91

Observation be6e1a6e-b594-4cd3-9ea2-f81b0733a5a0 · outbound

This paper cites Character-level Convolutional Networks for Text Classification.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Character-level Convolutional Networks for Text Classification

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:12.548796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.185463Z digest=sha256:f4da487cc2009f184250f80526fcbf0ed4098c93ea001b024d52d19467477e6c

Pith citing papers

Observation 2c191934-7008-4577-9c07-7205e8fa5185 · inbound

Modeling Multi-Task Model Merging as Adaptive Projective Gradient Descent cites this paper.

Modeling Multi-Task Model Merging as Adaptive Projective Gradient Descent Variational Model Merging for Pareto Front Estimation in Multitask Finetuning

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T22:38:49.074989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T22:38:48.719839Z digest=sha256:102b943effb38697a4430ac0c6687c61128dd637cf29aa3b3b56fdc81d45aba8