Pith. sign in

Paper Citation Record · LEDGER

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning

As of 15 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 1 inbound Pith citation observation for arXiv:2412.08147.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.08147 v2

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T18:15:11.185463Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:38:48.719839Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-10T22:38:49.070684Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact4
  • verified fuzzy41
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 395755ce-a8b7-4cf9-9dfb-08a3d245f1ba · outbound

This paper cites write newline.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:10.730523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:10.730523Z digest=sha256:d311373d4f82e29cec5c1db94c58e1d3324da08951ac911d927b23b6be57ebb5

Observation 964b5bda-5aeb-4d6e-8f45-272fa8877e9f · outbound

This paper cites Muppet: Massive multi-task representations with pre-finetuning.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Muppet: Massive multi-task representations with pre-finetuning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.917923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.737695Z digest=sha256:e20f2bfd1834b63d6f7f3e2e4797312b1647d51f5c2f59a3b282b5042a51404f

Observation 0e954d8c-e2b0-4e09-bfb8-138d5095de3f · outbound

This paper cites Safety-tuned LL a MA s: Lessons from improving the safety of large language models that follow instructions.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Safety-tuned LL a MA s: Lessons from improving the safety of large language models that follow instructions

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.877649Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.743431Z digest=sha256:224ad2db9a305222e2204a098b426f97a0a1af022b351d8fcb79b97c993c0b72

Observation 0b35ded3-ba47-40aa-b1b3-a6976ba82b26 · outbound

This paper cites Neural networks for pattern recognition.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Neural networks for pattern recognition

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.839602Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.750856Z digest=sha256:66397e5d73cac7113238016713029f36567def051a782ca63129b0fbff9901d4

Observation de994e55-adbb-4314-b309-acbb17e2bb7b · outbound

This paper cites an unresolved cited work.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:10.758533Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:10.758533Z digest=sha256:52c833174a3251efae0b7ea94727c48e997c3787f39ef714a526aa21aad4004e

Observation 3d774df1-e3d0-4e8c-9950-48a0ca900174 · outbound

This paper cites Mode-finding for mixtures of G aussian distributions.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Mode-finding for mixtures of G aussian distributions

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.809183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.765228Z digest=sha256:2c5869edca4997f9a75c29ca7aa57ec525b9ab6c9150ab05479d5f426911a656

Observation 18c311c5-d512-4450-b645-bc3a55f4abe4 · outbound

This paper cites Multitask learning.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Multitask learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:10.771952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:10.771952Z digest=sha256:7650d465c80b6844bfd6c8b8c2d5b7ca78c64ac45fb749a6be49b7c95d8c2135

Observation 84cd8a65-06e0-402c-855d-eb11a7205484 · outbound

This paper cites PAC-B ayesian supervised classification: The thermodynamics of statistical learning.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning PAC-B ayesian supervised classification: The thermodynamics of statistical learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.781424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.780121Z digest=sha256:1e8ab7d37e38acd996a9233f90f0a4397bc8a8517e0cd753b43e183a65ab2af7

Observation 6fa95463-b68c-4400-805a-98fa617405d3 · outbound

This paper cites Overview of the IWSLT 2017 evaluation campaign.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Overview of the IWSLT 2017 evaluation campaign

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.739210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.787548Z digest=sha256:d97b4b4c1e56e95ba8b013af53b3b0de6a1029b53c5e00f3dea2d2bf0b44b0d4

Observation 12efde92-4baf-4f4d-9937-f47678ccfd7f · outbound

This paper cites GradNorm : Gradient normalization for adaptive loss balancing in deep multitask networks.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning GradNorm : Gradient normalization for adaptive loss balancing in deep multitask networks

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.704671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.795884Z digest=sha256:e23fb10014872e56ad2c120c7f3be236b386d9ce64ae719bbf88373593dea310

Observation 020759e6-28e3-4368-8781-18109fdb74a8 · outbound

This paper cites Remote sensing image scene classification: Benchmark and state of the art.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Remote sensing image scene classification: Benchmark and state of the art

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:10.805911Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:10.805911Z digest=sha256:fa1aca0d502a34238e98ac456c81eb8c05fd4bb90c212fd0d7226e23ab18a434

Observation b746629e-d543-4e30-a57a-d7ff5ec90582 · outbound

This paper cites Zhao, Yanping Huang, Andrew M.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Zhao, Yanping Huang, Andrew M

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.679730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.812381Z digest=sha256:93ca3945dba70625bb89e7e4145f340f377a0b2a994e01270266901c1f6555a3

Observation 67ca8381-9653-4ef4-a298-33ec930f966a · outbound

This paper cites Model merging by uncertainty-based gradient matching.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Model merging by uncertainty-based gradient matching

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.655628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.820287Z digest=sha256:bcd3a691fa4ca271475cd7239548f9071bedb83bf93fb97c6477792bee0a5b52

Observation 7179bd36-8525-43e5-a84e-3d89372e9c13 · outbound

This paper cites ColD fusion: Collaborative descent for distributed multitask finetuning.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning ColD fusion: Collaborative descent for distributed multitask finetuning

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.623970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.829496Z digest=sha256:bc97f1545b2f56a32b9a612641c84c3e089b1d5ada54ed8f9babaeb792ef31a6

Observation c2553b3d-71bd-40d3-950a-2e55aaecec7c · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning An image is worth 16x16 words: Transformers for image recognition at scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:10.840168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:10.840168Z digest=sha256:e235d6b0a3d12555f674ab3fb522141f4fc6889629faddb957df4a5ac6c3ce08

Observation a46855af-b8df-4a75-b6c4-9bbb0d9283d2 · outbound

This paper cites GLaM: efficient scaling of language models with mixture-of-experts.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning GLaM: efficient scaling of language models with mixture-of-experts

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.576240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.850405Z digest=sha256:8f6a2671ff7125d39619027a14436cf92db5c1b5c809333c6515ef0a0e281e85

Observation 96aab1cc-4881-41ab-a5b4-8666dc84a3f4 · outbound

This paper cites Durrant-Whyte and Mike Stevens.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Durrant-Whyte and Mike Stevens

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.544903Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.858209Z digest=sha256:f1943f1419f97e085c44b616cff0b0384f1ea3dfa561cce40fe8a9de48309385

Observation 7db04345-154f-4cc2-9c3b-8e329aeeceaf · outbound

This paper cites Knowledge card: Filling LLM s' knowledge gaps with plug-in specialized language models.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Knowledge card: Filling LLM s' knowledge gaps with plug-in specialized language models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.519976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.865222Z digest=sha256:998c323f64080e0d4bceb5412520ae97af48da52d5ca2036f7945e397200f284

Observation 82690596-fd30-4d45-8362-8dae36fbafe3 · outbound

This paper cites Continual pre-training for cross-lingual LLM adaptation: Enhancing japanese language capabilities.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Continual pre-training for cross-lingual LLM adaptation: Enhancing japanese language capabilities

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.488184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.873089Z digest=sha256:d448eff8c01fb5345374c2607ef3692ed4fdd608d06038910ecf65e0a2ecb808

Observation 557a0e26-fc38-49c2-9823-140db3640a15 · outbound

This paper cites Gemma 2: Improving Open Language Models at a Practical Size.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Gemma 2: Improving Open Language Models at a Practical Size

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:10.879522Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:10.879522Z digest=sha256:5621b44dd06d96d1b2f2c79e550b31a815ec297008dbed1c6cb8577d25480006

Observation a6795780-5525-4f77-89a1-277feca12c6a · outbound

This paper cites Gemma: Open Models Based on Gemini Research and Technology.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Gemma: Open Models Based on Gemini Research and Technology

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:10.886248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:10.886248Z digest=sha256:c449e6b4672fe19b83a48145bd6cf3d26e92763e80fa2fa41d7f6028cf905fb3

Observation 81bc45e0-b04c-4bba-8969-4468630855b0 · outbound

This paper cites Multi-loss weighting with coefficient of variations.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Multi-loss weighting with coefficient of variations

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.461074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.894295Z digest=sha256:c727130ccc505ab229493db9d2faf60e4e47f3a062311cedd3834c6c42ca3d70

Observation 1183bd3c-63b1-4443-9971-171e25a6ffe4 · outbound

This paper cites Deep residual learning for image recognition.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Deep residual learning for image recognition

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:10.900181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:10.900181Z digest=sha256:157e3cb7192cd885fa11528f668b840f323ef3b902dcae232385153b09ad0b22

Observation c1a94f2a-443b-4dcd-944e-13ad64eb79ec · outbound

This paper cites Euro SAT : A novel dataset and deep learning benchmark for land use and land cover classification.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Euro SAT : A novel dataset and deep learning benchmark for land use and land cover classification

Reference 24

Resolution
verified exact
raw_fallback, observed 2026-08-11T18:15:12.306623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.907746Z digest=sha256:1fbd49e7cb972acece4ea266d01fbb2dc2e5950a43184463d933f83e1895dc10

Observation c0d203fc-989a-4dd8-86ef-e7c0ef3ce1d2 · outbound

This paper cites Detection of traffic signs in real-world images: The G erman T raffic S ign D etection B enchmark.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Detection of traffic signs in real-world images: The G erman T raffic S ign D etection B enchmark

Reference 25

Resolution
verified exact
raw_fallback, observed 2026-08-11T18:15:12.120608Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.913625Z digest=sha256:2d58f17aa8655907264a0fc307ca0ae6248fdb1d07e6b421cd30803c77d30902

Observation 572c2ceb-c933-4c12-ab1f-bd6df9ee96c8 · outbound

This paper cites Lo RA : Low-rank adaptation of large language models.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Lo RA : Low-rank adaptation of large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.434148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.919559Z digest=sha256:6c53a65af6a30c42c72145c6c46a9587e32604dec4d3e1944924658d1e30f415

Observation f5eff7f6-483d-418f-afb6-437de8b88dc5 · outbound

This paper cites Editing models with task arithmetic.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Editing models with task arithmetic

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.413070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.926694Z digest=sha256:1644fe8fe2480996345731a5677b89c657303138d3b622d67065bb9dc5b1f069

Observation 5c23cc78-7eea-49ca-9fe6-e439d8f52599 · outbound

This paper cites Multitask Mayhem: Unveiling and Mitigating Safety Gaps in LLMs Fine-tuning.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Multitask Mayhem: Unveiling and Mitigating Safety Gaps in LLMs Fine-tuning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:10.933257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:10.933257Z digest=sha256:ae83c15bf596d33069defd3184f7f30ce7287ee1311e4ba1b54a91f5d3048727

Observation 7159d9cb-459c-40fc-9ebd-39ce6e83583b · outbound

This paper cites ForkMerge : Mitigating negative transfer in auxiliary-task learning.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning ForkMerge : Mitigating negative transfer in auxiliary-task learning

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.380992Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.940131Z digest=sha256:a3853d3ce25f4cadc69df085958e0b8a2294238ae523fc2c35b324e48319b1e8

Observation 710b5c2c-3aca-4774-b597-a91a5cea8dca · outbound

This paper cites Dataless knowledge fusion by merging weights of language models.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Dataless knowledge fusion by merging weights of language models

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.342746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.947644Z digest=sha256:a0f4786baf7cfbd0c4122ceb64eba2164f71a82de3eb251f3fe33dec5f215279

Observation 52f24e98-7fbd-4fdc-9b91-53599770f4bb · outbound

This paper cites The B ayesian learning rule.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning The B ayesian learning rule

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.301085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.953903Z digest=sha256:03ce14273960e376ebba1742a619409720f8c32142e644e8f5efb3f203d816bd

Observation 8fda7099-3acd-4845-bf22-4f159185ae92 · outbound

This paper cites Fast and scalable bayesian deep learning by weight-perturbation in Adam.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Fast and scalable bayesian deep learning by weight-perturbation in Adam

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.272196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.962011Z digest=sha256:fa3d683d430236da672a281f35c2ca75852c877be7a41a74b2bf4d7bee185afd

Observation 485c367c-903d-41ef-a06c-9ff3e7bb601e · outbound

This paper cites 3D object representations for fine-grained categorization.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning 3D object representations for fine-grained categorization

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:10.970832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:10.970832Z digest=sha256:1053214bc346c357a9678ffc7bd117df5dba6db9e3038fb2c2ac5b6c34514fa4

Observation 9b844140-3071-4f67-aa83-a82d2194c4c7 · outbound

This paper cites Fast and simple natural-gradient variational inference with mixture of exponential-family approximations.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Fast and simple natural-gradient variational inference with mixture of exponential-family approximations

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.242617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:10.976303Z digest=sha256:1004a6a038bdc2f15dcb1f78705e6d723030f2287488b941880bd487f5caeb44

Observation f98c6fc2-88a7-4804-ace3-584a4bf49310 · outbound

This paper cites MFTCoder: Boosting Code LLMs with Multitask Fine-Tuning.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning MFTCoder: Boosting Code LLMs with Multitask Fine-Tuning

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:10.982181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:10.982181Z digest=sha256:c6ab0dc9589cadb980c669a2dd916eaf1ac71cda9448b87339d287ecfd9acf1b

Observation 74a37287-ab42-4f70-94a1-ed5ce9090c0a · outbound

This paper cites RoBERTa: A Robustly Optimized BERT Pretraining Approach.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning RoBERTa: A Robustly Optimized BERT Pretraining Approach

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:10.989779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:10.989779Z digest=sha256:44fc046a2969bbbb2ca0254f72a492cf9329da8ca0d3aee02cb0d7dcef278cdd

Observation 8c533a58-f88e-41f7-9cf2-fbe53f4d06d8 · outbound

This paper cites Decoupled weight decay regularization.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Decoupled weight decay regularization

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:10.996890Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:10.996890Z digest=sha256:4634e2d78465a5821b2d4f57eb7f751f05d167da855f177b1bad256c27ad0d9a

Observation fea2f85d-a94f-4872-ae6f-11c520ffe25c · outbound

This paper cites Maas, Raymond E.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Maas, Raymond E

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.182130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.004139Z digest=sha256:73ec670db6e32b3f0d8c4d055271cfef19895d9cd72d2924282d11a04bbb777b

Observation 0bcd4f09-389d-4ca9-a158-77b305fec2c1 · outbound

This paper cites EuroLLM: Multilingual Language Models for Europe.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning EuroLLM: Multilingual Language Models for Europe

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:11.010961Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:11.010961Z digest=sha256:97b5fbefad10490284b2aa98a54846443340ecc98d8a0a6e54338b55d6a2e9d7

Observation 3fa8acdd-d773-454f-862a-c982f0af8528 · outbound

This paper cites Merging models with F isher-weighted averaging.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Merging models with F isher-weighted averaging

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.145525Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.018531Z digest=sha256:5ddc4a4d782b15b2a98382e870a8bbbf5e14e8792e2b477f154c5fd8bc4dddb0

Observation f98b26df-4d7c-4dab-bb20-a2440de62d48 · outbound

This paper cites an unresolved cited work.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:11.024309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:11.024309Z digest=sha256:d638d1d5d4a846cbbd9a1fc4aed69992ab3c11aaec6478f78697f60b0e47802d

Observation e469d320-642d-4df6-a0ad-027c149e448c · outbound

This paper cites an unresolved cited work.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Unresolved cited work

Reference 42

Resolution
verified exact
raw_fallback, observed 2026-08-11T18:15:11.708518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.030817Z digest=sha256:3024797651dc12b3de2ee8aed7b9335719fafe6a0aa559522a9ed96367735aae

Observation ad718209-7434-42be-a598-4f2e7c83f624 · outbound

This paper cites Reading digits in natural images with unsupervised feature learning.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Reading digits in natural images with unsupervised feature learning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.115920Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.038823Z digest=sha256:d976fdb98caf94a06edfec3b915d131ccf82fab0b1c9ec71cdf115328f7cd835

Observation ac059a11-33d3-430f-8586-0df75f5a2bfb · outbound

This paper cites The variational gaussian approximation revisited.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning The variational gaussian approximation revisited

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:11.045381Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:11.045381Z digest=sha256:29747981c2567f45e44f0d3fa313f0f9446ea520f90771870c87d8e62a253c6b

Observation 5647c61b-0e2a-477b-8edc-4ffd8649f13d · outbound

This paper cites Task arithmetic in the tangent space: Improved editing of pre-trained models.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Task arithmetic in the tangent space: Improved editing of pre-trained models

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.070218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.050804Z digest=sha256:ad3a8631e98e11911188abb4ebb3fec00d6bb956fd0fe267e7e967e05f5a2e66

Observation 9f8a06f5-d6a4-4b8b-9474-c4fde8e006ee · outbound

This paper cites Training language models to follow instructions with human feedback.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Training language models to follow instructions with human feedback

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.042210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.059099Z digest=sha256:a0e4bb600481186b4381dff5c225ac2eb7802791c3d961a90ae8b65e720f5b47

Observation b8e9ae59-116f-4e3a-95ae-c74eb59673b0 · outbound

This paper cites Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Seeing stars: Exploiting class relationships for sentiment categorization with respect to rating scales

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:13.006404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.065388Z digest=sha256:4acea8b06f7687a1893dca44937dd444b0532062416a004f14d181b2fabd32a1

Observation c5436ae4-1b90-417c-ba7d-a69db06a26d5 · outbound

This paper cites Fine-tuning aligned language models compromises safety, even when users do not intend to! In International Conference on Learning Representations (ICLR), 2024.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Fine-tuning aligned language models compromises safety, even when users do not intend to! In International Conference on Learning Representations (ICLR), 2024

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:12.968269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.072473Z digest=sha256:a824211f768ee4887b667035ca59e87ca1ac656e56e7aa57d046e0c41c13f7ba

Observation 6baa1d97-31aa-4c8b-a2f8-46c822a4b443 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Learning transferable visual models from natural language supervision

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:12.943881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.078934Z digest=sha256:e978403f1f00f7b4981a41f8c20e318fe3a9aaf5b82442e290187dec4f05016e

Observation 9eec02c0-f56e-4f67-abef-976c2257cc97 · outbound

This paper cites an unresolved cited work.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-11T18:15:12.907624Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.085942Z digest=sha256:52a98f41d23984add476e5b2f6f1be2daa28480089115520c07f3bf7f1bca294

Observation b0c52ace-ab0d-4a77-bd1a-98b00360b063 · outbound

This paper cites Learning to reweight examples for robust deep learning.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Learning to reweight examples for robust deep learning

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:12.877091Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.093757Z digest=sha256:c625fddf939f2e60faad888a6590e53d37634dd7a5a8e76fdcd0112590dd3b4a

Observation bd4f4ced-f41a-442a-b3a5-228868419e38 · outbound

This paper cites An Overview of Multi-Task Learning in Deep Neural Networks.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning An Overview of Multi-Task Learning in Deep Neural Networks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:11.101238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:11.101238Z digest=sha256:9af494dd0deb051758ea10f1d0139f373060e8e3c1f71f900f09febebec2a958

Observation 19f589c9-c697-428a-a2d7-deafa1fc82c7 · outbound

This paper cites Variational learning is effective for large deep networks.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Variational learning is effective for large deep networks

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:12.841471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.107479Z digest=sha256:3e9c5c18df8b0dd7976e3bd5079e6bba59458e315a260b023df0a511c76767fc

Observation 5450d75e-ee04-4d28-bdd3-7c9bff97333b · outbound

This paper cites Manning, Andrew Ng, and Christopher Potts.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Manning, Andrew Ng, and Christopher Potts

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:12.801207Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.112906Z digest=sha256:12b94f1f643ec5bbfe97c8d67cea15b4b9cb1b035d2b67fba97319b965a98de6

Observation 8e03159c-696c-4e1c-b959-13f944dabfaf · outbound

This paper cites ZipIt! M erging models from different tasks without training.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning ZipIt! M erging models from different tasks without training

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:12.771096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.119063Z digest=sha256:831bec12cf9966ec1d523a2c9cd827d4340af1a01b9a31bca8a9ce4dd32e8c36

Observation 9a1db191-e886-4f29-84d9-efceb07a0918 · outbound

This paper cites Self-influence guided data reweighting for language model pre-training.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Self-influence guided data reweighting for language model pre-training

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:12.743621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.124939Z digest=sha256:a843084ab068f98e4f33853fdde3735f4f1db7c077fdbf219e3012e86d799eb9

Observation b02077f1-2c83-4e16-9927-6af9930972b8 · outbound

This paper cites A B ayesian committee machine.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning A B ayesian committee machine

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:12.706744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.131511Z digest=sha256:c88f7b9ce877213d66b3e8a405c2ab2f196eae2392c57835c471f4d828f8c3e9

Observation c2478687-5d2f-4d51-b170-201bdb0ea5cf · outbound

This paper cites Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:12.674500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.139994Z digest=sha256:fad46756c2956dc21e28b66be117f57c668bc4f745046d7a5c3016178e10fc46

Observation cfd7f9dc-8af2-40c3-b6f9-9428fea2ec56 · outbound

This paper cites Ehinger, Aude Oliva, and Antonio Torralba.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Ehinger, Aude Oliva, and Antonio Torralba

Reference 59

Resolution
verified exact
raw_fallback, observed 2026-08-11T18:15:11.511223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.146569Z digest=sha256:994b78996b6834fbc080b16b51802778d30d2e0867bea1681b7fbbf77a15c2c1

Observation 2585dbbb-1bae-4423-9ed7-200b9bfb7d54 · outbound

This paper cites Data selection for language models via importance resampling.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Data selection for language models via importance resampling

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:12.645460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.152454Z digest=sha256:3a39a2e8c5959774f8be4ec2c2c06945ada68562a289bf890c0b6bc0b891c0d6

Observation 97340a3b-6d14-49f4-b3a1-e7611ffe84e8 · outbound

This paper cites Towards few-shot adaptation of foundation models via multitask finetuning.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Towards few-shot adaptation of foundation models via multitask finetuning

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:12.612745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.158372Z digest=sha256:08e0d6bf3d05dbade6ebb3ea468a1ae4dba9351c1bd595814e31f29769fea144

Observation 602dd2c2-18fe-439b-97ff-1b0e5729ae93 · outbound

This paper cites FORML: Learning to Reweight Data for Fairness.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning FORML: Learning to Reweight Data for Fairness

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:11.167948Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:11.167948Z digest=sha256:11c1830aaafe725e08ef377b4517a30d1e4016fa179b12f0b3f467070e92f88e

Observation 223edb0f-e7c1-4a25-abae-67512cab4bec · outbound

This paper cites Adamerging: Adaptive model merging for multi-task learning.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Adamerging: Adaptive model merging for multi-task learning

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:12.582686Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.174240Z digest=sha256:8fef68e962fc7e8929f88cbcec15ac2c811306ef9ea3f95cae473b936b5ccfbd

Observation d3ace85e-fbb7-45a0-94b8-f3fafe13ce02 · outbound

This paper cites an unresolved cited work.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Unresolved cited work

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-11T18:15:11.179853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T18:15:11.179853Z digest=sha256:d14d4b9d72399abf3879a38727860d90b2c02d68b96f5ad2e42315e269963001

Observation be6e1a6e-b594-4cd3-9ea2-f81b0733a5a0 · outbound

This paper cites Character-level Convolutional Networks for Text Classification.

Variational Model Merging for Pareto Front Estimation in Multitask Finetuning Character-level Convolutional Networks for Text Classification

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T18:15:12.548796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-08-11T18:15:11.185463Z digest=sha256:a2c3d7bad4d09c6a15f76edd5a2b28d1365dc4d21cb3b46350d17c8b3feaa352

Pith citing papers

Observation 2c191934-7008-4577-9c07-7205e8fa5185 · inbound

Modeling Multi-Task Model Merging as Adaptive Projective Gradient Descent cites this paper.

Modeling Multi-Task Model Merging as Adaptive Projective Gradient Descent Variational Model Merging for Pareto Front Estimation in Multitask Finetuning

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T22:38:49.074989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-10T22:38:48.719839Z digest=sha256:7ab70c3c386c236ce95daae906f6d6fd3d0bcc20b82549f9f5caa8faec1ccf6c