Pith. sign in

Paper Citation Record · LEDGER

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning

As of 17 August 2026, this Paper Citation Record lists 100 of 111 outbound references and 5 inbound Pith citation observations for arXiv:2411.10928.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.10928 v1

Coverage vector

measured 100 of 111 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:14:24.308625Z

measured 105 of 105 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:46:51.193753Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T16:26:35.777715Z

Reference resolution

100 of 111 outbound references displayed

  • verified exact0
  • verified fuzzy46
  • unresolved54
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation be30b2e9-9e4a-4868-8b6c-02f732616e5c · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:23.860786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:23.860786Z digest=sha256:5539e4965473dfe363a1fb023158d98fcdad62db87efde1306abad87423b8477

Observation a3fa0501-a09c-48c6-a31d-890763afe444 · outbound

This paper cites PaLM 2 Technical Report.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning PaLM 2 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:23.865150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:23.865150Z digest=sha256:e9e696e06b20a308c838f6c6875bb297862e332ffc57409ce60a098f22d1f779

Observation 8a7cb665-c351-49fb-99a2-8fbedfa91bb2 · outbound

This paper cites Composable sparse fine-tuning for cross-lingual transfer.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Composable sparse fine-tuning for cross-lingual transfer

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:23.870108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:23.870108Z digest=sha256:089089e8c76b6af58a63f75f5141be3ac4e93d54668d5a8687b3c2678c1dac55

Observation c2e0e617-22a8-48b4-9a91-54987970165d · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:23.875663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:23.875663Z digest=sha256:2034c278f8d69bc5412c4a23f61efca1e56ff648c47840458c6d0ea818c3858d

Observation 3299ecbb-9ea8-4653-851a-45e03e749649 · outbound

This paper cites Dual lottery ticket hypothesis.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Dual lottery ticket hypothesis

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:23.880372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:23.880372Z digest=sha256:35fdd556b3c30515fc41132db2498c2931ebcb517950cf5098a81faa4d3cd9f6

Observation 34a2b509-13fb-437c-8645-e43ed2b7a846 · outbound

This paper cites Language models are few-shot learners.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Language models are few-shot learners

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:23.884781Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:23.884781Z digest=sha256:1af57ed88325a1b15d31fe2c9b593245d9e44801602edfb74a9fb9ba8e2bb4ac

Observation 41adf560-1d0e-42e8-a83f-51931ff1b7b2 · outbound

This paper cites Dark experience for general continual learning: a strong, simple baseline.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Dark experience for general continual learning: a strong, simple baseline

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:23.891187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:23.891187Z digest=sha256:6cfd83ba6d7cde6774ffba13bfab7db07b635254835fc656b1490d389e26dee7

Observation 140d1801-0c5e-4cf6-87c4-2895c5d12839 · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Honeybee: Locality-enhanced projector for multimodal llm

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:23.896335Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:23.896335Z digest=sha256:dfc5b388cf1617f7ead48a4da8f75cd74993924e4e62f0f0db29a4122ed90407

Observation bcbb0e57-99de-4119-8c4e-33e7801215ed · outbound

This paper cites A unified lottery ticket hypothe- sis for graph neural networks.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning A unified lottery ticket hypothe- sis for graph neural networks

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:23.900351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:23.900351Z digest=sha256:2ea65fdf41225fa0f5a04d60f1398d280612faa5ac99f63962d9b4e2b112db48

Observation 0e2fdcb3-d5ef-4808-b817-ac4f94ad194e · outbound

This paper cites Improved Baselines with Momentum Contrastive Learning.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Improved Baselines with Momentum Contrastive Learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:23.904008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:23.904008Z digest=sha256:07c6c509633212d8c678fc32e3c5f4b48ddc6f01849e63d48ea4de3f55f0dd35

Observation 49e7361f-f641-43e2-bcbd-e3931cee507f · outbound

This paper cites An empiri- cal study of training self-supervised vision transformers.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning An empiri- cal study of training self-supervised vision transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:23.908291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:23.908291Z digest=sha256:256502d2c383d05da130c63a36ab074922e71852905ff3213ad4e91a3b39025b

Observation 2fca8490-ca60-4669-8aa5-17cb9a548933 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:23.912578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:23.912578Z digest=sha256:424905673565f9abb6f54f2f0985fac88ab98fc88868e2e9c7635b65468ba2a0

Observation 880f13ef-c1b8-40d7-a08f-a48ed0189da7 · outbound

This paper cites PaLM: Scaling Language Modeling with Pathways.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning PaLM: Scaling Language Modeling with Pathways

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:23.916966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:23.916966Z digest=sha256:42682cd8f93082e259bd78d63c542523f241fab7eb1ea81e296491c660c0aafb

Observation 7d35fbd7-bf2e-4955-8a99-d7629b36e81c · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Instructblip: Towards general- purpose vision-language models with instruction tuning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:23.921181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:23.921181Z digest=sha256:7bdecbe204f12a69a4a87e3c662ab38638609c1a8e06b14ffd933a108d9d57b5

Observation f67ce704-e37f-4287-90c2-f544b005eb97 · outbound

This paper cites A tutorial on the cross-entropy method.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning A tutorial on the cross-entropy method

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:23.925121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:23.925121Z digest=sha256:682586be9b1a1644946ee70de27f28ff5fc7860e9ad353d123aaa6175449bc18

Observation 8d8bc2a5-c8fe-4345-b03f-003c251cfa24 · outbound

This paper cites An image is worth 16x16 words: Trans- formers for image recognition at scale.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning An image is worth 16x16 words: Trans- formers for image recognition at scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:23.929185Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:23.929185Z digest=sha256:17e1a605d56f4c887e49d755f26db6d8cda14e14fc62ee5836e1427e42a3780d

Observation 165d2a00-4f6f-4736-a360-e570f8084fcd · outbound

This paper cites On the mathematical foundations of theoretical statistics.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning On the mathematical foundations of theoretical statistics

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:23.933616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:23.933616Z digest=sha256:173f6116e92fc1459681a164784060fdcd0aff7f4d46fb34c3517ffe77849e24

Observation 5cfd363e-cb44-497b-a2a8-02550c207f03 · outbound

This paper cites The lottery ticket hypothesis: Finding sparse, trainable neural networks.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning The lottery ticket hypothesis: Finding sparse, trainable neural networks

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:23.939112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:23.939112Z digest=sha256:fb4ef11b9b45bc1f06e5d0a1beaf2674e5c3b36a484b17a0ca216defbaeaf9b7

Observation 0332478c-1441-4cf0-8f4e-ba2dff6475e7 · outbound

This paper cites Sparsegpt: Massive lan- guage models can be accurately pruned in one-shot.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Sparsegpt: Massive lan- guage models can be accurately pruned in one-shot

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:23.945402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:23.945402Z digest=sha256:5134618dd1d3cf05768295a4c46dd34852effb736eb9c1f2504038ea56db5ff6

Observation 51d60eac-ee0a-4656-82aa-691c3e4137ef · outbound

This paper cites Catastrophic forgetting in connection- ist networks.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Catastrophic forgetting in connection- ist networks

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:23.949441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:23.949441Z digest=sha256:e361ad5c02a3779b0e007570f1f1b31feaf82d675ba315d77f84d7540e014ff5

Observation 4b089e94-a2c1-4be0-8eb8-bbdc62f4d484 · outbound

This paper cites Clip-adapter: Better vision-language models with feature adapters.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Clip-adapter: Better vision-language models with feature adapters

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:23.954292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:23.954292Z digest=sha256:f5fd2a98370d734f80be47fd46560f78845686fa37bca785d76e35ebd6c3b7b1

Observation befe1434-20bc-4b71-b35f-a9ddb18c7612 · outbound

This paper cites Making the V in VQA matter: El- evating the role of image understanding in Visual Question Answering.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Making the V in VQA matter: El- evating the role of image understanding in Visual Question Answering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:23.959881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:23.959881Z digest=sha256:4df65d407bc583f908407f92149b72421ff3003c012b3341008c27ea152edd85

Observation 9be2f9d0-4809-473f-9a93-3dc0348ca446 · outbound

This paper cites On calibration of modern neural networks.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning On calibration of modern neural networks

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:23.966177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:23.966177Z digest=sha256:3cd93fb9b9dd4c6e1a6b798ab970c29adeaf76439171924f2c5d527e7332c6b9

Observation f89077bb-50fb-48f0-a856-5b94d0956cee · outbound

This paper cites Learn- ing both weights and connections for efficient neural net- work.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Learn- ing both weights and connections for efficient neural net- work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:23.970084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:23.970084Z digest=sha256:cbe3b9d3251d8f6ee9995308e8ae8391f3a337806a6907f5f78aec54a12311c2

Observation fd1e060d-8e21-4ba9-b1b2-5c579a43095e · outbound

This paper cites Towards attack-tolerant federated learning via critical parameter analysis.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Towards attack-tolerant federated learning via critical parameter analysis

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:23.973605Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:23.973605Z digest=sha256:8ca3a83b0f3a37dca9872759204f18cd32ef5aaf93be007dc4a11ccd43a92f14

Observation 93fd71a6-c836-41cf-bc58-b8192c2735c8 · outbound

This paper cites Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Parameter-Efficient Fine-Tuning for Large Models: A Comprehensive Survey

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:23.976995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:23.976995Z digest=sha256:ae3d54e4842a95649292afc548be1435c29c1ed87d91955530b895bcf552eb17

Observation 8f03c5c4-2d37-4a57-8ab4-bfe23922360e · outbound

This paper cites Flora: Low- rank adapters are secretly gradient compressors.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Flora: Low- rank adapters are secretly gradient compressors

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:23.982093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:23.982093Z digest=sha256:450f35617175b4af54598cc2183a1323762f8e9ec5d1dc8b3a9b3ad0bfddb52b

Observation 185b2326-7f64-44da-9af3-8d8b982c4e9c · outbound

This paper cites Deep residual learning for image recognition.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Deep residual learning for image recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:23.987548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:23.987548Z digest=sha256:9284e7d9aaadd7387b026548d780296b0a7803a81aa2a6c24ef287f862428579

Observation 337672d2-638b-4eb1-8e87-076bd4306a66 · outbound

This paper cites Momentum contrast for unsupervised visual rep- resentation learning.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Momentum contrast for unsupervised visual rep- resentation learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:23.991693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:23.991693Z digest=sha256:56481edaab9ce28b8f2422b4b1fa5c806cd9a391dcaebdb75ab01699cebc4c3a

Observation 3a0bc9ea-ceb7-4f1c-a077-a3aa2399a1f9 · outbound

This paper cites Parameter-efficient transfer learning for nlp.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Parameter-efficient transfer learning for nlp

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:23.996942Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:23.996942Z digest=sha256:46b08b5696a84b8583f61fb4d66664ef8d3c60e5cdcd74a43a0540ad0001df94

Observation ea4a647b-790b-4f42-a317-ddfe672e06fd · outbound

This paper cites Lora: Low-rank adaptation of large language models.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Lora: Low-rank adaptation of large language models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:24.000892Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:24.000892Z digest=sha256:e207028b3f8a34f0eb6f764b266726c86f973307690349c02cfd69f34b252f0c

Observation a601b069-6bb0-491d-8283-bb5cb8bea9e0 · outbound

This paper cites Multi-metrics adaptively identifies backdoors in fed- erated learning.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Multi-metrics adaptively identifies backdoors in fed- erated learning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:24.004962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:24.004962Z digest=sha256:2fa694beed0473f82c18596feb7b51d51f582b5d59b1d5102d80c8d2920043ef

Observation 11fcdbaf-ae16-428f-a158-017cc4decd65 · outbound

This paper cites Fisher calibration for backdoor-robust heterogeneous federated learning.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Fisher calibration for backdoor-robust heterogeneous federated learning

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:24.009524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:24.009524Z digest=sha256:2a0c174508e0ecd9af020b6ab812e80a22b14fd9c554cd081f05d67e93680dbf

Observation a904015f-6503-4336-bf6b-e98200e23d68 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:24.014761Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:24.014761Z digest=sha256:ce75d3c51848dba71c3a94db513eae36da5f60b9e5bc5b7cdfd1ec8d1f3dfe44

Observation 6af8b3fd-360d-452d-af1f-3017f6844728 · outbound

This paper cites HFT: Half Fine-Tuning for Large Language Models.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning HFT: Half Fine-Tuning for Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:24.019028Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:24.019028Z digest=sha256:ed95ac5126486d52a1548a496146c322be71dda4c0a5437286c523de82ea4b1c

Observation 9381ef20-114d-4f06-bd9b-ffa7531fd1dd · outbound

This paper cites Editing models with task arith- metic.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Editing models with task arith- metic

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:24.023164Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:24.023164Z digest=sha256:d322a1d8b4f59a88701b008a49f742dc491fd02cea859aa4e7479029c5f73d14

Observation 9e310663-2297-4bba-b988-38d3234b7658 · outbound

This paper cites Overcoming catastrophic forgetting in neural networks.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Overcoming catastrophic forgetting in neural networks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:24.027262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:24.027262Z digest=sha256:3a8047b88680c840bc13b32e31cc24d45c38a6a147d36a037d426de2e6eae265

Observation 2b856149-b159-4b50-86dc-382c54d2e071 · outbound

This paper cites Simple and scalable predictive uncertainty esti- mation using deep ensembles.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Simple and scalable predictive uncertainty esti- mation using deep ensembles

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.587175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.031099Z digest=sha256:a9117322f5a8cf5148d4e16adb085a2ff5c9d48c3083be57ff453759f4bfa1b0

Observation 562f8f40-ff9b-41a9-bd58-ab8014178959 · outbound

This paper cites Snip: Single-shot network pruning based on connec- tion sensitivity.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Snip: Single-shot network pruning based on connec- tion sensitivity

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.572479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.034944Z digest=sha256:89a3a02b8fd08cc718cf1070801381c7f0ab0245318e64067c581cae964357d1

Observation 0045bbf7-dfac-4dc8-82de-93e7600c256d · outbound

This paper cites The Power of Scale for Parameter-Efficient Prompt Tuning.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning The Power of Scale for Parameter-Efficient Prompt Tuning

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:24.038803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:24.038803Z digest=sha256:4be2c911268ce54c475308162955d4295ad079ac1d41f5bda699c477fa8b8984

Observation 301dd61b-058a-4929-9768-a126095cfd9f · outbound

This paper cites Pruning filters for efficient convnets.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Pruning filters for efficient convnets

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.557945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.043791Z digest=sha256:ebc1fd750d5ac00dffcd2198d276cc6e0258e67b5d30fb6e429b47bdd0a5f108

Observation c9eb82c0-9552-46e5-8530-a6d701c5e0ee · outbound

This paper cites Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Blip: Bootstrapping language-image pre-training for uni- fied vision-language understanding and generation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.543493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.047854Z digest=sha256:9d5a56968b9205a1dc0c54083d835beb78f0cfc7b1baa8b763295a6bb3ec2ace

Observation 549af629-12ce-4edb-bc5a-3b72977545e5 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.528167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.052788Z digest=sha256:6fd84e5106fbdc3e097c3e5af070d61b0dbe8d3c80a2d6d55055f15cd332afb5

Observation 00490ce4-46ba-496b-bd77-5922ecdc46e8 · outbound

This paper cites Federated optimiza- tion in heterogeneous networks.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Federated optimiza- tion in heterogeneous networks

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.513902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.056194Z digest=sha256:922cb7aef45f72a348f12f15309e77622112316a87e038326df7ab556951c62f

Observation b079f96d-50e4-4d81-9655-f32dd53cd4c4 · outbound

This paper cites Graphadapter: Tuning vision-language models with dual knowledge graph.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Graphadapter: Tuning vision-language models with dual knowledge graph

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.500721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.059855Z digest=sha256:747945e8705fb6321b6bff5f491f47c3bdda5d16ea5500c3002fa579955fcfd7

Observation 89d21f91-bdbc-4cbd-9381-6afc0a0b55d7 · outbound

This paper cites Parameter-efficient sparsity for large language models fine-tuning.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Parameter-efficient sparsity for large language models fine-tuning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.487734Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.063108Z digest=sha256:c66900038dd1ab5a3d2026599aaab5e10e92b9eca8e4188360634f6d2d8e5f03

Observation 84f31009-f671-493a-b2ee-63d166aa398a · outbound

This paper cites Losparse: Structured compression of large language models based on low-rank and sparse approximation.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Losparse: Structured compression of large language models based on low-rank and sparse approximation

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.474171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.067029Z digest=sha256:8d14e3e128f293ad835b28a2f5d62c49e6acc78e6f6b7938453f881421e2d313

Observation 332fb438-a9d6-4a27-9e92-00402b29ee05 · outbound

This paper cites Vila: On pre-training for visual language models.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Vila: On pre-training for visual language models

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.460135Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.070563Z digest=sha256:dbc5fde95eec99ad97cf0b20d6d6f90931225463d34f48727c61fd6c1d47d72c

Observation 6b6a2f1c-bbfd-4631-a0c2-eb65c8299cb6 · outbound

This paper cites Microsoft coco: Common objects in context.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Microsoft coco: Common objects in context

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.445693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.075932Z digest=sha256:6a8023c0923ebaea062abe73cbdab51625f00c9ad50d5a738cba2ee26a81a45c

Observation fd29f960-b542-44a2-aa3d-f20932af4828 · outbound

This paper cites Mitigating the Alignment Tax of RLHF.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Mitigating the Alignment Tax of RLHF

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:24.080019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:24.080019Z digest=sha256:0294b911418915c5c1bb1f7d5d549206a582fe83fe5204e1cf640a9517e8d120

Observation e103b47a-a826-465b-80fa-0130d43983a7 · outbound

This paper cites Sphinx: The joint mixing of weights, tasks, and visual embeddings for multi-modal large language models.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Sphinx: The joint mixing of weights, tasks, and visual embeddings for multi-modal large language models

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.431680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.085326Z digest=sha256:bd64a7730bc4ac184ca5be8be86e7b3b4f42091bca2a7e609bc79bb8ec5bc403

Observation 5c2b8b42-ec89-4abb-9da5-513b6624c8d8 · outbound

This paper cites Improved baselines with visual instruction tuning.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Improved baselines with visual instruction tuning

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.417688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.089494Z digest=sha256:ad468f9f294d54f41e44c814e4003d53818b4b1b368f2852ef521e4987e8f322

Observation 68c8f028-249c-4105-b49a-e75d3e075dc0 · outbound

This paper cites Visual instruction tuning.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Visual instruction tuning

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.404358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.094875Z digest=sha256:53f9ca6f2e3363105fd127d99d57501e54e6ddbeb387de415c6d1a88a17849de

Observation 258bab50-016e-4b9d-a221-07b917e04c24 · outbound

This paper cites Dora: Weight-decomposed low-rank adaptation.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Dora: Weight-decomposed low-rank adaptation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.390440Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.099136Z digest=sha256:f88006f6f0d46c17a03cb39cdaecbe5320acfb350ed67cb0803c69fed9528bcc

Observation f6aee091-95ce-462e-9861-0b11508cf13c · outbound

This paper cites Late Prompt Tuning: A Late Prompt Could Be Better Than Many Prompts.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Late Prompt Tuning: A Late Prompt Could Be Better Than Many Prompts

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:24.103422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:24.103422Z digest=sha256:5a3618502f2c8919850628519b7fd15d28f879977a9ad06646c1b4bd178a6426

Observation 77da13c2-7ec8-47fc-b5b8-2ca3b75416cf · outbound

This paper cites Rethinking the value of network pruning.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Rethinking the value of network pruning

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.376626Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.108985Z digest=sha256:2800e887cb5283ed16abef2b054237cd9debbb7e87f8218048190112a0dd98ab

Observation 0bc02d01-d039-47cd-bc38-556cfc4c89d6 · outbound

This paper cites Iconqa: A new benchmark for abstract diagram understand- ing and visual language reasoning.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Iconqa: A new benchmark for abstract diagram understand- ing and visual language reasoning

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.363400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.114015Z digest=sha256:1b1010b964f9ae33bd53bee9b3b5f55ea0a52d6c69c808a7678710db4d1c4af6

Observation b3b39024-5e68-4197-9fe4-ce8ef4a7f311 · outbound

This paper cites Learn to explain: Multimodal reason- ing via thought chains for science question answering.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Learn to explain: Multimodal reason- ing via thought chains for science question answering

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.350282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.119130Z digest=sha256:945ba6d76e09e236ec0d3380821b54518c81b9d2f343d4887e8a230132fc9490

Observation 4391fdf1-0229-494c-8037-e29e67388c21 · outbound

This paper cites Twin-merging: Dynamic integration of modular expertise in model merging.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Twin-merging: Dynamic integration of modular expertise in model merging

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.335056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.123753Z digest=sha256:f0e94b4860ec3db0eedd32d9adc06f1e19b1f97923786af22ea144ba8ee0af5d

Observation 0e935590-da46-4926-8170-3a7215b88e97 · outbound

This paper cites Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Feast Your Eyes: Mixture-of-Resolution Adaptation for Multimodal Large Language Models

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:24.127043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:24.127043Z digest=sha256:3a0f90a97a33af10339c450affd40ed3606c0b70272a6ef44a8b8f76e3ef5fed

Observation 7de2cd66-b31e-47eb-bfff-5504a00c291c · outbound

This paper cites An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning An Empirical Study of Catastrophic Forgetting in Large Language Models During Continual Fine-tuning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:24.130957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:24.130957Z digest=sha256:4201bcb35060ad2abefb5534a145f2862c557657417af91c431eb511078366ed

Observation a4c0d579-5a5a-4b7c-8be6-ea3c65cffef2 · outbound

This paper cites Ok-vqa: A visual question answering benchmark requiring external knowledge.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Ok-vqa: A visual question answering benchmark requiring external knowledge

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.320173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.134572Z digest=sha256:780f69b538c0e4b1a7a8a1e75bc95e501181a60f0093bc24183d234d30f624b3

Observation 659ba5d9-b44c-4fda-b031-1bfc9eeae357 · outbound

This paper cites Merging models with fisher-weighted averaging.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Merging models with fisher-weighted averaging

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.305845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.137798Z digest=sha256:57b588cabcfb7302357ad92d5877da6f4259c75ab2773b1831584251a270717f

Observation c23f4089-3453-467f-b01e-a889e9b213ce · outbound

This paper cites Catastrophic inter- ference in connectionist networks: The sequential learning problem.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Catastrophic inter- ference in connectionist networks: The sequential learning problem

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.291607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.140973Z digest=sha256:870ebd4a7b98b2a70a9368f78f58fd96174a9c284c52f376774e18480c0087c8

Observation f8253858-c3ce-416f-bccb-69c14390841f · outbound

This paper cites Understanding the role of training regimes in continual learning.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Understanding the role of training regimes in continual learning

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.276639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.146222Z digest=sha256:9d12ad3c8da0488ac616c2bb7a4622388b40f8307b6b0984c34ab37f6e139563

Observation e5c153ec-350a-4853-bd07-65f9d3f6475f · outbound

This paper cites Ocr-vqa: Visual question answering by reading text in images.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Ocr-vqa: Visual question answering by reading text in images

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:24.150694Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:24.150694Z digest=sha256:c593eecf4a944f114d4f58d7159817894fe7943f46c894798bf8ecb65eb63884

Observation 51dc8c0b-2199-49cd-9f64-fe115f83a077 · outbound

This paper cites GPT-4 Technical Report.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning GPT-4 Technical Report

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:24.155962Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:24.155962Z digest=sha256:eace671a2ca2660abcb7f32cc5561b142aa5eab4f02148e3a4a7e4a1090216d2

Observation 161c8c1b-d6f2-4e7b-93ec-60d500ea7ab5 · outbound

This paper cites Task-specific skill localization in fine-tuned language models.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Task-specific skill localization in fine-tuned language models

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.253910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.160799Z digest=sha256:efbcae1cce99ab7beee8acae92ad88b47cb99b7846a69cb80559ae435ac6b125

Observation 42759223-e75a-43ad-9cfa-8872dfbf07ee · outbound

This paper cites Revisiting Natural Gradient for Deep Networks.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Revisiting Natural Gradient for Deep Networks

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:24.165923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:24.165923Z digest=sha256:88c8620ee4dbf473499b9da79b0c1ec9bc876aa64ee56a6c99ceec902e4efea5

Observation f0a71c9a-b633-4a84-a052-1153a1b2a924 · outbound

This paper cites Language models are unsupervised multitask learners.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Language models are unsupervised multitask learners

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:24.170904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:24.170904Z digest=sha256:554fb9b395ca4cc71459d2893ca9a5e76393e072ae70f24ed219a153be5ecfec

Observation d5d5807c-02ef-47ec-9caa-6b014fafba66 · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Learn- ing transferable visual models from natural language super- vision

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.228255Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.175583Z digest=sha256:0f9a95ecf2e5971c59f158239405cf7cf97b69e5d5e54d82eda431b5b9f84c62

Observation c9b88928-d685-4dbb-b953-6b4a32e06d71 · outbound

This paper cites Fishr: Invariant gradient variances for out-of-distribution generalization.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Fishr: Invariant gradient variances for out-of-distribution generalization

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.214156Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.180318Z digest=sha256:1d20aca1cac401b6cee16b64b2abf75e267dc75d9539576f44eaf9e44b5d9573

Observation da5346c4-dca3-44f7-be7e-fe421c66043c · outbound

This paper cites Connectionist models of recognition mem- ory: constraints imposed by learning and forgetting func- tions.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Connectionist models of recognition mem- ory: constraints imposed by learning and forgetting func- tions

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.201104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.184508Z digest=sha256:a5634da318bed28611cc2a70d328a53b8285c3ccc2f0288bd0eed202586f9712

Observation 0f0cd4ac-e925-48d5-87ee-06ab97000735 · outbound

This paper cites Online structured laplace approximations for overcoming catastrophic forgetting.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Online structured laplace approximations for overcoming catastrophic forgetting

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.187066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.188822Z digest=sha256:6a3debc985518d12f3466e811efe92bfa93163eb2e2b87dd7bb4bb4b4ce9f7c9

Observation e6bfe690-af15-4444-90ed-29188005dc07 · outbound

This paper cites Move- ment pruning: Adaptive sparsity by fine-tuning.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Move- ment pruning: Adaptive sparsity by fine-tuning

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.171801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.193557Z digest=sha256:2a31aae75eedee835e6606a4db395cbce4b808085e9594f233625f83e61fc916

Observation 0f9e9fab-cd57-4678-80f2-f171aa1798b1 · outbound

This paper cites Test- time prompt tuning for zero-shot generalization in vision- language models.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Test- time prompt tuning for zero-shot generalization in vision- language models

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.156150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.198950Z digest=sha256:2c39564faed92e7e5529204add9c09fbf75a81387c6df236b0659887d29ecaae

Observation c4ec4285-c1e3-47f8-ba22-f0cf516fd282 · outbound

This paper cites Towards vqa models that can read.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Towards vqa models that can read

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.141176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.204147Z digest=sha256:3c1dcf1f353783c17925d25ba7fcd6faa7dcc1168cb2dd7386bcf51d73f3df5c

Observation aad301ce-1e0c-4346-8195-88d0f7562d5a · outbound

This paper cites Sparse is enough in fine-tuning pre-trained large language models.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Sparse is enough in fine-tuning pre-trained large language models

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.125730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.208629Z digest=sha256:b7b8b41c6629bb1da48b489a000895f6ed9ed97b790fcd5a8cc436dd2ccc4178

Observation c315ace5-49c5-4ac7-b587-388a49c94ef5 · outbound

This paper cites PandaGPT: One Model To Instruction-Follow Them All.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning PandaGPT: One Model To Instruction-Follow Them All

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:24.213581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:24.213581Z digest=sha256:882b17bce436960d3d8d357323c19d13e70ffe36cb460ca2e17282756ac83c31

Observation 75069c6b-31c7-4d28-bffa-ef6ba6af21b8 · outbound

This paper cites Prompt Tuning based Adapter for Vision-Language Model Adaption.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Prompt Tuning based Adapter for Vision-Language Model Adaption

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:24.217685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:24.217685Z digest=sha256:d712deccc329b094b7a6bc96436fc26002a187e7d7071c253083e06d23e76479

Observation 1de278e2-41ee-4b83-aefc-3b4d437ee914 · outbound

This paper cites A simple and effective pruning approach for large language models.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning A simple and effective pruning approach for large language models

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.108651Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.221736Z digest=sha256:02ddac91bc7b40ef34f7b3dc68172ec8c65b7f56e8505a1d39e5ce94f103543e

Observation f7fabc27-e9a9-44fb-a4fc-8d4ca641fca1 · outbound

This paper cites Training neu- ral networks with fixed sparse masks.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Training neu- ral networks with fixed sparse masks

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.092396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.225192Z digest=sha256:cdbd3ca6f551ea4fe736d799c76ea58a83fcd5040b7059a06da277aad92da3c5

Observation 2a4f92b0-30e4-4f66-89e2-d71200d24d72 · outbound

This paper cites A Unified View of Delta Parameter Editing in Post-Trained Large-Scale Models.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning A Unified View of Delta Parameter Editing in Post-Trained Large-Scale Models

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:24.228826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:24.228826Z digest=sha256:52c7d77e3c2d8790d0b31bc1bd68dc7c30ba70080359148bc703104254968a46

Observation 9096f3eb-5b03-4b6e-890f-e810978ac5dd · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning LLaMA: Open and Efficient Foundation Language Models

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:24.233723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:24.233723Z digest=sha256:6a9af8a80a2238e3b0b794b17ee9fea833a77beb265240e0f01dd69ec35e7a41

Observation 4ec43a31-c095-45ee-8ccb-909729af90ab · outbound

This paper cites Learning to grow pretrained models for efficient trans- former training.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Learning to grow pretrained models for efficient trans- former training

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.077826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.237852Z digest=sha256:bb78d02126350ebffccd9ba32a8f8dd138f49eee62b405ac23436162b66c0c4c

Observation de609817-0839-46b2-9d83-a8d1e942ecb0 · outbound

This paper cites Sharpness-aware gradient matching for domain generaliza- tion.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Sharpness-aware gradient matching for domain generaliza- tion

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.060595Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.242021Z digest=sha256:6dd203723d305e0c9af74875bed7ba6bbeab76c6f0483ebb5d8fe4fe0226e3a8

Observation 3ddb9813-26ff-435b-8e1f-6d79e2bf3baf · outbound

This paper cites Image as a foreign language: BEiT pretraining for vision and vision-language tasks.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Image as a foreign language: BEiT pretraining for vision and vision-language tasks

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:24.249327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:24.249327Z digest=sha256:64668af7aa0c5e36815a0773c6200471f158ade9e9b975f6e91529dbc615f208

Observation b72930da-b31f-42d5-880b-eff22fe309d6 · outbound

This paper cites Dualprompt: Complementary prompting for rehearsal-free continual learning.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Dualprompt: Complementary prompting for rehearsal-free continual learning

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.036631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.253364Z digest=sha256:a8a539b360950d834fa5c62a75a96f36cae09fb533ece140014f3f94e2a04f30

Observation 27a1ae25-7c70-4e80-9ef1-b6fa5a021559 · outbound

This paper cites Parameter-efficient fine-tuning for pre-trained vision models: A survey.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Parameter-efficient fine-tuning for pre-trained vision models: A survey

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:24.258651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:24.258651Z digest=sha256:e09869ef2b76ceabc5f7efba7093be1cad156445896bf2dd454be2667f275df0

Observation 2906a46b-4204-4e61-b7c0-c106caff9c57 · outbound

This paper cites Explicit inductive bias for transfer learning with convolutional net- works.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Explicit inductive bias for transfer learning with convolutional net- works

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.022669Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.262711Z digest=sha256:4ca1c88e4e25838c5cbf33de72ede84790d6cdfe30432da9efee8ea7a3e57c8c

Observation 353e0553-6540-4302-989c-a0d4a9095cb6 · outbound

This paper cites Lst: Ladder side- tuning for parameter and memory efficient transfer learn- ing.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Lst: Ladder side- tuning for parameter and memory efficient transfer learn- ing

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:25.008628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.266896Z digest=sha256:f72421ff471ea4a346c7d40eb839f143e54d3b10f4f12f453f840cd41ace6294

Observation 0840e68e-4177-462e-bde1-d2b36c2c2616 · outbound

This paper cites Dynamic sparsity is channel-level sparsity learner.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Dynamic sparsity is channel-level sparsity learner

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:24.994241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.272878Z digest=sha256:051bbce5e94c03f7abd7b720e81e982a3fc777dd7ca5eaf4a2772cef0a6a54dd

Observation d0ad8404-cdb8-422b-b269-96e68718d788 · outbound

This paper cites From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:24.979533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.277224Z digest=sha256:09df91711006db582d7788fafd5f734919b11059a1712e02af8ac94c1f05af4a

Observation 446780d8-b57c-42fa-9074-d7772702d193 · outbound

This paper cites Language models are super mario: Absorbing abilities from homologous models as a free lunch.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Language models are super mario: Absorbing abilities from homologous models as a free lunch

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:24.964678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.281783Z digest=sha256:08db0c25620b43b5e04c31122ac4edeb723dde8091361eab389f7d5a24393e8f

Observation 7fa3b420-754e-499b-9489-2fb4b47b4d9c · outbound

This paper cites Unified Vision and Language Prompt Learning.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Unified Vision and Language Prompt Learning

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:24.286924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:24.286924Z digest=sha256:f9a89e8e2162e7a61a658ef37d3226b4a428ac0530cb6b0a06bd518c99c4dfb9

Observation 5b082696-b4a9-49c7-b5a1-4e60a6d6c91a · outbound

This paper cites Contin- ual learning through synaptic intelligence.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Contin- ual learning through synaptic intelligence

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:24.950273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.291959Z digest=sha256:bb9fb5bb18f22075c147685e589cd02d42dccb14d30ae1a5ba1af840f6a3dfc9

Observation 3c36d4ce-b68a-4183-b5fb-c6174babc09e · outbound

This paper cites Investigating the catastrophic forgetting in multimodal large language model fine-tuning.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Investigating the catastrophic forgetting in multimodal large language model fine-tuning

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:24.936431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.295885Z digest=sha256:444997fcf836e89ab4d5a2128a2787997a30ba6a120ba1125a336a47d1c7173a

Observation 47ee7f28-ae11-4bea-8aef-79e8bd5fc203 · outbound

This paper cites Vision-language models for vision tasks: A survey.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Vision-language models for vision tasks: A survey

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:24.922038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.299477Z digest=sha256:0aac12bca59a232b379ae5e4e53da375ae2d83cd47826724901d8f0e701b1865

Observation d2be3cac-1bbf-4a46-8f3f-ace0536753c2 · outbound

This paper cites AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning AdaLoRA: Adaptive Budget Allocation for Parameter-Efficient Fine-Tuning

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-12T19:14:24.303307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T19:14:24.303307Z digest=sha256:5bff39706129e31f26795a694d9195d0b5ee87c364eab4ec7271c80082c4c4ff

Observation 966d5176-9ede-4b53-9aa7-092a323193af · outbound

This paper cites Tip-adapter: Training-free clip-adapter for better vision- language modeling.

Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning Tip-adapter: Training-free clip-adapter for better vision- language modeling

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T19:14:24.908283Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-12T19:14:24.308625Z digest=sha256:a6a973fad0a67b0ba9fdbd01246c2c12992bd0a81daf4eec88776fc2edb816bc

Pith citing papers

Observation c3e44a21-d900-479f-90f7-46f8a9105a47 · inbound

Unleashing the Power of Continual Learning on Non-Centralized Devices: A Survey cites this paper.

Unleashing the Power of Continual Learning on Non-Centralized Devices: A Survey Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning

Reference 275

Resolution
unresolved
no resolver link, observed 2026-08-11T12:46:51.193753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:46:51.193753Z digest=sha256:039b5e72bc6aa1cff4c6b3c6b2e3882cf80eb4c2518b87c60f631690421df615

Observation 50aef3b6-ccc3-44ca-8ad4-2df5516d0ba3 · inbound

A Unified Gradient-based Framework for Task-agnostic Continual Learning-Unlearning cites this paper.

A Unified Gradient-based Framework for Task-agnostic Continual Learning-Unlearning Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T15:28:55.962111Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T15:28:55.962111Z digest=sha256:678ebb004db58581bd78261ea4c00467ea33cad0a2b5c8bd8a56948aa14ef216

Observation e27cb4aa-2d05-4bd2-9e3f-da99a794f8e3 · inbound

Backdoor Cleaning without External Guidance in MLLM Fine-tuning cites this paper.

Backdoor Cleaning without External Guidance in MLLM Fine-tuning Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:56:51.027820Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:56:51.027820Z digest=sha256:5f8fd2df24f99552b48b62deb20d6ec6ee4ef8606b4202060da9f19f219c04d4

Observation 4f5cc45e-8a91-4853-8278-83229e352999 · inbound

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond cites this paper.

Continual Learning for Generative AI: From LLMs to MLLMs and Beyond Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T00:40:23.310658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:40:23.310658Z digest=sha256:5616a203b6564e395c847c6fb4fb9f43892ed4e9bc3a8b41fb0cff451c3a3c96

Observation 3dd389d7-4e46-4f77-95f2-713ec43ba291 · inbound

LoRA-Loop: Closing the Synthetic Replay Cycle for Continual VLM Learning cites this paper.

LoRA-Loop: Closing the Synthetic Replay Cycle for Continual VLM Learning Learn from Downstream and Be Yourself in Multimodal Large Language Model Fine-Tuning

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-06T16:26:35.785326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T16:26:35.174592Z digest=sha256:e87e1c812b6ebde83a9eeb4407dd160296d54ff0e726d0534e85d50590e47c20