Pith. sign in

Paper Citation Record · LEDGER

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models

As of 17 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 0 inbound Pith citation observations for arXiv:2508.00260.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.00260 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T10:21:21.500180Z

measured 55 of 55 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

55 of 55 outbound references displayed

  • verified exact1
  • verified fuzzy39
  • unresolved15
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation a2cd95b2-bd47-4f87-863c-e5546ad593d3 · outbound

This paper cites GPT-4 Technical Report.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models GPT-4 Technical Report

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T10:21:21.310768Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:21:21.310768Z digest=sha256:9f4d7e0c74839664cf0eac8def1a8cb50ef195a50d2347cb9e2cfa982156289e

Observation 8c6088e5-da7c-4619-bf4c-78bc5c5ff8be · outbound

This paper cites Expert gate: Lifelong learning with a network of experts.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Expert gate: Lifelong learning with a network of experts

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:22.067100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.315076Z digest=sha256:b0cf472fe568f6f31d8d2ecdc6bb082e4ed70c44853925ccf14f6afa7339542b

Observation 99ba7021-4be1-4c8b-ab97-3ef594dff23e · outbound

This paper cites Memory aware synapses: Learning what (not) to forget.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Memory aware synapses: Learning what (not) to forget

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:22.057628Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.318768Z digest=sha256:ce2c227391bc92ca909aa10aeceafe5952b6947ba1039af964b966d0d307695b

Observation f62ad9ad-b7c2-4400-94b4-e6169e020360 · outbound

This paper cites Generative multi-modal models are good class incremental learners.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Generative multi-modal models are good class incremental learners

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:22.047700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.322689Z digest=sha256:0ff17b4472626bc429a76bfdf04bbf0142bea1030db84ea105793d0cbc64fa60

Observation 8e8e1114-3b47-4343-a4fc-008ddfec2dd0 · outbound

This paper cites Honeybee: Locality-enhanced projector for multimodal llm.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Honeybee: Locality-enhanced projector for multimodal llm

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:22.037476Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.326048Z digest=sha256:823e6e99a630f75106ad41bffe2aaee1e839519ad15015bc8547cb44a92526fa

Observation 570d82c6-cebe-42ef-a1e6-21476761a094 · outbound

This paper cites Riemannian walk for incremen- tal learning: Understanding forgetting and intransigence.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Riemannian walk for incremen- tal learning: Understanding forgetting and intransigence

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:22.027727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.329716Z digest=sha256:534489f344b090e16bb25115333c1ff93177952d73a12483c603c24801c28ddc

Observation 30788f19-cb92-480b-8fbd-a5ac6691d02f · outbound

This paper cites CoIN: A benchmark of continual instruction tuning for multimodel large language models.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models CoIN: A benchmark of continual instruction tuning for multimodel large language models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:22.017546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.333291Z digest=sha256:1c3484770931d48d32e134e6ce79ce661933aabc91fa47cc11e025f2d58455bc

Observation dd7b02c6-1581-4797-8a2c-1351ec226a56 · outbound

This paper cites Lifelong language pretraining with distribution-specialized experts.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Lifelong language pretraining with distribution-specialized experts

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:22.007441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.336761Z digest=sha256:0b5b9b7acdfb54a04e8094c8d07f7f7d932095dd699ead94ed13dd0af186b62a

Observation 81a25914-8eaa-4039-a0dc-17fafc13e101 · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, march 2023.URL https://lmsys.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality, march 2023.URL https://lmsys

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.997478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.340156Z digest=sha256:eb949f753e7eb09d4ff2ae58d96c258ae0fa48c18e475bcf7fd51af4e75db295

Observation 945b7b39-0945-4280-a61a-bdaa7a41fb4d · outbound

This paper cites V ocabulary-free image classification.Advances in Neural Information Processing Systems, 36:30662–30680, 2023.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models V ocabulary-free image classification.Advances in Neural Information Processing Systems, 36:30662–30680, 2023

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.987564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.343385Z digest=sha256:26f6c8f80c601c6f84503921836179e07881c58d1fdadc160336596844ea780a

Observation 2ab47047-23b1-4284-bff6-c519341f81e7 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T10:21:21.347451Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:21:21.347451Z digest=sha256:58b4bbf4399d2e51da47ad438bb16d4864544591a162993ba65adeae34fbb21d

Observation 06a9d7c0-b14e-4747-883b-cdd02f613800 · outbound

This paper cites RATT: Recurrent attention to transient tasks for continual image captioning.Advances in Neural Information Processing Systems, 33:16736–16748,.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models RATT: Recurrent attention to transient tasks for continual image captioning.Advances in Neural Information Processing Systems, 33:16736–16748,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.976443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.351302Z digest=sha256:693b8973117e2d76e05183d3467c269cf7e03c48a5017e221b78502cfbb1789d

Observation a6ea476f-f29f-493b-a247-ed1bf7a12bc6 · outbound

This paper cites DyTox: Transformers for continual learning with dynamic token expansion.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models DyTox: Transformers for continual learning with dynamic token expansion

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.965024Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.355162Z digest=sha256:81ddae00cd93ab636c6f37a7a720b7cd81ca000e1fd2fefc6453664523c3a902

Observation 45fca5ca-69f9-48be-8db1-23abe7e24f42 · outbound

This paper cites EV A: Exploring the limits of masked visual representa- tion learning at scale.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models EV A: Exploring the limits of masked visual representa- tion learning at scale

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.954607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.359574Z digest=sha256:6f08a027992832c70a9df7697ae8fd34255c098bdf43f226253e66f3889b33d0

Observation b22df089-9ead-47e1-a943-6e6380b1d531 · outbound

This paper cites Beyond prompt learning: Continual adapter for efficient rehearsal-free continual learning.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Beyond prompt learning: Continual adapter for efficient rehearsal-free continual learning

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.944646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.363466Z digest=sha256:7ffd39b3256fd57df603020e1c90c78d365b8ac32bfac22c7e51f7508cdf857b

Observation 02372aff-844f-4e71-8370-72ec0917cdeb · outbound

This paper cites Continual Instruction Tuning for Large Multimodal Models.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Continual Instruction Tuning for Large Multimodal Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T10:21:21.366736Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:21:21.366736Z digest=sha256:27808fbed201c047a595a3db8d5a0500d8158f8b0c5de8a520ba744ca7983b95

Observation 0bee0757-8a2a-4ee7-a506-1be9dca4f475 · outbound

This paper cites The many faces of robust- ness: A critical analysis of out-of-distribution generalization.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models The many faces of robust- ness: A critical analysis of out-of-distribution generalization

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.934456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.370202Z digest=sha256:ba0b274db4ca9a3af9197cf07a2b0c325467e6a3aa09e32fdfa7d6f9a01c636d

Observation b20eeacf-6977-4ed0-ad91-0c2f538a8ce3 · outbound

This paper cites CLIPScore: A reference-free evaluation metric for image captioning.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models CLIPScore: A reference-free evaluation metric for image captioning

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.924153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.373331Z digest=sha256:1812cc5c839c4bd8bb935ffbe1062026a690969ba4bbe272d3665917c5f23c5f

Observation 3bdd2268-3957-48f4-88f7-8181ab064744 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models LoRA: Low-Rank Adaptation of Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T10:21:21.376708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:21:21.376708Z digest=sha256:838304055cf0c2b29e6e76f41c822a28fbe5d003ad7391d4965865c4be7693ac

Observation 5326a5b6-36e1-4a81-92ee-3c77860b2220 · outbound

This paper cites Adaptive mixtures of local experts.Neu- ral computation, 3(1):79–87, 1991.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Adaptive mixtures of local experts.Neu- ral computation, 3(1):79–87, 1991

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T10:21:21.380456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:21:21.380456Z digest=sha256:cee7642c1a927f7c08b5c146a1f665b7255c1659d97033fb08de5cb79a69489c

Observation 8a380d33-a13b-4e35-94f8-41dbb652a36c · outbound

This paper cites Vi- sual prompt tuning.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Vi- sual prompt tuning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.908229Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.384416Z digest=sha256:e7bb0084fc14b482f99a60eea97b2edbcc3c99f98b637c0de1038bdffa252ee3

Observation 4046f619-d310-4c43-916e-7e154fa85702 · outbound

This paper cites Helpful or harmful: Inter- task association in continual learning.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Helpful or harmful: Inter- task association in continual learning

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.898082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.387520Z digest=sha256:8ff7e2f9212536ffe9d3d037ed7f920831be0bebf20e4a7f5dcf8234e2cc9ed8

Observation c2627989-9973-45ca-8420-7fcadd3372a3 · outbound

This paper cites Growing a brain with sparsity-inducing genera- tion for continual learning.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Growing a brain with sparsity-inducing genera- tion for continual learning

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.888147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.391038Z digest=sha256:7c2cef7e981313c6b7dce8bfcd9ab5f1dbc43a5417fac9ca546f5c94b173d2b5

Observation 263528cf-4529-4fb3-92b3-b84b4b146685 · outbound

This paper cites Deep visual-semantic align- ments for generating image descriptions.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Deep visual-semantic align- ments for generating image descriptions

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.877582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.394628Z digest=sha256:4b6ff86465791da1862484f74e5c09858a0f3bec46001cdfc7a7318a2535166f

Observation 6ad4d4fa-f5d2-4bd1-a32d-e55d25b24242 · outbound

This paper cites Overcoming catastrophic forgetting in neu- ral networks.Proceedings of the National Academy of Sci- ences of the United States of America, 114(13):3521–3526,.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Overcoming catastrophic forgetting in neu- ral networks.Proceedings of the National Academy of Sci- ences of the United States of America, 114(13):3521–3526,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.867258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.397882Z digest=sha256:35c2529c9959420fc9142442e4540b8571a41ca9598660b3559279ccdfeba28e

Observation 2d94a6d3-33dc-493a-b215-ac027f494ba4 · outbound

This paper cites Quantifying the Carbon Emissions of Machine Learning.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Quantifying the Carbon Emissions of Machine Learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T10:21:21.401198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:21:21.401198Z digest=sha256:5978f5df3b551651ef0386519acc09e05d18bb2ee122f0c2a3d5104e0f828dd5

Observation e97382b7-74b5-40ff-9acb-2aa6f612a842 · outbound

This paper cites GShard: Scaling giant models with conditional computation and automatic sharding.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models GShard: Scaling giant models with conditional computation and automatic sharding

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.857066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.404660Z digest=sha256:d2cd497e4b051927f57c520c11a34df9997713fe60db8454e5c67f78e0ffb560

Observation e1280ef7-e5aa-4e2e-afc9-d1a6581f1682 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T10:21:21.408225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:21:21.408225Z digest=sha256:18a0f7e0cb8cd121651d6ce7e4e1a4229848cd2a1657013d4a58a376ac19a77e

Observation ba8b8395-1e75-4d1a-b2e2-4241ffd20f49 · outbound

This paper cites BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models BLIP-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.845884Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.411788Z digest=sha256:65b36255c96b6ad092c92f303340f3c27546c54ad8fc033cd80ee74b5fe1dd34

Observation 4f8163e8-6e42-49a3-ae1f-3b14c7fd1d8b · outbound

This paper cites Learning without forgetting.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Learning without forgetting

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.835870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.414960Z digest=sha256:e71204463f60be7171f3b470d8f5d25a70acb8d4a76982b10fcaf2783458d776

Observation 3ea7bd6f-3ff7-43e2-baac-83a22dff776b · outbound

This paper cites InfLoRA: Interference-free low-rank adaptation for continual learning.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models InfLoRA: Interference-free low-rank adaptation for continual learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.825621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.418162Z digest=sha256:a0d0dac1cf73ada62d553ccffa9a2c35dd478f0ae7949d809726215f42dafd40

Observation 444189ce-f650-444a-8d50-5e51921f735a · outbound

This paper cites Improved baselines with visual instruction tuning.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Improved baselines with visual instruction tuning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T10:21:21.421578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:21:21.421578Z digest=sha256:54fff6233cb9d2581393e9fa68b2aad47f3baf775040712c6d5ccf4e62f100c8

Observation 1e4fc8c2-bcee-4400-86e1-ce01a0297c89 · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T10:21:21.424576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:21:21.424576Z digest=sha256:6a0bb79395615258b47fdfd60ae5a11af2a743547c6727156f72a93146a4d247

Observation 2d6433d5-c150-47ed-b400-6d06c5ff7eab · outbound

This paper cites Adaptive aggregation networks for class-incremental learning.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Adaptive aggregation networks for class-incremental learning

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.802726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.427663Z digest=sha256:f33b3e5a43c79ed9cd4eb164c83fcad39d2400bd7ea234e4aabaeea3adb88ac4

Observation 11518c7f-6194-43b9-b38f-a2c4eae7dfa6 · outbound

This paper cites Unified-IO 2: Scaling autoregressive multimodal models with vision language audio and action.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Unified-IO 2: Scaling autoregressive multimodal models with vision language audio and action

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.791017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.430584Z digest=sha256:4256036b7ecd4dd81d44f1141e70e494d5373d88a0b7779c935d8d713883fa8b

Observation 61fee644-8cda-467b-9d08-33eb3dd9dfbb · outbound

This paper cites Not all ex- perts are equal: Efficient expert pruning and skipping for mixture-of-experts large language models.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Not all ex- perts are equal: Efficient expert pruning and skipping for mixture-of-experts large language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.779647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.433965Z digest=sha256:bc596992739df043f2959b187bc6cda84615c170f532b643ee51d26ec9989fe3

Observation cf2acef6-ec07-4ca3-a677-6d4a9ef772d6 · outbound

This paper cites Catastrophic inter- ference in connectionist networks: The sequential learning problem.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Catastrophic inter- ference in connectionist networks: The sequential learning problem

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.768403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.437007Z digest=sha256:9183524957f06da8c3b53f0bc50f5d1de97af3fc7072cfe10d7831e836b062af

Observation 512af70e-e130-4b8b-a7f1-7c739dff4727 · outbound

This paper cites Multimodal contrastive learn- ing with limoe: the language-image mixture of experts.Ad- vances in Neural Information Processing Systems, 35:9564– 9576, 2022.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Multimodal contrastive learn- ing with limoe: the language-image mixture of experts.Ad- vances in Neural Information Processing Systems, 35:9564– 9576, 2022

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.758311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.440162Z digest=sha256:6757893f3b8bba7702c61d23c879c421213092df5f9df1e15be0ad00313b2549

Observation c68e684d-ac33-46b2-aa90-aa2bb7345d87 · outbound

This paper cites Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models.Interna- tional Journal of Computer Vision, 123:74–93, 2017.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Flickr30k entities: Collecting region-to-phrase corre- spondences for richer image-to-sentence models.Interna- tional Journal of Computer Vision, 123:74–93, 2017

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.748323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.443261Z digest=sha256:e5133edb82920e950f0d9818317edf74cd99906880b03b259cda8293d74bb8b2

Observation fa114b31-03a8-4d3c-bf60-47870e7f9285 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Learning transferable visual models from natural language supervi- sion

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T10:21:21.446302Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:21:21.446302Z digest=sha256:fc1a5894b76dee86d88c2a007a2be25d0f3dbec138076083e924dbcdd6811869

Observation 7dc32692-beca-432e-8e4b-17fde4ca2554 · outbound

This paper cites iCaRL: Incremental clas- sifier and representation learning.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models iCaRL: Incremental clas- sifier and representation learning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.730959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.449652Z digest=sha256:4483245217a74e81f15f6cec1a396cd5dd2024262e17c395f8a421916833ab5a

Observation b83644da-00c9-4940-8e7b-e4593c362396 · outbound

This paper cites Exploring models and data for image question answering.Advances in neural information processing systems, 28, 2015.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Exploring models and data for image question answering.Advances in neural information processing systems, 28, 2015

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.720656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.452679Z digest=sha256:8eacc478252dc881a3545a800dab93e404519aec1d3f6284afd190fcb232617a

Observation aa84c749-4020-40f3-b455-7c2192d7ba52 · outbound

This paper cites Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Outrageously Large Neural Networks: The Sparsely-Gated Mixture-of-Experts Layer

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T10:21:21.455817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:21:21.455817Z digest=sha256:d767cc72db2a2130bb4b09427d43da6f7d96b3a0e8ba7354b0a4e819d05b8c48

Observation 90f5229b-c3f4-48f0-8f8f-30b2dbec25f2 · outbound

This paper cites CODA-prompt: Con- tinual decomposed attention-based prompting for rehearsal- free continual learning.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models CODA-prompt: Con- tinual decomposed attention-based prompting for rehearsal- free continual learning

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.710769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.459384Z digest=sha256:6230f21c8d4ca5c69e2456e66d3d16c16b2bd5fe71976e192cd61491dd4aded1

Observation b5ad4abf-dce0-4b27-9427-29e05c0955ae · outbound

This paper cites Constrained contrastive distribution learning for unsupervised anomaly detection and localisation in med- ical images.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Constrained contrastive distribution learning for unsupervised anomaly detection and localisation in med- ical images

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.700593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.463488Z digest=sha256:8c55c45d23d5266f0532533ea2db5983f80af7adfa8405cfa62862f47bb727fb

Observation 1655feca-8306-4c3c-b6c9-47273b294f71 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T10:21:21.467333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:21:21.467333Z digest=sha256:75a9af2d6c0fdea49f33a5b6bdbf0322aa871650cee2157c3832aba7aafae3ce

Observation 3da1cda1-5e60-4ce3-85fe-caed156eca0f · outbound

This paper cites DualPrompt: Complementary prompting for rehearsal-free continual learning.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models DualPrompt: Complementary prompting for rehearsal-free continual learning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.690408Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.470608Z digest=sha256:372992c11b8ba000b67a3317292ff55376b138d95ca1c1c984e398094ed35d99

Observation e34c4a4b-b655-4c20-8fb9-5dba204c4ec0 · outbound

This paper cites Learning to prompt for con- tinual learning.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Learning to prompt for con- tinual learning

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T10:21:21.474330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:21:21.474330Z digest=sha256:4351502ccc381bf1663e0c02aaa886262337152e3c28b85bce9b5d59e76789aa

Observation 85a0656a-f59d-4f0a-96d5-9010d31056c4 · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-06T10:21:21.478032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:21:21.478032Z digest=sha256:ec0cc20fe558b019f4e1e4a73a188276742e6f63cc52ecb5e6fba412b9ed18bd

Observation 0e2c69be-6c75-40ee-a6d0-ba572582c434 · outbound

This paper cites Boosting continual learning of vision-language models via mixture-of-experts adapters.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Boosting continual learning of vision-language models via mixture-of-experts adapters

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.674647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.481597Z digest=sha256:b4bc809e6d98792f9b24cde7214f43f6dcad7b2740dd1e4a1703251ed008b543

Observation 15c32d4d-2942-4925-a2f2-e79a44d2b491 · outbound

This paper cites Contin- ual learning through synaptic intelligence.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Contin- ual learning through synaptic intelligence

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.664041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.485352Z digest=sha256:63839f8c3cc6c1fc0017de205405af59851b3d38c02438012bb91d17ec0ad8de

Observation 34554ce1-1655-4f89-9000-5f55e38d0500 · outbound

This paper cites Investigating the catastrophic for- getting in multimodal large language models.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Investigating the catastrophic for- getting in multimodal large language models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.653638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.489745Z digest=sha256:c7d64c3c6883a552ae19b4092b900ba7c03b114634a8a698c6aa949d0d611774

Observation 128e48f9-fff9-4e9f-b946-457d78234640 · outbound

This paper cites Prompt-Aware Adapter: Towards Learning Adaptive Visual Tokens for Multimodal Large Language Models.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Prompt-Aware Adapter: Towards Learning Adaptive Visual Tokens for Multimodal Large Language Models

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-08-06T10:21:21.546029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.493130Z digest=sha256:00bf2a07baaa055da3d3e214eaac5bef30ad4292853bae1f9862643c17dfa4c1

Observation 7893dd5b-5c4d-451e-a26d-18c1465b73f8 · outbound

This paper cites Preventing zero-shot transfer degradation in continual learning of vision-language models.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models Preventing zero-shot transfer degradation in continual learning of vision-language models

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:21:21.642002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T10:21:21.496493Z digest=sha256:b11cf9a7d10251fecf5935b0383c756f5562fcc27dcd2e79ac27681aba615152

Observation 904ad81f-2db6-4e47-bd45-2d68bd2ff969 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

Instruction-Grounded Visual Projectors for Continual Learning of Generative Vision-Language Models MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T10:21:21.500180Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:21:21.500180Z digest=sha256:1b5b1c295acad7f4d6c37d9799926607db5d157cd6a151fa58806dfc219e5d10

Pith citing papers

No inbound Pith citation observations are available.