Pith. sign in

Paper Citation Record · LEDGER

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL

As of 15 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2608.13385.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.13385 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T12:25:20.568189Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fc017263-2b7e-4fe4-93c7-c253e23ee134 · outbound

This paper cites What learning algorithm is in-context learning? Investigations with linear models.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL What learning algorithm is in-context learning? Investigations with linear models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.392114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.392114Z digest=sha256:5541acc52dbcdf52dbd8c8367574d9eded6e4f10803cc4f41f8a14452ff600bd

Observation bcc7d6d6-f8cd-4873-928b-0406c42d5a51 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.398327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.398327Z digest=sha256:aeab1d866ec2ecde8767c34c9bb2bc3a8ecd4e5d73ed424f6c3a61db5cba04bb

Observation c8e11e3a-5d18-462f-bff5-ce47c2745a9b · outbound

This paper cites an unresolved cited work.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:25:21.110026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T12:25:20.404342Z digest=sha256:64d8c7ae5d97da8b0f8fd37a880240c39a4fa1a4df2e934b5dfef2507331318d

Observation 52e97647-f601-4bbc-a8ab-cd8b628a1c2a · outbound

This paper cites Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.408621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.408621Z digest=sha256:b37839deb29334ed83a0320eb9a14833c49f091c517023379516b86ccefcfcc2

Observation aba977d2-d585-44a4-b3fc-ef3e7e08c83d · outbound

This paper cites Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.413368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.413368Z digest=sha256:3ed4acf87b7629474a99cdfdc15ece4d3037434a9d3e9025057a1a3154604553

Observation 2c95a73f-4b83-49f8-8fbd-c469f901babc · outbound

This paper cites Towards Multimodal In-Context Learning for Vision & Language Models.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Towards Multimodal In-Context Learning for Vision & Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.419246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.419246Z digest=sha256:23920af2afa7989dcb0c27f20c4613775d52b33b7765f896c6cf87502140d0fe

Observation bbaafca0-6a2d-4307-b0b9-f085588084db · outbound

This paper cites Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.428086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.428086Z digest=sha256:03aa57816e68ed01177f2f133e2948c84def73f64692e4a3a28af4e7b802d3b2

Observation 3d011afc-5792-4370-b720-190f88fc958f · outbound

This paper cites In-Context Learning Creates Task Vectors.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL In-Context Learning Creates Task Vectors

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.433188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.433188Z digest=sha256:a08d9fb6b934e9235e0cab00aba219fa88f084809c6eda84540f7fb7e7505cae

Observation 6f03b30b-636b-4216-abed-a42feae33873 · outbound

This paper cites Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.439241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.439241Z digest=sha256:23ee73310cbeb1350a63cbc08c36dd01de319a592284a73e2ae0d5c0ebf4b256

Observation 537b46df-ed04-42af-9955-4bd386dc975d · outbound

This paper cites Dissecting Multimodal In-Context Learning: Modality Asymmetries and Circuit Dynamics in modern Transformers.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Dissecting Multimodal In-Context Learning: Modality Asymmetries and Circuit Dynamics in modern Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.444342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.444342Z digest=sha256:9353c85e4daac4c2b14d7f69577df66b674ceaf18cec66dad31f23fa568eef4c

Observation bab64fb7-04fb-414b-ba00-4fea119a867a · outbound

This paper cites an unresolved cited work.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.450842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.450842Z digest=sha256:26052b86c7ffaff430fd8a30f5d2430a1ff0be149e60c3ab3fec0ef4e00a7e20

Observation d7a316ab-4e05-4b8d-a21a-67efc54b3ae6 · outbound

This paper cites Mimic In-Context Learning for Multimodal Tasks.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Mimic In-Context Learning for Multimodal Tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.456349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.456349Z digest=sha256:7ce1e23baedde731a7bb4c88309ef6ea703344429dd75c58233fbaf144470b52

Observation d2eacccb-6752-4d98-ba71-ae2b603cd968 · outbound

This paper cites What matters when building vision-language models?.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL What matters when building vision-language models?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.460988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.460988Z digest=sha256:b1d9282de1c5cb98d54d4ff1cd22f33d8d8b59337a2fdcf45f6fd708822fb016

Observation 26f14349-9cc1-4e46-a011-705cd973b0d0 · outbound

This paper cites STARE at the Structure: Steering ICL Exemplar Selection with Structural Alignment.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL STARE at the Structure: Steering ICL Exemplar Selection with Structural Alignment

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-14T12:25:20.927913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T12:25:20.467307Z digest=sha256:d35cc790bd35fb6ea5a5d80dbf9fc01c0072649f20dfdec63a5aa4ce32061a3c

Observation ea781fbe-d942-4f7f-9b39-d617e33c5649 · outbound

This paper cites Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.472516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.472516Z digest=sha256:ba0c85bf5ee09b1c49849176c9ac3b35005d1e8af06eb9215ea90e904f801b80

Observation 0f75de65-340c-4a1b-98e6-78db09d28c07 · outbound

This paper cites an unresolved cited work.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.477568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.477568Z digest=sha256:d8ae6ff16fe071bf356cafb1e9ef25426ea811c43a7a51bc57376109a2f41702

Observation d09bbd4b-e638-4401-972c-b897c86dccb0 · outbound

This paper cites M$^2$IV: Towards Efficient and Fine-grained Multimodal In-Context Learning via Representation Engineering.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL M$^2$IV: Towards Efficient and Fine-grained Multimodal In-Context Learning via Representation Engineering

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.485623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.485623Z digest=sha256:dffd1db5c139f3e7016a0a79337b3a8050796ad1a06efa2b707bebd6c3e38c2d

Observation ae6b7285-7ab3-4029-9ffd-8267a5cb9212 · outbound

This paper cites Implicit In-context Learning.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Implicit In-context Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.491489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.491489Z digest=sha256:de61efef73e135f6ee5044d2fdaae2b4646aeb4a5140f57117cb74f17c981a53

Observation 7846e799-ecc7-4547-a4f9-c8ec4728d78f · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Improved Baselines with Visual Instruction Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.497916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.497916Z digest=sha256:dad854999d0727981ce2132eb95ca6315adb4c65aeee313bb06d2f3fc484d27c

Observation 6a1adcf5-6235-4762-bb77-b65a4d00852a · outbound

This paper cites In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.507829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.507829Z digest=sha256:3b126d9308aa98622709996ba12110566bcca80578b6f2f1b247c3ddaec7dfab

Observation 39524009-c21d-46c3-87dd-5d5c22842be3 · outbound

This paper cites OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.515072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.515072Z digest=sha256:305b89fbe06300539a98f5b8e9beb45c9f4aed94f2521624442aa53c3ff1400b

Observation e344edda-438f-4548-81ef-49e5d3c7ce66 · outbound

This paper cites In-context Learning and Induction Heads.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL In-context Learning and Induction Heads

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.519581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.519581Z digest=sha256:bb9c02e6a94b2b1bfa99e87f48668bfa97797e9f6faefcbb0ed196c7c86c0473

Observation daf18c9b-1bbb-450d-88d6-8fd8d53d52c1 · outbound

This paper cites Transformers learn in-context by gradient descent.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Transformers learn in-context by gradient descent

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.524771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.524771Z digest=sha256:96b7fc9e40331dec20151786d16dff16cc7ff6b43ed58defb8ab4aa747397063

Observation 28ca7e54-5561-4bde-baa1-cc503637035f · outbound

This paper cites LIVE: Learnable In-Context Vector for Visual Question Answering.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL LIVE: Learnable In-Context Vector for Visual Question Answering

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.530439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.530439Z digest=sha256:a1394e754262ecccd0b5e953a976be3e53e3c53a55c24a1d9a81eec2c9f72e65

Observation 22fa7297-7e1c-41fe-bdfc-d4b0bfee6443 · outbound

This paper cites What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.535764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.535764Z digest=sha256:4a55d0c4e9588d1b9f6f50e2dc2b17571ce4fe4a245a12e5089579732e9cd05e

Observation f7681ec2-f2ad-47ef-8814-7fa80d1444a0 · outbound

This paper cites CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.540281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.540281Z digest=sha256:59ace0629f87437ae15a6da558963820ca70dae12083353d5f589d230367e7c9

Observation 754e5f61-90a1-4e01-9a12-dd61ba3d9f87 · outbound

This paper cites What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.546289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.546289Z digest=sha256:f3610decc46efadf1d485f54748f0a38271b9c26a10475cc9e2283fe686f159e

Observation 7021379d-49e2-442e-912b-daba235122a2 · outbound

This paper cites Generative Multimodal Models are In-Context Learners.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Generative Multimodal Models are In-Context Learners

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.550940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.550940Z digest=sha256:009db5b957478f306d3b205a3f9bf00473b4b1ba7149ae8de380a820c208bb1a

Observation a2c8165b-bd6f-4367-9fb7-a22a2fc4fe5d · outbound

This paper cites Link-Context Learning for Multimodal LLMs.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Link-Context Learning for Multimodal LLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.557658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.557658Z digest=sha256:bd6b67887a62fa275f7e87fd3f29cb1ce67da4454a64e58afc2c96b9890bd09a

Observation 6b5a5a94-d062-4fc1-83f6-48d27875411a · outbound

This paper cites Function Vectors in Large Language Models.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Function Vectors in Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.562359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.562359Z digest=sha256:1f4cb7efc56921a8eed797c80f86985e7ec7a5a882e1e90f0de80b748ff6dab8

Observation 9af55475-d233-430c-8a2a-cd24eec435ac · outbound

This paper cites An Explanation of In-context Learning as Implicit Bayesian Inference.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL An Explanation of In-context Learning as Implicit Bayesian Inference

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.568189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.568189Z digest=sha256:e18a1633a0fad1b86eefc512104142fc92737d98a4ec70ef9bf8c96dc6c6f2d7

Pith citing papers

No inbound Pith citation observations are available.