Pith. sign in

Paper Citation Record · LEDGER

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL

As of 15 August 2026, this Paper Citation Record lists 31 of 31 outbound references and 0 inbound Pith citation observations for arXiv:2608.13385.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.13385 v1

Coverage vector

measured 31 of 31 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-14T12:25:20.568189Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

31 of 31 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fc017263-2b7e-4fe4-93c7-c253e23ee134 · outbound

This paper cites What learning algorithm is in-context learning? Investigations with linear models.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL What learning algorithm is in-context learning? Investigations with linear models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.392114Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.392114Z digest=sha256:b16edaa277741ddc956bd9cc3567ca28bf51f1a5e86e3fe16eb85e07fdbd0bdb

Observation bcc7d6d6-f8cd-4873-928b-0406c42d5a51 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.398327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.398327Z digest=sha256:90ce263789cd48c8d76ffb2ac0ebc6e382758b91925b88201a3be2efcf808a39

Observation c8e11e3a-5d18-462f-bff5-ce47c2745a9b · outbound

This paper cites an unresolved cited work.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-14T12:25:21.110026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T12:25:20.404342Z digest=sha256:d1942570fc8f6f8a1a090195a847d3825a56edcf87d961935e36fba1ef090bd0

Observation 52e97647-f601-4bbc-a8ab-cd8b628a1c2a · outbound

This paper cites Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.408621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.408621Z digest=sha256:ee6b7ec41966ed84beb595a6a5f41b25229aacda5b0888fc71ebe9003271d8a6

Observation aba977d2-d585-44a4-b3fc-ef3e7e08c83d · outbound

This paper cites Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Why Can GPT Learn In-Context? Language Models Implicitly Perform Gradient Descent as Meta-Optimizers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.413368Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.413368Z digest=sha256:2f3dca1262f6ebdf383a06dd85a4b3605566cd3fd3bfa552c398531a463409a0

Observation 2c95a73f-4b83-49f8-8fbd-c469f901babc · outbound

This paper cites Towards Multimodal In-Context Learning for Vision & Language Models.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Towards Multimodal In-Context Learning for Vision & Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.419246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.419246Z digest=sha256:07b6602e3213aa5b10503ca8b62d213ef22c127ee2ee6fb21b4d3931a4b645e1

Observation bbaafca0-6a2d-4307-b0b9-f085588084db · outbound

This paper cites Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Making the V in VQA Matter: Elevating the Role of Image Understanding in Visual Question Answering

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.428086Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.428086Z digest=sha256:d7bf32e6b5fb3d5e583f8315c9a0ace67f06c2eb0c1f4b987f4533c23cb09bc7

Observation 3d011afc-5792-4370-b720-190f88fc958f · outbound

This paper cites In-Context Learning Creates Task Vectors.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL In-Context Learning Creates Task Vectors

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.433188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.433188Z digest=sha256:57534410521772149b28f511b94ae35fac82af1163a62d3b4bb01c880d012775

Observation 6f03b30b-636b-4216-abed-a42feae33873 · outbound

This paper cites Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Multimodal Task Vectors Enable Many-Shot Multimodal In-Context Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.439241Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.439241Z digest=sha256:39d4796e0a8dc362626de6936b8974c9947337f81d7e929e9d4b8b0860ef4836

Observation 537b46df-ed04-42af-9955-4bd386dc975d · outbound

This paper cites Dissecting Multimodal In-Context Learning: Modality Asymmetries and Circuit Dynamics in modern Transformers.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Dissecting Multimodal In-Context Learning: Modality Asymmetries and Circuit Dynamics in modern Transformers

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.444342Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.444342Z digest=sha256:cfb1a7f29af1a2ce765d34d42bb9b07a1801f288f7a7983e2d8e89d80ea02fb1

Observation bab64fb7-04fb-414b-ba00-4fea119a867a · outbound

This paper cites an unresolved cited work.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.450842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.450842Z digest=sha256:a05110fc9a3d26effa2d30b6358e4847d0c43ba2c3ca03ccd9b6f43a29acf75a

Observation d7a316ab-4e05-4b8d-a21a-67efc54b3ae6 · outbound

This paper cites Mimic In-Context Learning for Multimodal Tasks.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Mimic In-Context Learning for Multimodal Tasks

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.456349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.456349Z digest=sha256:fff14ff4b72a68c7013dec5279b29f14a7d209da1904177d27aa231f56abc40b

Observation d2eacccb-6752-4d98-ba71-ae2b603cd968 · outbound

This paper cites What matters when building vision-language models?.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL What matters when building vision-language models?

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.460988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.460988Z digest=sha256:55a21f7b152a002d0d5847b13a1e3853817231b9ffd82aaa95e9f8c97d46cdca

Observation 26f14349-9cc1-4e46-a011-705cd973b0d0 · outbound

This paper cites STARE at the Structure: Steering ICL Exemplar Selection with Structural Alignment.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL STARE at the Structure: Steering ICL Exemplar Selection with Structural Alignment

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-14T12:25:20.927913Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-14T12:25:20.467307Z digest=sha256:8a6f06a7fdf44f656151fd402dcd64064ed23bb1757ef358213ea70924396cb5

Observation ea781fbe-d942-4f7f-9b39-d617e33c5649 · outbound

This paper cites Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Train Once, Reuse Everywhere: Generalizable Implicit In-Context Learning by Routing Attention

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.472516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.472516Z digest=sha256:969258a59987405e08007091f746ecce8df8c4a7a7356fb57952f3ce83b7e92e

Observation 0f75de65-340c-4a1b-98e6-78db09d28c07 · outbound

This paper cites an unresolved cited work.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.477568Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.477568Z digest=sha256:766f577ddc02a2d4b02dfaa74006831639cde0247a365848ef0fd0d4e510e419

Observation d09bbd4b-e638-4401-972c-b897c86dccb0 · outbound

This paper cites M$^2$IV: Towards Efficient and Fine-grained Multimodal In-Context Learning via Representation Engineering.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL M$^2$IV: Towards Efficient and Fine-grained Multimodal In-Context Learning via Representation Engineering

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.485623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.485623Z digest=sha256:4c95e59d39c8e1316656a4b0418e3676cd05f95071d87c985038fc2115d7058f

Observation ae6b7285-7ab3-4029-9ffd-8267a5cb9212 · outbound

This paper cites Implicit In-context Learning.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Implicit In-context Learning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.491489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.491489Z digest=sha256:aad5fd848b09711e62b4aeaa4e4573dece821d329ca4d804d565b00f25c63a53

Observation 7846e799-ecc7-4547-a4f9-c8ec4728d78f · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Improved Baselines with Visual Instruction Tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.497916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.497916Z digest=sha256:3cd2fe3cf2066bde699c10d5078f7d07777432dcf7534bf5205e05b371198614

Observation 6a1adcf5-6235-4762-bb77-b65a4d00852a · outbound

This paper cites In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL In-context Vectors: Making In Context Learning More Effective and Controllable Through Latent Space Steering

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.507829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.507829Z digest=sha256:0ae201f853f8edfde0fa0fc61272c595b1255c08984e220863036dc26eafaeaa

Observation 39524009-c21d-46c3-87dd-5d5c22842be3 · outbound

This paper cites OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL OK-VQA: A Visual Question Answering Benchmark Requiring External Knowledge

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.515072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.515072Z digest=sha256:bb3591c2cf905c3a0aa24cf03b9bbad0e192854cae0e6c10cf34c600ab8fe996

Observation e344edda-438f-4548-81ef-49e5d3c7ce66 · outbound

This paper cites In-context Learning and Induction Heads.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL In-context Learning and Induction Heads

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.519581Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.519581Z digest=sha256:22a59f91c0b65aca282ff9ac95b51988e021d87b260eaad84f898f1b6ca61899

Observation daf18c9b-1bbb-450d-88d6-8fd8d53d52c1 · outbound

This paper cites Transformers learn in-context by gradient descent.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Transformers learn in-context by gradient descent

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.524771Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.524771Z digest=sha256:e25d8c19ce8928b26c412758167784d993adcaf5832304b8c44e79f8c57e42f5

Observation 28ca7e54-5561-4bde-baa1-cc503637035f · outbound

This paper cites LIVE: Learnable In-Context Vector for Visual Question Answering.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL LIVE: Learnable In-Context Vector for Visual Question Answering

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.530439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.530439Z digest=sha256:e6a51151702dfa4b0a6f02067055fc7573e10543fbf14a174862db80f1448b21

Observation 22fa7297-7e1c-41fe-bdfc-d4b0bfee6443 · outbound

This paper cites What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.535764Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.535764Z digest=sha256:0569f4fa369c1ad33ad9ce94048a341b2a6ab71a2f7e8736d8ec1d5e2fe4c157

Observation f7681ec2-f2ad-47ef-8814-7fa80d1444a0 · outbound

This paper cites CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL CVQA: Culturally-diverse Multilingual Visual Question Answering Benchmark

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.540281Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.540281Z digest=sha256:83a2ab5728c18d9d0d922cc571dcf92c247b4eb93f6a5239690215d63a618405

Observation 754e5f61-90a1-4e01-9a12-dd61ba3d9f87 · outbound

This paper cites What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL What needs to go right for an induction head? A mechanistic study of in-context learning circuits and their formation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.546289Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.546289Z digest=sha256:a98a4b59ff9274d7940946f1677c263ac8639aef0b6324ba44d4e47c52ad2cec

Observation 7021379d-49e2-442e-912b-daba235122a2 · outbound

This paper cites Generative Multimodal Models are In-Context Learners.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Generative Multimodal Models are In-Context Learners

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.550940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.550940Z digest=sha256:c1405fe822abbf5eb794f3b40e0d0d0851fdf28de6301e8f8886254b2339f4ff

Observation a2c8165b-bd6f-4367-9fb7-a22a2fc4fe5d · outbound

This paper cites Link-Context Learning for Multimodal LLMs.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Link-Context Learning for Multimodal LLMs

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.557658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.557658Z digest=sha256:22339a21a413a5abcfaf14bf36822a1c3c40a234da5047c1987c1d0c40c949b7

Observation 6b5a5a94-d062-4fc1-83f6-48d27875411a · outbound

This paper cites Function Vectors in Large Language Models.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL Function Vectors in Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.562359Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.562359Z digest=sha256:3f399df86d2443180916b1e3bd3f99604a61afaaf362fb1271c4445913df8a70

Observation 9af55475-d233-430c-8a2a-cd24eec435ac · outbound

This paper cites An Explanation of In-context Learning as Implicit Bayesian Inference.

When Is a Task Vector Enough? An Empirical Theory of Implicit Multimodal ICL An Explanation of In-context Learning as Implicit Bayesian Inference

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-14T12:25:20.568189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-14T12:25:20.568189Z digest=sha256:91a5cdccbc1f7410155510a83119f54fc59b25df39fdf730bd65cfcc2b72e7fe

Pith citing papers

No inbound Pith citation observations are available.