Pith. sign in

Paper Citation Record · LEDGER

CF-VLM:CounterFactual Vision-Language Fine-tuning

As of 20 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 4 inbound Pith citation observations for arXiv:2506.17267.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17267 v1

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:01:07.332466Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T05:51:29.406876Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:09:44.501635Z

Reference resolution

81 of 81 outbound references displayed

  • verified exact7
  • verified fuzzy33
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation db04e06f-f11b-45e7-a857-2b6f14e0b7e2 · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation.

CF-VLM:CounterFactual Vision-Language Fine-tuning BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:06.943940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:06.943940Z digest=sha256:dc5e1433b976d64d90d1fa42ec54b9629814156a2873af0bc0b7287e351a0b31

Observation dfbd2932-fb48-49b2-b13d-92c274c9d958 · outbound

This paper cites A Survey of Vision-Language Pre-Trained Models.

CF-VLM:CounterFactual Vision-Language Fine-tuning A Survey of Vision-Language Pre-Trained Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:06.949603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:06.949603Z digest=sha256:df5840bd4473ab5ad3d6f4dc37cd3febbe10a72fbcd8f4845bc51d1ccf840010

Observation 37109915-1310-47d6-9543-ff27ed9b2992 · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Flamingo: a visual language model for few-shot learning,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:06.954808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:06.954808Z digest=sha256:2a6063f693ed70fa0bb820c183d247f9632a6cb47ecde79e5b2d92f2c6181366

Observation fcc8c1a0-c985-4ed2-b869-d631b3815a97 · outbound

This paper cites Causal Inference with Large Language Model: A Survey.

CF-VLM:CounterFactual Vision-Language Fine-tuning Causal Inference with Large Language Model: A Survey

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:06.959963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:06.959963Z digest=sha256:99b01d3bd4ad0dfe6ad9539690674a8a64e7368408e3e1627d8a77ee94c745b3

Observation 60fea556-4575-45ff-9b62-2d850b2a616b · outbound

This paper cites CELLO: Causal Evaluation of Large Vision-Language Models.

CF-VLM:CounterFactual Vision-Language Fine-tuning CELLO: Causal Evaluation of Large Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:06.966031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:06.966031Z digest=sha256:0ab5d6400a0a99b0a7bab4ea0d2926ab7170639f475be51446e1ad8b0d5e9f12

Observation 99555b6c-afd1-4d6d-8953-d6324ebaf0e1 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

CF-VLM:CounterFactual Vision-Language Fine-tuning Learning Transferable Visual Models From Natural Language Supervision

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:06.971026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:06.971026Z digest=sha256:dc4c90ea5c1a1d39f812a3b4a97a757f5460306dba4f16ad5d67a6c68c22a8d2

Observation 8ff0d1b4-07be-48bd-ae98-24a0b4f88ee3 · outbound

This paper cites Measuring progress in fine-grained vision-and-language understanding,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Measuring progress in fine-grained vision-and-language understanding,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:09.057731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:06.977108Z digest=sha256:d5db1fa08d12b141d44509944e609ba3acad275d697910ef8d1b3188b5dd36c1

Observation 8300da25-7d91-4df9-abed-eaece9b8617e · outbound

This paper cites Synthesize, diagnose, and optimize: Towards fine-grained vision-language understanding,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Synthesize, diagnose, and optimize: Towards fine-grained vision-language understanding,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.985378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:06.981847Z digest=sha256:c734a42f3932c66aaab6e18918d85a4b60a6af74ea3310145cbc79a1ba09113e

Observation e3947761-44ac-4821-adb9-1a20a4b3d0b9 · outbound

This paper cites Finer: Investigating and enhancing fine-grained visual concept recognition in large vision language models,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Finer: Investigating and enhancing fine-grained visual concept recognition in large vision language models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.966532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:06.986025Z digest=sha256:c4d1faf9714f93ff50e27ccd86d31e59f40ae4a19b91182bdfad3b323bbcd9c0

Observation 7d2999e9-fa44-4ff1-b40a-137410a2ded7 · outbound

This paper cites Benchmarking zero-shot recognition with vision-language models: Challenges on granularity and specificity,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Benchmarking zero-shot recognition with vision-language models: Challenges on granularity and specificity,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.947634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:06.990755Z digest=sha256:c418d5f0a8c4531545c9458ae22f758d5291296c1c747c933d1aacab1b433e7a

Observation ab451ce1-26b1-4cb7-81d3-32137ef45aaa · outbound

This paper cites Vilta: Enhancing vision-language pre-training through textual augmentation,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Vilta: Enhancing vision-language pre-training through textual augmentation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.931918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:06.995224Z digest=sha256:f4d39c8d4b385d9765269d43ffebf062385e45e6851d4618c0ed28297e525a73

Observation b637aff9-66da-4a03-bcb8-bc90540662f9 · outbound

This paper cites Towards vision-language mechanistic interpretability: A causal tracing tool for blip,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Towards vision-language mechanistic interpretability: A causal tracing tool for blip,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.914022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:06.999305Z digest=sha256:385e0a35bd76fb3f4392d6f0de75728aa326e1d7462527397c65b9f204d25dd6

Observation a1d558dd-dfec-4602-8c5c-0d506ba8ad8f · outbound

This paper cites What matters when building vision-language models?.

CF-VLM:CounterFactual Vision-Language Fine-tuning What matters when building vision-language models?

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.895947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.003354Z digest=sha256:ae734add6c0847470edc620eb27535e15c7b5bc26a69e38b5a206bc8fef6d9f7

Observation b9f42f80-8a62-4c16-b903-fd65508ef49a · outbound

This paper cites Fine-grained alignment for cross-modal recipe retrieval,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Fine-grained alignment for cross-modal recipe retrieval,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.878291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.013604Z digest=sha256:31bbcf03f6bb30c5e5f57fb2987c483675f3d6974f745d9589da2a4c5d23770b

Observation 0c8a2ede-f477-4e9f-bc3e-ac43428afc04 · outbound

This paper cites Localized triplet loss for fine-grained fashion image retrieval,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Localized triplet loss for fine-grained fashion image retrieval,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.858955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.018556Z digest=sha256:e798f844ddf1176e11a22450a29a9b66f30f41dc7beb2adb6fd86f55132f2c1b

Observation 110a4311-512f-4704-9fe0-b9aaca3441f3 · outbound

This paper cites TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives.

CF-VLM:CounterFactual Vision-Language Fine-tuning TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:01:08.073739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.023061Z digest=sha256:7a59f003cdd1b93f2ee84ea5680452de297023aece7d3709c72d9b1d9e6804ad

Observation 3c99b13f-56ef-4af6-9ab3-ca42fbd4ccb9 · outbound

This paper cites Facenet: A unified embedding for face recognition and clustering,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Facenet: A unified embedding for face recognition and clustering,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.028168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.028168Z digest=sha256:cf0d4cf18f9a7aa0fe6bc10cbc1325bf09b72477fab4b33fb1e8db7da8e9a7e0

Observation 8bbcaeb3-aba4-445b-8ba4-599906698382 · outbound

This paper cites CPL: Counterfactual prompt learning for vision and language models,.

CF-VLM:CounterFactual Vision-Language Fine-tuning CPL: Counterfactual prompt learning for vision and language models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.825807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.037577Z digest=sha256:cc6b118e296b2ed4d5cfdda0c9bf99f76ff92a2530a10c6337ead7a5ce9f321e

Observation f2cdb927-37fb-41da-80bd-bbd97f3162c2 · outbound

This paper cites Attention Is All You Need.

CF-VLM:CounterFactual Vision-Language Fine-tuning Attention Is All You Need

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.046597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.046597Z digest=sha256:e7c4b6505f5b9729c89f10b15229494b073d51f98b805899a25500540c005658

Observation b666d97b-07a6-4e8f-afd5-494e7f79f3db · outbound

This paper cites A survey on evaluation of large language models,.

CF-VLM:CounterFactual Vision-Language Fine-tuning A survey on evaluation of large language models,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.051054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.051054Z digest=sha256:71c7ac0f119c0e9e7e163188a7ee5d5d7bc3980fd8a2ed77bc6947d118b10250

Observation ba152358-f5c9-440e-9f5e-efe7a008b601 · outbound

This paper cites A survey of visual transformers,.

CF-VLM:CounterFactual Vision-Language Fine-tuning A survey of visual transformers,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.808874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.055484Z digest=sha256:01319aedd646d5fee2af7b7d43e1b17039cbb0596e7e008c15038bb8d20d29cd

Observation 846e6f77-4295-47f2-ba90-0ee581c751d9 · outbound

This paper cites A simple framework for contrastive learning of visual representations,.

CF-VLM:CounterFactual Vision-Language Fine-tuning A simple framework for contrastive learning of visual representations,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.059548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.059548Z digest=sha256:598c3bc109749c922c9c71d5098e15022f2bca16226c30d950875ae39a35ff2e

Observation e4f02ba1-a538-42dd-8408-84b1e0268399 · outbound

This paper cites Understanding contrastive representation learning through alignment and uniformity on the hypersphere,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Understanding contrastive representation learning through alignment and uniformity on the hypersphere,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.781994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.064955Z digest=sha256:fdde4d25659e75b5404f674d29d6bc0fe90c5f0a1b551dd47cfe488d961e8fe3

Observation 02c019f0-7f4e-4b7e-b6b5-d13924477716 · outbound

This paper cites Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision.

CF-VLM:CounterFactual Vision-Language Fine-tuning Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.069602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.069602Z digest=sha256:d29ff9266518019102335176e9021d37140f27f39bb9eacac95016b647815df7

Observation 3b9e9e6a-efdb-4d44-a706-1028d67ab742 · outbound

This paper cites Cogs: A compositional generalization challenge based on semantic interpretation,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Cogs: A compositional generalization challenge based on semantic interpretation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.764968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.074611Z digest=sha256:443203db15ab333ce59572dd6cebe9012b7c544f6f498bb374367832f6f0fae3

Observation 43a53351-d77c-4251-b4fd-cea102c36b13 · outbound

This paper cites Learning what makes a difference from counterfactual examples and gradient supervision,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Learning what makes a difference from counterfactual examples and gradient supervision,

Reference 28

Resolution
verified exact
doi, observed 2026-08-07T05:01:07.374921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.085658Z digest=sha256:e45443fd948001e1f6a3b324b213d1c0f2c16f62ab4c31df54e7b3b3a1320b2d

Observation 08b38393-f8f0-44a3-b187-ce2bd65a3274 · outbound

This paper cites Teaching clip to count to ten,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Teaching clip to count to ten,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.748372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.090433Z digest=sha256:7bde12d58757ad3a64d7f11f8ef48a701de908f1cfa4b9c114ed40d642e0b704

Observation e06c7165-c2c9-4efa-b367-550c11e41768 · outbound

This paper cites DISCO: Distilling Counterfactuals with Large Language Models.

CF-VLM:CounterFactual Vision-Language Fine-tuning DISCO: Distilling Counterfactuals with Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.094854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.094854Z digest=sha256:7fbbcc52ecfae3b6dadde1631df821881294da1adc3602d0906e0fbce161a104

Observation 2bad5841-e5b9-4482-a9f0-64aaf70d9a71 · outbound

This paper cites Counterfactually measuring and eliminating social bias in vision-language pre-training models,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Counterfactually measuring and eliminating social bias in vision-language pre-training models,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.100956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.100956Z digest=sha256:6fc598af78a2b0bdea2f96c6ab9fffbb81c2c6c3abd905c485aa856ae6c5624d

Observation e09ab765-4ce8-4e63-9431-0f3250b93549 · outbound

This paper cites Counterfactual attention learning for fine-grained visual categorization and re-identification,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Counterfactual attention learning for fine-grained visual categorization and re-identification,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.842305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.105617Z digest=sha256:fa90d87e5fa5e8cfe178aad0f4c5a2bc42c35ce4b354b5e354c6df43e7ca474f

Observation af420799-f337-4ec0-b5f4-19b03960bb1c · outbound

This paper cites Counterfactual samples synthesizing and training for robust visual question answering,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Counterfactual samples synthesizing and training for robust visual question answering,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.730534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.110453Z digest=sha256:5ed4c187e3b7908bb1764c169302d5879331ead09d9e2bc16f06bdb4cce535ff

Observation 5733a09a-3318-489a-98d1-376fb7498080 · outbound

This paper cites Qwen2.5 Technical Report.

CF-VLM:CounterFactual Vision-Language Fine-tuning Qwen2.5 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.114672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.114672Z digest=sha256:93c2d5c28f111f40964a8053c314ce8a811cef17a9d8084ec75894e38da9c36a

Observation f4593c00-6d8a-4f6d-aa1d-a6a6c4f83ef0 · outbound

This paper cites Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.711237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.119470Z digest=sha256:105177bb7a587ab6de4e98c4f2af5b6f96035468892fe8679c19a8f0e927b6bc

Observation fda6020a-f7a7-4d80-8212-4c0c609522ea · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.692774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.123742Z digest=sha256:3d683423942074f6da3f3831a469abd9d577681981e6a121d9856aeda756ba49

Observation 060bdcea-5d8c-4cec-9572-47b5edeed4be · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

CF-VLM:CounterFactual Vision-Language Fine-tuning Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.128034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.128034Z digest=sha256:a0f073138784848fa3daa6902c92dc931cef27bf933c7defeb3f79db6ce6bedf

Observation efa34ba6-5906-48d1-9a18-de51b7693da2 · outbound

This paper cites ConMe: Rethinking Evaluation of Compositional Reasoning for Modern VLMs.

CF-VLM:CounterFactual Vision-Language Fine-tuning ConMe: Rethinking Evaluation of Compositional Reasoning for Modern VLMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.132424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.132424Z digest=sha256:9e95b4a161dbb7d05df9a9ef1683cdd6c2c6aacd6bd3bebb429afcbaa8984458

Observation 866c9125-f995-4dfe-bd87-feb0efc052a2 · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it?.

CF-VLM:CounterFactual Vision-Language Fine-tuning When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.674987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.136591Z digest=sha256:baa95016bd9533de5773c3960930cd14c1d485c962e4c9772c4c6637ab7ccddd

Observation e0bd20ed-6479-47b0-bb65-0026b8d5f9d2 · outbound

This paper cites VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations.

CF-VLM:CounterFactual Vision-Language Fine-tuning VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.140504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.140504Z digest=sha256:b44aa733b927dc7c6ca2bc8f92f847460a1ffb6d286ccc2ca0e931781f5db024

Observation f12b6b44-3b06-4d04-a7a4-ca0e8df4fa74 · outbound

This paper cites ImageNet Large Scale Visual Recognition Challenge,.

CF-VLM:CounterFactual Vision-Language Fine-tuning ImageNet Large Scale Visual Recognition Challenge,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.657970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.145231Z digest=sha256:0ebcc7c5a30687c8277416bbd6a807d113b0f2f6ec6ae65824191b224be7bb04

Observation 799bc5f2-8288-4427-b732-aea50db657ef · outbound

This paper cites From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,.

CF-VLM:CounterFactual Vision-Language Fine-tuning From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.640639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.149485Z digest=sha256:133087e96f5d7ff52266a8508b7e0153cb3f67c7345dabbd19dbefc857ec62ca

Observation ed03be52-36d3-4416-990c-4cafb667f01d · outbound

This paper cites Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional Understanding.

CF-VLM:CounterFactual Vision-Language Fine-tuning Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional Understanding

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:01:07.741382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.153634Z digest=sha256:6b5f3f664b57e2fdc0d1c659159f245e9044a3108ae93ffbd99e2334cea6dade

Observation c6d5e468-171b-47c6-a514-c88c7e26d27c · outbound

This paper cites Structure-CLIP: Towards Scene Graph Knowledge to Enhance Multi-modal Structured Representations.

CF-VLM:CounterFactual Vision-Language Fine-tuning Structure-CLIP: Towards Scene Graph Knowledge to Enhance Multi-modal Structured Representations

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:01:07.697666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.163971Z digest=sha256:5636d34c9b0e132684bfaa7103f27532906d69ae02c7177d3cb9653ca3af7045

Observation 036de79c-a746-4d6b-8d94-91e7d50d6c8f · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

CF-VLM:CounterFactual Vision-Language Fine-tuning Improved Baselines with Visual Instruction Tuning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.168343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.168343Z digest=sha256:64eb9a7b277af895b8c2b84294e9f6c2bd141cbf7fbcf53fa2a715ddf7a05bd1

Observation 556d253d-ee8b-4367-bd2b-bd573994ec36 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

CF-VLM:CounterFactual Vision-Language Fine-tuning InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.173286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.173286Z digest=sha256:7837d9136819431305d8de0ce1a6561e072281235561da9a44269926f63e3428

Observation a3d19d3a-9461-4759-90de-978b5b273bc3 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

CF-VLM:CounterFactual Vision-Language Fine-tuning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.178097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.178097Z digest=sha256:5df425b941f2e61e8dded95b5bf2870d0a107969d1b54ef34139f82868e7322f

Observation 3838e51a-9024-44e9-8b6c-4840d6b2b301 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

CF-VLM:CounterFactual Vision-Language Fine-tuning SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.182593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.182593Z digest=sha256:9e56b1d180ffdc7af56714cacdb36b5f383599e1c989fe252d9d78586f6eacb9

Observation 4d6eedb6-ee1e-4ec7-a636-f31252676b0a · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

CF-VLM:CounterFactual Vision-Language Fine-tuning Evaluating Object Hallucination in Large Vision-Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.187178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.187178Z digest=sha256:d2a16eec10a6766a0aff0b47ac89f357dff2f51d4b032e2e94cb0a0e4f5672c1

Observation 8903ba8e-e6da-43b9-b714-2c24b5ee9f14 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

CF-VLM:CounterFactual Vision-Language Fine-tuning MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.191460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.191460Z digest=sha256:6a7771fabdbc28bab45d62b60a889287602ffb51120e8e4fc3f22fd0856e8aef

Observation ce3d1c0b-f348-4bdc-a973-23c90907feb9 · outbound

This paper cites Vision-and-Language Pretrained Models: A Survey.

CF-VLM:CounterFactual Vision-Language Fine-tuning Vision-and-Language Pretrained Models: A Survey

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.195795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.195795Z digest=sha256:fdfde512f2788d8c98e9e3b6a39f6ffedd60e2df369b485c79e8c07fb0526345

Observation 6f46b262-c841-40b7-afca-bda312bc53a6 · outbound

This paper cites Exploring the frontier of vision-language models: A survey of current methodologies and future directions,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Exploring the frontier of vision-language models: A survey of current methodologies and future directions,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.200219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.200219Z digest=sha256:d18b6a50407bd10bb5731acf46484afb60cd9ca0a21416c1865540a663b1272e

Observation 04362e3a-ca1b-4aa9-b3c2-5162f4df9c23 · outbound

This paper cites Vision-language models for vision tasks: A survey,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Vision-language models for vision tasks: A survey,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.204431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.204431Z digest=sha256:be65e27d6aa29dcc8e469fcfeebc230cd025f1c4cffa6553aca45b3feedb5615

Observation 1064166d-50a6-44ac-a19e-7dead8da82db · outbound

This paper cites Counterfactual vision and language learning,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Counterfactual vision and language learning,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.613419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.208880Z digest=sha256:ec757b13abfde8f03f391dfec6e19263a41e0894186864cc2fc5c63fe919cdfb

Observation 5256965e-acca-49f2-83fb-b036e5e749f4 · outbound

This paper cites Counterfactual reasoning for multi-label image classification via patching-based training,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Counterfactual reasoning for multi-label image classification via patching-based training,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.597922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.213460Z digest=sha256:98011ed396d08e2dd784e6a1b920ae5bc032fa7af2b758fda6b4a3600f48d12c

Observation 75ad03d2-c6ca-4845-a852-39b0a661e028 · outbound

This paper cites Causal Graphical Models for Vision-Language Compositional Understanding.

CF-VLM:CounterFactual Vision-Language Fine-tuning Causal Graphical Models for Vision-Language Compositional Understanding

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:01:07.719377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.218215Z digest=sha256:4d7eb7679e67250347b9dd282a0dc1ff5a604c429087d92ef3c3fcc9e3fb6bd2

Observation f4e29da0-3367-43e2-b371-50cbf99390a1 · outbound

This paper cites CPL: Counterfactual Prompt Learning for Vision and Language Models.

CF-VLM:CounterFactual Vision-Language Fine-tuning CPL: Counterfactual Prompt Learning for Vision and Language Models

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:01:07.433090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.222289Z digest=sha256:5755b01796ce216060eefd867982812a514b27c3359b5ba506e81c2ae6c1ca75

Observation 82b5f82b-87b4-4f61-8a8e-38e472eac53b · outbound

This paper cites Counterfactual Visual Explanations.

CF-VLM:CounterFactual Vision-Language Fine-tuning Counterfactual Visual Explanations

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.227257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.227257Z digest=sha256:fdef23fd4521ee8d64565abf54b499e5751c32721351dde33fd06b94be95f715

Observation 8e4108df-7b85-443d-af3d-3ff07f794393 · outbound

This paper cites Counterfactual Attention Learning for Fine-Grained Visual Categorization and Re-identification.

CF-VLM:CounterFactual Vision-Language Fine-tuning Counterfactual Attention Learning for Fine-Grained Visual Categorization and Re-identification

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:01:07.410653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.231457Z digest=sha256:33b094a3a693ff29530af7d82c3202be0193947fe674ea81faf4996a963d9bb3

Observation af0ade69-265d-4dfe-87f4-5bac436fca09 · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.580872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.236954Z digest=sha256:6eecb024f8bf1dfddbaf67d7bbf9bcd2ab777ea5caf87ef159a41db343b5ff3e

Observation 38178384-683d-4c49-b0d8-2d9434e28870 · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.565173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.241166Z digest=sha256:8f606d11094ac07423bb9558742580e7e5ab8424cf2b8855fac6c050c762bb38

Observation b4b2c7b5-42f4-4330-97b6-da6084a79ff5 · outbound

This paper cites Rules: • Modifyonly one thing(either one attribute or one causal link).

CF-VLM:CounterFactual Vision-Language Fine-tuning Rules: • Modifyonly one thing(either one attribute or one causal link)

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.547774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.245312Z digest=sha256:053fe42557f4d90fef92dcfd586de81a1897ebc4d49c37174131d16759ee8231

Observation a125a548-442b-48f1-8e97-5a7935069a43 · outbound

This paper cites Example Input: A young woman holding a racket hit the ball, and the ball flew outward.

CF-VLM:CounterFactual Vision-Language Fine-tuning Example Input: A young woman holding a racket hit the ball, and the ball flew outward

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.528325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.250450Z digest=sha256:f39d494ba5bb8d82713a7c6418ed9c64f531f02257e2244a99c86c5a7bce7501

Observation 98563f15-079f-4b01-877a-75ae0aa6db83 · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.512290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.255371Z digest=sha256:a114625017437aa5854ee50b0595821caab99ebaaa081e857d6e2afd74cf5b48

Observation 87f61e64-b74d-41b0-90d6-958147e1b9ef · outbound

This paper cites Output (causal):.

CF-VLM:CounterFactual Vision-Language Fine-tuning Output (causal):

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.495564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.260715Z digest=sha256:fb26cd99b54d053ba01194bb37ce66299bf7a2efd321aa34591b2817bd363c49

Observation b3fa6ee5-d2a5-4657-a4b6-db4745d88fe1 · outbound

This paper cites dirt” with “paved roads.

CF-VLM:CounterFactual Vision-Language Fine-tuning dirt” with “paved roads

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.474047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.265588Z digest=sha256:7bdd18cfb48ca2757a5ecd5bc7d4edba0d8776d2bbcce93ae999647e50c119c4

Observation 37f1a45c-cfb8-476f-877b-b90457564489 · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.446664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.269999Z digest=sha256:22b2deaf8c38d63ad59356fbe5588e87cdd1940ef123e460c81cd6b67a9aab2d

Observation b49d51f6-6c25-4ba7-827c-78703bf04ab2 · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.428518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.273979Z digest=sha256:55df3720ee5d3c076606f35c792b1b51bcb4c76c5b221faa84f8c23f5cdb8fd4

Observation efcd3bf2-24ca-48b3-95c4-1d7b15b217ad · outbound

This paper cites causal decision points.

CF-VLM:CounterFactual Vision-Language Fine-tuning causal decision points

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.412119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.277833Z digest=sha256:197b7cde67fc93467df704e0613f582ac11c32acb015a986b3516f8d80be123d

Observation 1a1cde11-e036-45e2-b861-38daf8f12ede · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.393260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.282997Z digest=sha256:61d24fb52b9afad928d98d6e4871824fbc9998356487f395a959008764352963

Observation 5504e3cd-2e3b-4cfa-854b-3e4a660c462e · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.375706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.287322Z digest=sha256:d0adcaadf47402608f8cb0a427200551d5b80865ff09254548b2b4f0b7adb67e

Observation 72662f1d-9f16-4588-bd8f-3ce329cae31a · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.357028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.292067Z digest=sha256:10d8e2fbd8459fc57ecbfa775ea0a9268f58f810a4f59e0135796d0ffa099b12

Observation b4b55706-9edf-4850-8635-25e55d99d1eb · outbound

This paper cites Two people on motorcycles riding them.

CF-VLM:CounterFactual Vision-Language Fine-tuning Two people on motorcycles riding them

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.338390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.297273Z digest=sha256:ef53b1c181ac5e2a32d9fa284221b89458b937c1ca720721064f9ce3470dde10

Observation ef86dff8-a5a1-4677-9571-c66da1bcdad5 · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.316820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.302045Z digest=sha256:73653fcef024dbaccec18a11b0804c95fc7d822ffe9b80aef6564a90df3435c0

Observation c2267686-b567-446f-825c-7242bcfec410 · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.298808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.306355Z digest=sha256:6c2de13cf4264d3cd95cfcafc96cb22499a8dbbb8963192507a209ee5e375dc6

Observation dc45ff25-b0fd-42b3-9c54-1e74ba3d9c7a · outbound

This paper cites causality.

CF-VLM:CounterFactual Vision-Language Fine-tuning causality

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.281488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.310308Z digest=sha256:cb534a6787038569be644c0d98157dcb41f73232e27c7b61de172ab29ea0fb01

Observation 5e3b4834-cbbe-463d-ad1c-215fa2526338 · outbound

This paper cites parallel realities.

CF-VLM:CounterFactual Vision-Language Fine-tuning parallel realities

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.260724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.314460Z digest=sha256:d53a6b09bf4c203c0beba47d024b8b3c287e3386ef4f6d85c90b0b39252b325c

Observation 443bbec6-8f01-47e4-a298-1b68d0e77f84 · outbound

This paper cites These are employed in Lcsd to help the model learn semantic scene boundaries.

CF-VLM:CounterFactual Vision-Language Fine-tuning These are employed in Lcsd to help the model learn semantic scene boundaries

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.243881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.319372Z digest=sha256:cdc656522eaaf7a1f27c04adbf63b59be38d77d23d3ff8835bd1d33a70c6568c

Observation 7aa85fba-ece3-4611-85b4-63191fb81046 · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.224933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.323874Z digest=sha256:d9c76bd09e0d2682cb870ac822f57dc91b68cd80e4782c020c99f6f80e2d7f4f

Observation 527bc56b-0dc5-4d77-bf56-110aaa7a9aa8 · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.210112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.327832Z digest=sha256:daef5b6cf6939bad2b23974add9f4b8dbb23f8f01dd80f90864e7c7e0edcd2cd

Observation 39db48cf-7edf-4131-b359-6611f6b84ee1 · outbound

This paper cites kicking a ball.

CF-VLM:CounterFactual Vision-Language Fine-tuning kicking a ball

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.194421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T05:01:07.332466Z digest=sha256:f20c05b8f21e574c6137f6e3e46abceacbbc41370717c7abb61745facb6af341

Observation d6f96fae-5544-4683-a0ec-c5087db8e7fc · outbound

This paper cites COGS: A Compositional Generalization Challenge Based on Semantic Interpretation.

CF-VLM:CounterFactual Vision-Language Fine-tuning COGS: A Compositional Generalization Challenge Based on Semantic Interpretation

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.080997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.080997Z digest=sha256:3d07795a7d0791e9b9b624a7b7f11904a8c87a3b99c39547ba30f3d2dcc4718f

Observation f17ba9ae-74f7-42d7-bef8-edc49c67d258 · outbound

This paper cites What matters when building vision-language models?.

CF-VLM:CounterFactual Vision-Language Fine-tuning What matters when building vision-language models?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.007888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.007888Z digest=sha256:dfef6245a690f6109acc68b6aa86cbc9f145a201fcb3424d0168c8658f2ff2fe

Pith citing papers

Observation ed3a27c0-7fbc-410e-b5e1-f1305cb5945a · inbound

OSC: Cognitive Orchestration through Dynamic Knowledge Alignment in Multi-Agent LLM Collaboration cites this paper.

OSC: Cognitive Orchestration through Dynamic Knowledge Alignment in Multi-Agent LLM Collaboration CF-VLM:CounterFactual Vision-Language Fine-tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T05:51:29.406876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:51:29.406876Z digest=sha256:c3388e173eea16a92a50e1d3f998bb2868411415219483b59a35345863c092da

Observation f99d24fa-3e34-4deb-bb1e-d630323c6d50 · inbound

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs cites this paper.

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs CF-VLM:CounterFactual Vision-Language Fine-tuning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T22:14:31.005765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:14:31.005765Z digest=sha256:6cce514f02148b2b4bc58ef57fd43b04e6d4418ddba8abd22bb8b42801af3ad5

Observation e43294dc-c047-41b0-97e7-665272a05f21 · inbound

CFPO: Counterfactual Policy Optimization for Multimodal Reasoning cites this paper.

CFPO: Counterfactual Policy Optimization for Multimodal Reasoning CF-VLM:CounterFactual Vision-Language Fine-tuning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:09:44.503204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-06-26T09:09:09.716984Z digest=sha256:b6b309c03c98e0707ce646dc5c9e592432cd0bc12d034b584365b639470a9dd3

Observation 03766e84-061f-41f4-9881-13016b986e34 · inbound

SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning cites this paper.

SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning CF-VLM:CounterFactual Vision-Language Fine-tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T03:20:41.148082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:20:41.148082Z digest=sha256:a934837c317c845e2a09457ebe1dd70621dedaa4a2c0b868b08e4bab01035729