Pith. sign in

Paper Citation Record · LEDGER

CF-VLM:CounterFactual Vision-Language Fine-tuning

As of 9 August 2026, this Paper Citation Record lists 81 of 81 outbound references and 4 inbound Pith citation observations for arXiv:2506.17267.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.17267 v1

Coverage vector

measured 81 of 81 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:01:07.332466Z

measured 85 of 85 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T05:51:29.406876Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T10:09:44.501635Z

Reference resolution

81 of 81 outbound references displayed

  • verified exact7
  • verified fuzzy33
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation db04e06f-f11b-45e7-a857-2b6f14e0b7e2 · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation.

CF-VLM:CounterFactual Vision-Language Fine-tuning BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:06.943940Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:06.943940Z digest=sha256:26146a2e36bf58faed33e455f58294f5472cd37cddae9fc514da4571b44e14ab

Observation dfbd2932-fb48-49b2-b13d-92c274c9d958 · outbound

This paper cites A Survey of Vision-Language Pre-Trained Models.

CF-VLM:CounterFactual Vision-Language Fine-tuning A Survey of Vision-Language Pre-Trained Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:06.949603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:06.949603Z digest=sha256:6c9d6fd76170cc830ca6bbf7741fe1fe5f96aeacc040cb459005605a56acb3f5

Observation 37109915-1310-47d6-9543-ff27ed9b2992 · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Flamingo: a visual language model for few-shot learning,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:06.954808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:06.954808Z digest=sha256:c65450ef5955d3084dd5b6c4b3816dab73efe8b1f67da675e407eeb8697e7899

Observation fcc8c1a0-c985-4ed2-b869-d631b3815a97 · outbound

This paper cites Causal Inference with Large Language Model: A Survey.

CF-VLM:CounterFactual Vision-Language Fine-tuning Causal Inference with Large Language Model: A Survey

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:06.959963Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:06.959963Z digest=sha256:504fc92e3bd58abf4260d622c07c5e3a0eacd5f006d1e328001513aae3d773ad

Observation 60fea556-4575-45ff-9b62-2d850b2a616b · outbound

This paper cites CELLO: Causal Evaluation of Large Vision-Language Models.

CF-VLM:CounterFactual Vision-Language Fine-tuning CELLO: Causal Evaluation of Large Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:06.966031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:06.966031Z digest=sha256:86151cc37f5ea378b1fd5c7e0ea63db736f4bd9917b33ea1a763956de503f4b3

Observation 99555b6c-afd1-4d6d-8953-d6324ebaf0e1 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

CF-VLM:CounterFactual Vision-Language Fine-tuning Learning Transferable Visual Models From Natural Language Supervision

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:06.971026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:06.971026Z digest=sha256:c329ae786b122c2bd0067ea056d61ae30ceb317510fdb5b8a3d54c9d4f1ef25c

Observation 8ff0d1b4-07be-48bd-ae98-24a0b4f88ee3 · outbound

This paper cites Measuring progress in fine-grained vision-and-language understanding,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Measuring progress in fine-grained vision-and-language understanding,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:09.057731Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:06.977108Z digest=sha256:52e0f20bcb3a23d5797060c02b49964258730059bc86198cb12e99deac69ef55

Observation 8300da25-7d91-4df9-abed-eaece9b8617e · outbound

This paper cites Synthesize, diagnose, and optimize: Towards fine-grained vision-language understanding,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Synthesize, diagnose, and optimize: Towards fine-grained vision-language understanding,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.985378Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:06.981847Z digest=sha256:24a354538d8cc3eb92d68aff057df61f570062738238386b0dc2e7ae13c6c5d5

Observation e3947761-44ac-4821-adb9-1a20a4b3d0b9 · outbound

This paper cites Finer: Investigating and enhancing fine-grained visual concept recognition in large vision language models,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Finer: Investigating and enhancing fine-grained visual concept recognition in large vision language models,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.966532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:06.986025Z digest=sha256:c68e3f264418a50d61c707ecb3ffc0aa4149023ad13a1ab86b16189640812e7b

Observation 7d2999e9-fa44-4ff1-b40a-137410a2ded7 · outbound

This paper cites Benchmarking zero-shot recognition with vision-language models: Challenges on granularity and specificity,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Benchmarking zero-shot recognition with vision-language models: Challenges on granularity and specificity,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.947634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:06.990755Z digest=sha256:3a1afafd86f23477807d1a7cda51f03627df4541da250f484eab3c65808fa2f6

Observation ab451ce1-26b1-4cb7-81d3-32137ef45aaa · outbound

This paper cites Vilta: Enhancing vision-language pre-training through textual augmentation,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Vilta: Enhancing vision-language pre-training through textual augmentation,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.931918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:06.995224Z digest=sha256:16ff5d51b1b673b8e1c7decc0b6a13ff0dae04aa8d32fd993ca3f22caadcd1bf

Observation b637aff9-66da-4a03-bcb8-bc90540662f9 · outbound

This paper cites Towards vision-language mechanistic interpretability: A causal tracing tool for blip,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Towards vision-language mechanistic interpretability: A causal tracing tool for blip,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.914022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:06.999305Z digest=sha256:400f536a102d3e3321de2a9d2ea9b3d1cadfafeef9b97b4bd3addd786a53c10b

Observation a1d558dd-dfec-4602-8c5c-0d506ba8ad8f · outbound

This paper cites What matters when building vision-language models?.

CF-VLM:CounterFactual Vision-Language Fine-tuning What matters when building vision-language models?

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.895947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.003354Z digest=sha256:ce208937a5afa78426390f8fd6de29cde51f755471cb8203340db8d2f737ce22

Observation b9f42f80-8a62-4c16-b903-fd65508ef49a · outbound

This paper cites Fine-grained alignment for cross-modal recipe retrieval,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Fine-grained alignment for cross-modal recipe retrieval,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.878291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.013604Z digest=sha256:af057674df15021489ad812a00fdaf1781dc2891b2734d3ed115930b2209f0e9

Observation 0c8a2ede-f477-4e9f-bc3e-ac43428afc04 · outbound

This paper cites Localized triplet loss for fine-grained fashion image retrieval,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Localized triplet loss for fine-grained fashion image retrieval,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.858955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.018556Z digest=sha256:120eacff44c84d0b98a96fd4d574d098f97bcab1f599e981251ce93adc2e8d55

Observation 110a4311-512f-4704-9fe0-b9aaca3441f3 · outbound

This paper cites TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives.

CF-VLM:CounterFactual Vision-Language Fine-tuning TripletCLIP: Improving Compositional Reasoning of CLIP via Synthetic Vision-Language Negatives

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:01:08.073739Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.023061Z digest=sha256:4156a98b6ad8c5d6ab4977c3fc065fe1ae72f4fc45c68b9f749e81c1e33aefc1

Observation 3c99b13f-56ef-4af6-9ab3-ca42fbd4ccb9 · outbound

This paper cites Facenet: A unified embedding for face recognition and clustering,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Facenet: A unified embedding for face recognition and clustering,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.028168Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.028168Z digest=sha256:99228c8881e2aca4cc2a98b128282bc0ae4cf18b02ace9c1225986f788732eb4

Observation 8bbcaeb3-aba4-445b-8ba4-599906698382 · outbound

This paper cites CPL: Counterfactual prompt learning for vision and language models,.

CF-VLM:CounterFactual Vision-Language Fine-tuning CPL: Counterfactual prompt learning for vision and language models,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.825807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.037577Z digest=sha256:75054cc9c5da6e88fc8686878e6a36ded9e09c53570a675ad292af27841485db

Observation f2cdb927-37fb-41da-80bd-bbd97f3162c2 · outbound

This paper cites Attention Is All You Need.

CF-VLM:CounterFactual Vision-Language Fine-tuning Attention Is All You Need

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.046597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.046597Z digest=sha256:cddc60facd6eff4f3d0fa5b01d553e9d76de904d67e60b35554cd2724095ee5b

Observation b666d97b-07a6-4e8f-afd5-494e7f79f3db · outbound

This paper cites A survey on evaluation of large language models,.

CF-VLM:CounterFactual Vision-Language Fine-tuning A survey on evaluation of large language models,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.051054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.051054Z digest=sha256:49894d5d32d2f5e4e32d581e94e67fd3f3479b35dcf8a0ad3072d180283963a5

Observation ba152358-f5c9-440e-9f5e-efe7a008b601 · outbound

This paper cites A survey of visual transformers,.

CF-VLM:CounterFactual Vision-Language Fine-tuning A survey of visual transformers,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.808874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.055484Z digest=sha256:cf8b304584f7cb02e47b2d0cafbc6be0ff2f99536c0e5217b64c945c4678ecea

Observation 846e6f77-4295-47f2-ba90-0ee581c751d9 · outbound

This paper cites A simple framework for contrastive learning of visual representations,.

CF-VLM:CounterFactual Vision-Language Fine-tuning A simple framework for contrastive learning of visual representations,

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.059548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.059548Z digest=sha256:a4cf76272a5c5667a8c24243f4cab926de98f4a3e544febc21c2b26e2caee0b7

Observation e4f02ba1-a538-42dd-8408-84b1e0268399 · outbound

This paper cites Understanding contrastive representation learning through alignment and uniformity on the hypersphere,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Understanding contrastive representation learning through alignment and uniformity on the hypersphere,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.781994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.064955Z digest=sha256:9742f0d53d3a599b154ebad3c648f5918af1b70a0b58bc1c9a5e12e6e083d758

Observation 02c019f0-7f4e-4b7e-b6b5-d13924477716 · outbound

This paper cites Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision.

CF-VLM:CounterFactual Vision-Language Fine-tuning Scaling Up Visual and Vision-Language Representation Learning With Noisy Text Supervision

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.069602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.069602Z digest=sha256:27bb74ca32bf720fe1daa5e0c02073bc1a212b3625100e99351f8a28fe3d9428

Observation 3b9e9e6a-efdb-4d44-a706-1028d67ab742 · outbound

This paper cites Cogs: A compositional generalization challenge based on semantic interpretation,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Cogs: A compositional generalization challenge based on semantic interpretation,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.764968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.074611Z digest=sha256:dd26d98ab93bd4690b9f4da84e1eaaac1164ac7f5fdc9aa0aad0b785f5d25a16

Observation 43a53351-d77c-4251-b4fd-cea102c36b13 · outbound

This paper cites Learning what makes a difference from counterfactual examples and gradient supervision,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Learning what makes a difference from counterfactual examples and gradient supervision,

Reference 28

Resolution
verified exact
doi, observed 2026-08-07T05:01:07.374921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.085658Z digest=sha256:cfa3ae80566a42158adf749c6e2c810c04a2d0daa551ac0b594ab6846f154a8d

Observation 08b38393-f8f0-44a3-b187-ce2bd65a3274 · outbound

This paper cites Teaching clip to count to ten,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Teaching clip to count to ten,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.748372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.090433Z digest=sha256:fdf1ab1ce7b038f7d8733ee982aa60977dc733b7eed7f025056260b688291c3a

Observation e06c7165-c2c9-4efa-b367-550c11e41768 · outbound

This paper cites DISCO: Distilling Counterfactuals with Large Language Models.

CF-VLM:CounterFactual Vision-Language Fine-tuning DISCO: Distilling Counterfactuals with Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.094854Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.094854Z digest=sha256:abdbadab629f366feeee99698bd7bf1c0abf522d21fe6bb928c5158364f237d3

Observation 2bad5841-e5b9-4482-a9f0-64aaf70d9a71 · outbound

This paper cites Counterfactually measuring and eliminating social bias in vision-language pre-training models,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Counterfactually measuring and eliminating social bias in vision-language pre-training models,

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.100956Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.100956Z digest=sha256:6fcceb501b64d07a5596794aecee19a3be4ff146ed8414bb655aaf9b3e9ad44d

Observation e09ab765-4ce8-4e63-9431-0f3250b93549 · outbound

This paper cites Counterfactual attention learning for fine-grained visual categorization and re-identification,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Counterfactual attention learning for fine-grained visual categorization and re-identification,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.842305Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.105617Z digest=sha256:198a8f3239d6bc6f4007e0b96558ea859549ce9bd771d5bb8668d6138ba3ebe7

Observation af420799-f337-4ec0-b5f4-19b03960bb1c · outbound

This paper cites Counterfactual samples synthesizing and training for robust visual question answering,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Counterfactual samples synthesizing and training for robust visual question answering,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.730534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.110453Z digest=sha256:aa30cff3af7885555616ea03d4e05d8fe5b2e2d6750380f70604ef615e09b995

Observation 5733a09a-3318-489a-98d1-376fb7498080 · outbound

This paper cites Qwen2.5 Technical Report.

CF-VLM:CounterFactual Vision-Language Fine-tuning Qwen2.5 Technical Report

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.114672Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.114672Z digest=sha256:0ae6122877a99185488fe0b37e88cd23db184623da9515e848ba1d78d15d8a0b

Observation f4593c00-6d8a-4f6d-aa1d-a6a6c4f83ef0 · outbound

This paper cites Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Conceptual 12M: Pushing web-scale image-text pre-training to recognize long-tail visual concepts,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.711237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.119470Z digest=sha256:050f035c941eb52793e77bdcc207a52f168c95e119f8010600fc425f363bb066

Observation fda6020a-f7a7-4d80-8212-4c0c609522ea · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.692774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.123742Z digest=sha256:5827c3983bb3ad138a9e02f975857afc848f126fb3e5ff800826f3fca9fc2c3f

Observation 060bdcea-5d8c-4cec-9572-47b5edeed4be · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

CF-VLM:CounterFactual Vision-Language Fine-tuning Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.128034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.128034Z digest=sha256:bdfc8145b41091d09f7101987a5a4acb8e9d417a18a2ba073f9af227fec19139

Observation efa34ba6-5906-48d1-9a18-de51b7693da2 · outbound

This paper cites ConMe: Rethinking Evaluation of Compositional Reasoning for Modern VLMs.

CF-VLM:CounterFactual Vision-Language Fine-tuning ConMe: Rethinking Evaluation of Compositional Reasoning for Modern VLMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.132424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.132424Z digest=sha256:0b3ee7858915e93e62ca38823a955cdb18ed91fd2fa14fc9428c084edac54631

Observation 866c9125-f995-4dfe-bd87-feb0efc052a2 · outbound

This paper cites When and why vision-language models behave like bags-of-words, and what to do about it?.

CF-VLM:CounterFactual Vision-Language Fine-tuning When and why vision-language models behave like bags-of-words, and what to do about it?

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.674987Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.136591Z digest=sha256:acf2a1982caef88258cec21b5282c6ef1ab6aa1aae94081b4e5eb9b8ea2853b1

Observation e0bd20ed-6479-47b0-bb65-0026b8d5f9d2 · outbound

This paper cites VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations.

CF-VLM:CounterFactual Vision-Language Fine-tuning VL-CheckList: Evaluating Pre-trained Vision-Language Models with Objects, Attributes and Relations

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.140504Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.140504Z digest=sha256:dbba314d51e46b4b22c456cf1665d0f89769ca5008d338408845aa3cf7f76b6c

Observation f12b6b44-3b06-4d04-a7a4-ca0e8df4fa74 · outbound

This paper cites ImageNet Large Scale Visual Recognition Challenge,.

CF-VLM:CounterFactual Vision-Language Fine-tuning ImageNet Large Scale Visual Recognition Challenge,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.657970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.145231Z digest=sha256:39629b25be574aa5b4f59b75fdeda0b067a72367ebde86c15d811b95bf1ee2b5

Observation 799bc5f2-8288-4427-b732-aea50db657ef · outbound

This paper cites From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,.

CF-VLM:CounterFactual Vision-Language Fine-tuning From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.640639Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.149485Z digest=sha256:cd5c714e1ba497e58227b3f721e7605d83b563f9dfb0e34996cacb55d5ee2fde

Observation ed03be52-36d3-4416-990c-4cafb667f01d · outbound

This paper cites Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional Understanding.

CF-VLM:CounterFactual Vision-Language Fine-tuning Contrasting Intra-Modal and Ranking Cross-Modal Hard Negatives to Enhance Visio-Linguistic Compositional Understanding

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:01:07.741382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.153634Z digest=sha256:8a5f08f1a58f2e93ec7b92b31a335b0c7bceb0330b495339abf62e5a811e5b8b

Observation c6d5e468-171b-47c6-a514-c88c7e26d27c · outbound

This paper cites Structure-CLIP: Towards Scene Graph Knowledge to Enhance Multi-modal Structured Representations.

CF-VLM:CounterFactual Vision-Language Fine-tuning Structure-CLIP: Towards Scene Graph Knowledge to Enhance Multi-modal Structured Representations

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:01:07.697666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.163971Z digest=sha256:5c2cb7333a49b7b2c72bb307f605657d673f9ab9a9d5e3d9560d0f8ad9c6247e

Observation 036de79c-a746-4d6b-8d94-91e7d50d6c8f · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

CF-VLM:CounterFactual Vision-Language Fine-tuning Improved Baselines with Visual Instruction Tuning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.168343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.168343Z digest=sha256:251561718c4da827de84255e887b94003d03d1d5f8a7b6666c4409fc243fca39

Observation 556d253d-ee8b-4367-bd2b-bd573994ec36 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

CF-VLM:CounterFactual Vision-Language Fine-tuning InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.173286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.173286Z digest=sha256:19fd853b3eefda2ac04b049ed5677ebf4315918c46a50df108951cb63cde2ac1

Observation a3d19d3a-9461-4759-90de-978b5b273bc3 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

CF-VLM:CounterFactual Vision-Language Fine-tuning MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.178097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.178097Z digest=sha256:967530ee66dd327057406813760f9ac7b117b5963293b678d2c94f92a7dab4c0

Observation 3838e51a-9024-44e9-8b6c-4840d6b2b301 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

CF-VLM:CounterFactual Vision-Language Fine-tuning SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.182593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.182593Z digest=sha256:b35ee025a9785df33818166078a15e60beb10859b25572af960157cd69ffac7d

Observation 4d6eedb6-ee1e-4ec7-a636-f31252676b0a · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

CF-VLM:CounterFactual Vision-Language Fine-tuning Evaluating Object Hallucination in Large Vision-Language Models

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.187178Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.187178Z digest=sha256:bdf0d72c534d8df079e6d61dace1bb17a75fc395c0399b790196b8d8a899fe4d

Observation 8903ba8e-e6da-43b9-b714-2c24b5ee9f14 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

CF-VLM:CounterFactual Vision-Language Fine-tuning MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.191460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.191460Z digest=sha256:48559d4c030e8fb889a5aa99ed347487a1c60b91ba9ee960e02f2745d0fc1b6a

Observation ce3d1c0b-f348-4bdc-a973-23c90907feb9 · outbound

This paper cites Vision-and-Language Pretrained Models: A Survey.

CF-VLM:CounterFactual Vision-Language Fine-tuning Vision-and-Language Pretrained Models: A Survey

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.195795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.195795Z digest=sha256:42f0dce5c5337ba28479678c048ecebcc288d3519ae0f1ce6802bfd3f3aeed01

Observation 6f46b262-c841-40b7-afca-bda312bc53a6 · outbound

This paper cites Exploring the frontier of vision-language models: A survey of current methodologies and future directions,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Exploring the frontier of vision-language models: A survey of current methodologies and future directions,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.200219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.200219Z digest=sha256:2e8fed9f2bc183b8d2fdd6b6fe9cc39c96c2a2f39af890164ea7f384562d66b6

Observation 04362e3a-ca1b-4aa9-b3c2-5162f4df9c23 · outbound

This paper cites Vision-language models for vision tasks: A survey,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Vision-language models for vision tasks: A survey,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.204431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.204431Z digest=sha256:832ada27d51ac41da565caf1c11dd19923d937034520dc52622cb0ed03be202b

Observation 1064166d-50a6-44ac-a19e-7dead8da82db · outbound

This paper cites Counterfactual vision and language learning,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Counterfactual vision and language learning,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.613419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.208880Z digest=sha256:2a772b72a465284bb2a75ed194fae84de76f632f3f821ee5f68070f9ba84e99c

Observation 5256965e-acca-49f2-83fb-b036e5e749f4 · outbound

This paper cites Counterfactual reasoning for multi-label image classification via patching-based training,.

CF-VLM:CounterFactual Vision-Language Fine-tuning Counterfactual reasoning for multi-label image classification via patching-based training,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.597922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.213460Z digest=sha256:548215f8a45d862a6590eb60d3ce5cd0217a0ab94d8e5fb44eb33292ad97e7dc

Observation 75ad03d2-c6ca-4845-a852-39b0a661e028 · outbound

This paper cites Causal Graphical Models for Vision-Language Compositional Understanding.

CF-VLM:CounterFactual Vision-Language Fine-tuning Causal Graphical Models for Vision-Language Compositional Understanding

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:01:07.719377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.218215Z digest=sha256:6ad4a1fb3bc5366c35df5bdaa0df5748c3979af02c2795fcc1721fd87186f0ae

Observation f4e29da0-3367-43e2-b371-50cbf99390a1 · outbound

This paper cites CPL: Counterfactual Prompt Learning for Vision and Language Models.

CF-VLM:CounterFactual Vision-Language Fine-tuning CPL: Counterfactual Prompt Learning for Vision and Language Models

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:01:07.433090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.222289Z digest=sha256:3a9362771ddeeb96c3b9bcad50f6b1835c0b269d6a62cfca897cf19c7aee68b6

Observation 82b5f82b-87b4-4f61-8a8e-38e472eac53b · outbound

This paper cites Counterfactual Visual Explanations.

CF-VLM:CounterFactual Vision-Language Fine-tuning Counterfactual Visual Explanations

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.227257Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.227257Z digest=sha256:ba8ba32f088c90266b0389c0657ccacea84f576d0ed94f9397d358b9d3f83482

Observation 8e4108df-7b85-443d-af3d-3ff07f794393 · outbound

This paper cites Counterfactual Attention Learning for Fine-Grained Visual Categorization and Re-identification.

CF-VLM:CounterFactual Vision-Language Fine-tuning Counterfactual Attention Learning for Fine-Grained Visual Categorization and Re-identification

Reference 60

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:01:07.410653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.231457Z digest=sha256:e83269112b18f49321c1c73021bc1f3f501ebff93b87a8e20a22795aa166077a

Observation af0ade69-265d-4dfe-87f4-5bac436fca09 · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.580872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.236954Z digest=sha256:eb1743127ca69b1bef3abc7306867de995d40611f6b9f02980d4da46d9811f48

Observation 38178384-683d-4c49-b0d8-2d9434e28870 · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.565173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.241166Z digest=sha256:c16b199f0c6164a9c6f90dfa500697ce460bef9c8dabfd7403cff1ed441c9220

Observation b4b2c7b5-42f4-4330-97b6-da6084a79ff5 · outbound

This paper cites Rules: • Modifyonly one thing(either one attribute or one causal link).

CF-VLM:CounterFactual Vision-Language Fine-tuning Rules: • Modifyonly one thing(either one attribute or one causal link)

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.547774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.245312Z digest=sha256:e619f9c6bd183be429bbd96b48f8ec8a60df820bd40fa2749f969152a9507f95

Observation a125a548-442b-48f1-8e97-5a7935069a43 · outbound

This paper cites Example Input: A young woman holding a racket hit the ball, and the ball flew outward.

CF-VLM:CounterFactual Vision-Language Fine-tuning Example Input: A young woman holding a racket hit the ball, and the ball flew outward

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.528325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.250450Z digest=sha256:ab1279de02c2a98d061cb239b28370db061cc3aecd2e0d89b6a83b8f0e7f084b

Observation 98563f15-079f-4b01-877a-75ae0aa6db83 · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 67

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.512290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.255371Z digest=sha256:f25086dcf7b326340ee3f179113dce2efa061d3a821a40a5edfcf1daa4bdbc57

Observation 87f61e64-b74d-41b0-90d6-958147e1b9ef · outbound

This paper cites Output (causal):.

CF-VLM:CounterFactual Vision-Language Fine-tuning Output (causal):

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.495564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.260715Z digest=sha256:bc7e2a0a37207104c40c48f3306bcdaa525529a3043c8972f6a2b9cba231978f

Observation b3fa6ee5-d2a5-4657-a4b6-db4745d88fe1 · outbound

This paper cites dirt” with “paved roads.

CF-VLM:CounterFactual Vision-Language Fine-tuning dirt” with “paved roads

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.474047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.265588Z digest=sha256:19330c39d27b36121dc120287651e5cbd89a9f422fa3b08ac997fd7a87e70c9a

Observation 37f1a45c-cfb8-476f-877b-b90457564489 · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 70

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.446664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.269999Z digest=sha256:5e5d29119c57b3fd66d5b8be4a2b8efac196543d04fb480692cea513372b5161

Observation b49d51f6-6c25-4ba7-827c-78703bf04ab2 · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 71

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.428518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.273979Z digest=sha256:42d1500ec0c696d206bd870bace3f41dbb1f5381ecba9196255a47983564a02f

Observation efcd3bf2-24ca-48b3-95c4-1d7b15b217ad · outbound

This paper cites causal decision points.

CF-VLM:CounterFactual Vision-Language Fine-tuning causal decision points

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.412119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.277833Z digest=sha256:4b5e40ef416c585866b896e189e43b27e0d3df95c1f4b1973dcbe10b75e52199

Observation 1a1cde11-e036-45e2-b861-38daf8f12ede · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.393260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.282997Z digest=sha256:7cce908842d34c44da8836909e30fc243628513f8415a0204514b0bee5df6bd9

Observation 5504e3cd-2e3b-4cfa-854b-3e4a660c462e · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.375706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.287322Z digest=sha256:2af26f2857157cbbe88dc7f7ead92a88b8d9cfa012cdd83c886efe44fda7371b

Observation 72662f1d-9f16-4588-bd8f-3ce329cae31a · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 75

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.357028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.292067Z digest=sha256:3bdc0f76c7d0db9325c8b21fba2d7af68c1955bf94ae01192d457186c2d8f1c1

Observation b4b55706-9edf-4850-8635-25e55d99d1eb · outbound

This paper cites Two people on motorcycles riding them.

CF-VLM:CounterFactual Vision-Language Fine-tuning Two people on motorcycles riding them

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.338390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.297273Z digest=sha256:b3e9613ae3a33975d015e8a364443315d2e18db5d67199cf4bb8274047b470dd

Observation ef86dff8-a5a1-4677-9571-c66da1bcdad5 · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.316820Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.302045Z digest=sha256:a7b321ec5ff8de4836f2860e5af0100ad97c6c77ee66ad6d048098cf94c3231d

Observation c2267686-b567-446f-825c-7242bcfec410 · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 78

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.298808Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.306355Z digest=sha256:adf9482e8fbde46bcfc71b19092cc1aedfbaff5622bff9b01f232b584a812747

Observation dc45ff25-b0fd-42b3-9c54-1e74ba3d9c7a · outbound

This paper cites causality.

CF-VLM:CounterFactual Vision-Language Fine-tuning causality

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.281488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.310308Z digest=sha256:ff664da6381bde5a8d1e01d2444518aa20f7036c70ffbf8ec83e791254f9de67

Observation 5e3b4834-cbbe-463d-ad1c-215fa2526338 · outbound

This paper cites parallel realities.

CF-VLM:CounterFactual Vision-Language Fine-tuning parallel realities

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.260724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.314460Z digest=sha256:ef105c15dcd9059f4853402c683fd745b6e0ece0adb9034cdeaa8f9b77deef8e

Observation 443bbec6-8f01-47e4-a298-1b68d0e77f84 · outbound

This paper cites These are employed in Lcsd to help the model learn semantic scene boundaries.

CF-VLM:CounterFactual Vision-Language Fine-tuning These are employed in Lcsd to help the model learn semantic scene boundaries

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.243881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.319372Z digest=sha256:ebe33afaa49bbc0d90fb20f11448ed3bb635c06dcc882ac203f6d3b95b6f83e5

Observation 7aa85fba-ece3-4611-85b4-63191fb81046 · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 82

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.224933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.323874Z digest=sha256:838c7342b84eb651bd6ee6e16bc0929145d149ca9c4746da0962b4c15486a519

Observation 527bc56b-0dc5-4d77-bf56-110aaa7a9aa8 · outbound

This paper cites an unresolved cited work.

CF-VLM:CounterFactual Vision-Language Fine-tuning Unresolved cited work

Reference 83

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:01:08.210112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.327832Z digest=sha256:eeab9ea9bbb1bcedcbf91d6bcd56d21c3c28b97d5c81b06651e7f1a1ee939d5a

Observation 39db48cf-7edf-4131-b359-6611f6b84ee1 · outbound

This paper cites kicking a ball.

CF-VLM:CounterFactual Vision-Language Fine-tuning kicking a ball

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:01:08.194421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T05:01:07.332466Z digest=sha256:6981f5e397dfc4c25b7106add77948d39cf7021ccec525cabd843908ac60f4cc

Observation d6f96fae-5544-4683-a0ec-c5087db8e7fc · outbound

This paper cites COGS: A Compositional Generalization Challenge Based on Semantic Interpretation.

CF-VLM:CounterFactual Vision-Language Fine-tuning COGS: A Compositional Generalization Challenge Based on Semantic Interpretation

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.080997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.080997Z digest=sha256:efdbb270d9806548d25afed8f907e660f1ad35ef7b26707c1dc7659e54d207d1

Observation f17ba9ae-74f7-42d7-bef8-edc49c67d258 · outbound

This paper cites What matters when building vision-language models?.

CF-VLM:CounterFactual Vision-Language Fine-tuning What matters when building vision-language models?

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-07T05:01:07.007888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:01:07.007888Z digest=sha256:c7c24a86d6a147dd063664c54ef894894a4fbdaec59eef806070063b021b7fa7

Pith citing papers

Observation ed3a27c0-7fbc-410e-b5e1-f1305cb5945a · inbound

OSC: Cognitive Orchestration through Dynamic Knowledge Alignment in Multi-Agent LLM Collaboration cites this paper.

OSC: Cognitive Orchestration through Dynamic Knowledge Alignment in Multi-Agent LLM Collaboration CF-VLM:CounterFactual Vision-Language Fine-tuning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-05T05:51:29.406876Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T05:51:29.406876Z digest=sha256:f8be23d2036ae61b469786957f23d714431e0cd2833202f89a20fd0ed92b19b3

Observation f99d24fa-3e34-4deb-bb1e-d630323c6d50 · inbound

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs cites this paper.

When Vision Overrides Language: Evaluating and Mitigating Counterfactual Failures in VLAs CF-VLM:CounterFactual Vision-Language Fine-tuning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-02T22:14:31.005765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T22:14:31.005765Z digest=sha256:09895f00512d7dd295ccb962f2c339c847f812ea4b07eae27b521fe0e0ebb593

Observation e43294dc-c047-41b0-97e7-665272a05f21 · inbound

CFPO: Counterfactual Policy Optimization for Multimodal Reasoning cites this paper.

CFPO: Counterfactual Policy Optimization for Multimodal Reasoning CF-VLM:CounterFactual Vision-Language Fine-tuning

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-07-04T10:09:44.503204Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-26T09:09:09.716984Z digest=sha256:d0f5089d7d987ab4a79e57ac9dd6b0fd056cce4101d51a915ef1e59dacdd09ae

Observation 03766e84-061f-41f4-9881-13016b986e34 · inbound

SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning cites this paper.

SIVA-RL: Sensitivity-Invariance Visual Alignment for Multimodal Reinforcement Learning CF-VLM:CounterFactual Vision-Language Fine-tuning

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-02T03:20:41.148082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T03:20:41.148082Z digest=sha256:3531098276bc3c95c2006208c0f4a8cbd24cda255386a443bc11fba3730b94cb