Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T04:40:11.919275Z
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 1 inbound Pith citation observation for arXiv:2412.18675.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T04:40:11.919275Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-11T20:06:55.102341Z
A source-named dated measurement, never combined with another source.
Source: cited_works
82 of 82 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation ca505ee9-72fb-4e9a-a55d-ac784213d483 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models https : / / github
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 572d7d2e-9b5f-491a-b792-6fbe0580eb55 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models https://github.com/ Seth-Park/RobustChangeCaptioning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 52bae144-4cda-422f-a59e-5d923c86ecf9 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models https : / / github
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1faf6095-33c5-4b2a-a557-8ea8225e67ae · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models siglip needs registers for comparison, here’s dino-v2 with registers. it has five extra tokens for the model to work with: one cls token and four
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d2e3c8b0-ba72-4386-86f9-ea8e4aba3a9a · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Quantifying Attention Flow in Transformers
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5bf9054-d6c4-4664-a62f-1661062c4527 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Debugging Tests for Model Explanations
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6d8e999-7bb2-44ab-bf4f-1bf9d5dca8db · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Explaining image clas- sifiers by removing input features using generative models
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 144db66c-0baa-4de4-b9e4-2c6c7f9a1b98 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models METEOR: An auto- matic metric for MT evaluation with improved correlation with human judgments
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 05f03c3c-4620-4291-aacd-3398a9afe754 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Sam: The sensitivity of attribution methods to hyperparameters
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation ed4a48e5-0c7a-4dd5-bd92-f85d1f7880e5 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Rewriting a deep generative model
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8737a29c-99eb-40cf-8ece-427e7567a568 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Understanding the role of individual units in a deep neural network.Proceedings of the National Academy of Sciences, 117(48):30071–30078, 2020
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a0c0dede-a94b-4ca4-8ba1-95007dddd4ab · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Better plain ViT baselines for ImageNet-1k
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c2d1d0a-e1a8-481c-b063-7463c3ddbd9e · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Vixen: Visual text comparison network for image difference captioning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6b30731e-e5d6-47e8-98e8-98dfdc6f40e9 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Michigan man wrongfully accused with facial recognition urges Congress to act
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 49c6539d-a73e-4d1d-b093-0076c82a0dc1 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 14bab65c-f0d2-4059-9074-f6670e325244 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Emerg- ing properties in self-supervised vision transformers
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a5067b96-9852-4732-b00f-91c98e40de08 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 40341ed4-a55e-447a-a4aa-54dfcd54e1a5 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Crossvit: Cross-attention multi-scale vision transformer for image classification
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44d3f4f7-b95e-4f42-a53a-356718601b0b · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models gscorecam: What objects is clip looking at? In Proceedings of the Asian Conference on Computer Vision , pages 1959– 1975, 2022
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 1267bee9-71b5-4e73-b3b6-f43dbeeea742 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Concept whitening for interpretable image recognition
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation af7352a7-ba5f-4632-a13f-533e7d74791e · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Blender - a 3D modelling and rendering package
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 736222fb-6438-4532-8d55-203299ba9e7f · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Vision transformers need registers
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation bf3f9988-9c2a-424f-a7df-7e787b8ec2ea · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models An image is worth 16x16 words: Transformers for image recognition at scale
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7ce08f67-a395-4f4a-99c9-1fc86e29db59 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Towards a multimodal framework for remote sensing image change retrieval and captioning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c5e931e8-c616-4c4d-be32-c7d6d330c745 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Interpretable explana- tions of black boxes by meaningful perturbation
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fa8df934-cec8-4203-847e-8eb59d9ec776 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Openagi: When llm meets domain experts
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6c20fc50-e27a-419f-8cc9-e7e25106ee57 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Clip4idc: Clip for image difference captioning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 10445064-b551-4a21-9e49-c0eac24d3331 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Wrongfully arrested man sues Detroit police over false facial recognition match
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b2e26d52-3ea7-47a0-98de-c919a85fc0f7 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Which tokens to use? investi- gating token reduction in vision transformers
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6f1d1728-df68-4686-93a9-841a3a600b0a · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Flawed Facial Recognition Leads To Arrest and Jail for New Jersey Man
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 66abd07e-8b50-48e4-a3bf-457f03e1b8ab · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Change captioning: A new paradigm for multitem- poral remote sensing image analysis
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 45c94ae9-fc16-402b-8211-7254586c5c2d · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models OneDiff: A Generalist Model for Image Difference Captioning
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5b022a48-c6e7-4b77-9dd2-bb0a1b0408fb · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Summers, and Yingying Zhu
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8937f13c-a5ff-45a9-8869-fd583805823e · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Learning to describe differences between pairs of similar images
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 2ef8ce1a-999d-44fb-bdec-d0ef11d773ad · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d64106d8-b0f3-4233-834b-4fe7c4677643 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Now You See Me (CME): Concept-based Model Extraction
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7e1137a7-d478-4a28-8324-a5f2b80514da · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Vilt: Vision- and-language transformer without convolution or region su- pervision
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation acb038ef-8d9f-4e25-aa68-f565ace05d34 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Concept bottleneck models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation b78ace51-0a8e-4d00-9630-b6f806d8f6b6 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb05402c-8232-4edd-9bf2-f24441ffc57b · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Dynamic graph enhanced contrastive learning for chest x-ray report generation
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 70d59d8e-abcb-4198-b3eb-721d097a2876 · outbound
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 257477fb-1220-4192-835b-30b774815ffb · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models ROUGE: A package for automatic evaluation of summaries
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5d62f66c-5431-4e21-9d37-08600d125242 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Improved baselines with visual instruction tuning
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 53a7c641-d834-4a9e-89ca-f778bfba4cf4 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Interpretability Beyond Classification Output: Semantic Bottleneck Networks
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a1cb517a-7d4d-4927-85b6-aa45e5449c18 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models SGDR: Stochastic gradi- ent descent with warm restarts
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 121a2486-b8b7-4d04-996d-0c0d557d5cd8 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a788c71f-1001-40b8-913b-e3b3c820acae · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Spot the difference: Difference visual ques- tion answering with residual alignment
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 020238cf-4847-4ad4-96e5-1f5fe5cbd9a5 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Locating and editing factual associations in gpt
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dc535ae8-4efb-4fea-b512-aa420ea07542 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Memory-based model editing at scale
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0af686c2-2903-48ef-947b-f6676e1905a1 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models The ef- fectiveness of feature attribution methods and its correlation 11 with automatic evaluation scores
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation f602c512-5c6d-4331-bf93-54a9fb05c6a3 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Improving change detection by incor- porating correspondence information
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 006609de-5b9d-41b5-855c-282ccaf6df6e · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models How explainable are adversarially-robust CNNs?
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 04f704ac-538f-45c5-913c-a5ce1b336319 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a2cc38e9-c70d-4f94-9025-937ff13cd304 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Bleu: a method for automatic evaluation of machine translation
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 69ca58c8-6fa3-436f-a35f-e3c9c5b7f636 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Robust change captioning
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3369db3a-15af-4f37-8cbd-017abb361e0c · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models PEEB: Part-based image clas- sifiers with an explainable and editable language bottleneck
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c651e62f-0f01-40e2-aaf2-328d54eda638 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Deepface-emd: Re-ranking using patch-wise earth mover’s distance improves out-of- distribution face identification
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation fcc54815-77c0-48e6-9886-1897cb300b12 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Fast and interpretable face identification for out-of- distribution data using vision transformers
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d099e3ea-c21e-4983-a3de-738ec096d427 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Describing and localizing multiple changes with transform- ers
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a575f53f-daeb-4c8f-a819-4e6801a15dcd · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Learning transferable visual models from natural language supervi- sion
Reference 60
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 455668d7-94eb-4261-be5f-09a846fc08ae · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models The new lawsuit that shows facial recog- nition is officially a civil rights issue
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6ffd5800-7f6c-449c-96e5-b67d3897a2f4 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models The change you want to see
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c99281f6-7dd6-4ec6-b84d-d9af97f4911b · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Explaining deep neural networks and beyond: A review of methods and applications
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7767697d-892a-4b99-8615-ab0876f80a16 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models The pandemic is testing the limits of face recogni- tion
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4fd33787-fdc1-4522-9c73-43ca0576e1c4 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Reclip: A strong zero-shot baseline for referring expression compre- hension
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 25100f87-6b2a-4528-a6f4-e83acbbe0c03 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models The stvchrono dataset: Towards continuous change recognition in time
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 4a825cc5-9115-48bf-8852-d3b040e7d297 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Resolution-robust large mask inpainting with fourier convolutions
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8c1f87f8-280c-4bdc-81c4-fe8e5ed133f8 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Expressing visual relationships via language
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 3aaf7735-ca9e-435b-836a-048e6344c70e · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Eyes wide shut? exploring the visual shortcomings of multimodal llms
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0f2f387e-adac-44f4-8843-d400cb5a0b6d · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Attention is all you need
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8335ec1d-9c70-4ca6-a125-62c81907d01a · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Cider: Consensus-based image description evalua- tion
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e397a496-d379-4258-8df5-f49192b60329 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Learning bottleneck concepts in image classification
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation a3dd02bc-8e3e-41ef-909c-d3d422e36bf5 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Co-attention for conditioned image matching
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation d5f7cd9f-531a-4181-9fb7-a3ae3bfa65dc · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models L2C: Describing visual differences needs semantic under- 12 standing of individuals
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation e72e372e-33ec-4ddf-9756-728d21cf47d5 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Image difference captioning with pre-training and contrastive learning
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 380f9dac-d355-4e39-a4b4-e42913e5fef6 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Scaling vision transformers
Reference 76
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd6d6b9c-1bba-4c43-aac2-e3c8e67302c6 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Top-down neu- ral attention by excitation backprop
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7c33fa14-d549-4911-9c57-9ed1856c1fe4 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models BERTScore: Evaluating Text Generation with BERT
Reference 78
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fb05bac-2db6-426f-8c98-0f79bd294d30 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Learning deep features for discrimina- tive localization
Reference 79
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 06b2ccbf-3274-4bc7-bc81-c6a38ecabb64 · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Implementation details A.1
Reference 2016
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5a9f7c76-530c-4161-89d7-1f81053b20cc · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Unresolved cited work
Reference 2019
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 95fc221d-0d70-49d9-b8ec-01e9cebb77ec · outbound
TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Unresolved cited work
Reference 2024
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5079130-d1f2-4118-a6e1-627332b5c157 · inbound
Legible-by-Construction: Attention and End-to-End Transformers TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.