Pith. sign in

Paper Citation Record · LEDGER

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models

As of 11 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 1 inbound Pith citation observation for arXiv:2412.18675.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18675 v4

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:40:11.919275Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-11T20:06:55.102341Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

82 of 82 outbound references displayed

  • verified exact1
  • verified fuzzy57
  • unresolved23
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation ca505ee9-72fb-4e9a-a55d-ac784213d483 · outbound

This paper cites https : / / github.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models https : / / github

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T04:40:10.642258Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:40:10.642258Z digest=sha256:311711c6e3c547a0b9c883b590ee643c783eb69089f4346a71591fc022b5cd78

Observation 572d7d2e-9b5f-491a-b792-6fbe0580eb55 · outbound

This paper cites https://github.com/ Seth-Park/RobustChangeCaptioning.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models https://github.com/ Seth-Park/RobustChangeCaptioning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T04:40:10.651655Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:40:10.651655Z digest=sha256:8c0e06dfd10b93cf07e1f46b88ec9d762a01cb9a7114a40b85d886d24d8cfeb6

Observation 52bae144-4cda-422f-a59e-5d923c86ecf9 · outbound

This paper cites https : / / github.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models https : / / github

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:16.727100Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:10.663724Z digest=sha256:443888d3ec2890d3a2163b9a1038c1a048cddffb630465690050de2e8731cf94

Observation 1faf6095-33c5-4b2a-a557-8ea8225e67ae · outbound

This paper cites siglip needs registers for comparison, here’s dino-v2 with registers. it has five extra tokens for the model to work with: one cls token and four.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models siglip needs registers for comparison, here’s dino-v2 with registers. it has five extra tokens for the model to work with: one cls token and four

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:16.686292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:10.671891Z digest=sha256:de017d67aad5ef5338bcccb41ca8516179c491ffe23d9df9d3e578768b2b8df8

Observation d2e3c8b0-ba72-4386-86f9-ea8e4aba3a9a · outbound

This paper cites Quantifying Attention Flow in Transformers.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Quantifying Attention Flow in Transformers

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T04:40:10.680273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:40:10.680273Z digest=sha256:e8d9d5df7e5043381fef0ff8524408b96ffb6ecd720587bccc693dcf0cc0be1b

Observation f5bf9054-d6c4-4664-a62f-1661062c4527 · outbound

This paper cites Debugging Tests for Model Explanations.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Debugging Tests for Model Explanations

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T04:40:10.687978Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:40:10.687978Z digest=sha256:5c7d6f76a2cfe6da9eec6ec4286e7a2213aba6ae917f8c0e7d6b8df02e87045e

Observation e6d8e999-7bb2-44ab-bf4f-1bf9d5dca8db · outbound

This paper cites Explaining image clas- sifiers by removing input features using generative models.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Explaining image clas- sifiers by removing input features using generative models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:16.640981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:10.702027Z digest=sha256:78085ae387451e6ebb1196e156122bca82060b6cc3aed692ce19bb8543a5ed6a

Observation 144db66c-0baa-4de4-b9e4-2c6c7f9a1b98 · outbound

This paper cites METEOR: An auto- matic metric for MT evaluation with improved correlation with human judgments.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models METEOR: An auto- matic metric for MT evaluation with improved correlation with human judgments

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:16.480113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:10.711233Z digest=sha256:8749364868f24b75bbea4bd71a70b97e0c05b15df05fb58da757f538fa8f8b5f

Observation 05f03c3c-4620-4291-aacd-3398a9afe754 · outbound

This paper cites Sam: The sensitivity of attribution methods to hyperparameters.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Sam: The sensitivity of attribution methods to hyperparameters

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:16.424733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:10.719292Z digest=sha256:a769dd25a43a282bfa29871a24b2a7922f616da078ddd5e8ce32aa569486b6e6

Observation ed4a48e5-0c7a-4dd5-bd92-f85d1f7880e5 · outbound

This paper cites Rewriting a deep generative model.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Rewriting a deep generative model

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:16.389107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:10.724878Z digest=sha256:7e0656d23b2bed6e085a536b0e83ddb4768b3b78088f3bd5a0106054f89139e3

Observation 8737a29c-99eb-40cf-8ece-427e7567a568 · outbound

This paper cites Understanding the role of individual units in a deep neural network.Proceedings of the National Academy of Sciences, 117(48):30071–30078, 2020.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Understanding the role of individual units in a deep neural network.Proceedings of the National Academy of Sciences, 117(48):30071–30078, 2020

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:16.346017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:10.730663Z digest=sha256:53daa267ac0a72a1251435051d8e22473a816fedb956202f1e4aa0b2581b5b69

Observation a0c0dede-a94b-4ca4-8ba1-95007dddd4ab · outbound

This paper cites Better plain ViT baselines for ImageNet-1k.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Better plain ViT baselines for ImageNet-1k

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T04:40:10.735003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:40:10.735003Z digest=sha256:7900c665eb06728d49d90ed2eb8515ccfe1ee8f0ad2480dd14ff7ae65ee71c60

Observation 7c2d1d0a-e1a8-481c-b063-7463c3ddbd9e · outbound

This paper cites Vixen: Visual text comparison network for image difference captioning.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Vixen: Visual text comparison network for image difference captioning

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:16.312956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:10.741079Z digest=sha256:55f4d379a6a157a8cbd6a58ebf76d979bbf80d9d72ca914f95cc624928be5ed3

Observation 6b30731e-e5d6-47e8-98e8-98dfdc6f40e9 · outbound

This paper cites Michigan man wrongfully accused with facial recognition urges Congress to act.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Michigan man wrongfully accused with facial recognition urges Congress to act

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:16.257021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:10.747022Z digest=sha256:76add4ade094479c109a25dd3cc62a6b78b92bf5d3ab4b56195fe8254d28e6a1

Observation 49c6539d-a73e-4d1d-b093-0076c82a0dc1 · outbound

This paper cites an unresolved cited work.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:40:16.200732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:10.753586Z digest=sha256:54601d7a67d5776e925330d20c25d2daf2af519507e35913dd736774bbdb563f

Observation 14bab65c-f0d2-4059-9074-f6670e325244 · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Emerg- ing properties in self-supervised vision transformers

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:16.152467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:10.774438Z digest=sha256:fbbef57034c574eca2c2c219de01a346c4236e3509d70a3f830798c2fad00069

Observation a5067b96-9852-4732-b00f-91c98e40de08 · outbound

This paper cites Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:16.109275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:10.788975Z digest=sha256:16b3291bf6e9bed50feb774e1b97f3aeb8f40c77f2a8d1892e4dd0e8814edff3

Observation 40341ed4-a55e-447a-a4aa-54dfcd54e1a5 · outbound

This paper cites Crossvit: Cross-attention multi-scale vision transformer for image classification.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Crossvit: Cross-attention multi-scale vision transformer for image classification

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T04:40:10.798524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:40:10.798524Z digest=sha256:db209362940bd5d4bf2d91b74a6326ac77dee2096058f24adc2941db5cbc6cb8

Observation 44d3f4f7-b95e-4f42-a53a-356718601b0b · outbound

This paper cites gscorecam: What objects is clip looking at? In Proceedings of the Asian Conference on Computer Vision , pages 1959– 1975, 2022.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models gscorecam: What objects is clip looking at? In Proceedings of the Asian Conference on Computer Vision , pages 1959– 1975, 2022

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:16.007733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:10.804418Z digest=sha256:28f662ff3cdf97a2de684212f0fc7948ba58856e75b7c1db8e3e77323373c5bd

Observation 1267bee9-71b5-4e73-b3b6-f43dbeeea742 · outbound

This paper cites Concept whitening for interpretable image recognition.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Concept whitening for interpretable image recognition

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:15.944748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:10.817714Z digest=sha256:ec9d2d5eaf240d1bb041427ce10964668cc109568f9da422b54926929be34bca

Observation af7352a7-ba5f-4632-a13f-533e7d74791e · outbound

This paper cites Blender - a 3D modelling and rendering package.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Blender - a 3D modelling and rendering package

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:15.901961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:10.826606Z digest=sha256:5276cba3fa2c988e4214536de5c9c6a28516f8e7e93a54e6059909de837329c7

Observation 736222fb-6438-4532-8d55-203299ba9e7f · outbound

This paper cites Vision transformers need registers.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Vision transformers need registers

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:15.855838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:10.834158Z digest=sha256:f1c2b88d2bb90314e2fc5329092dc1a69bc121c4b05ccf1215b92710e34ef2a1

Observation bf3f9988-9c2a-424f-a7df-7e787b8ec2ea · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models An image is worth 16x16 words: Transformers for image recognition at scale

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:15.779515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:10.851533Z digest=sha256:4930bf643fd5f48ad4711c8f63971b560567ae5c64b634d91a78e9222bd0329a

Observation 7ce08f67-a395-4f4a-99c9-1fc86e29db59 · outbound

This paper cites Towards a multimodal framework for remote sensing image change retrieval and captioning.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Towards a multimodal framework for remote sensing image change retrieval and captioning

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:40:12.372706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:10.857273Z digest=sha256:c4dc6dbe4249cb7dce239b0cdde7125ab0a2d32a6d10c9f535f4adc9b979d4a4

Observation c5e931e8-c616-4c4d-be32-c7d6d330c745 · outbound

This paper cites Interpretable explana- tions of black boxes by meaningful perturbation.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Interpretable explana- tions of black boxes by meaningful perturbation

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:15.696518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:10.865683Z digest=sha256:5caea5ea6a82cd4a92a21eb72f0490e9fc1a020fe3a64532ee5c3c333ea830f1

Observation fa8df934-cec8-4203-847e-8eb59d9ec776 · outbound

This paper cites Openagi: When llm meets domain experts.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Openagi: When llm meets domain experts

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:15.644489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:10.879502Z digest=sha256:886decb94e972842261d9ea8b8e03a9bc441007bb97b1c551944eb6d5b49ebc7

Observation 6c20fc50-e27a-419f-8cc9-e7e25106ee57 · outbound

This paper cites Clip4idc: Clip for image difference captioning.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Clip4idc: Clip for image difference captioning

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:15.589182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:10.888222Z digest=sha256:5a18d79e80b65e52b4106c7730d78bf34f0c9a385f1a0405527549a9ee7c04fc

Observation 10445064-b551-4a21-9e49-c0eac24d3331 · outbound

This paper cites Wrongfully arrested man sues Detroit police over false facial recognition match.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Wrongfully arrested man sues Detroit police over false facial recognition match

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:15.531575Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:10.900520Z digest=sha256:51e2457e2f6073cbb5d595c72759a2b51dc72e7bf726294bca64841485cf98c6

Observation b2e26d52-3ea7-47a0-98de-c919a85fc0f7 · outbound

This paper cites Which tokens to use? investi- gating token reduction in vision transformers.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Which tokens to use? investi- gating token reduction in vision transformers

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:15.417015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:10.906979Z digest=sha256:e7abd7ad8804291272f343caba70eb7a89de1e5f56c2c718ebf2b0ddbea86365

Observation 6f1d1728-df68-4686-93a9-841a3a600b0a · outbound

This paper cites Flawed Facial Recognition Leads To Arrest and Jail for New Jersey Man.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Flawed Facial Recognition Leads To Arrest and Jail for New Jersey Man

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:15.371263Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:10.913648Z digest=sha256:39abaea58451b2199b0b9c259711ae8804992a2655fb755a5e036c77975c1310

Observation 66abd07e-8b50-48e4-a3bf-457f03e1b8ab · outbound

This paper cites Change captioning: A new paradigm for multitem- poral remote sensing image analysis.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Change captioning: A new paradigm for multitem- poral remote sensing image analysis

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:15.325479Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:10.922351Z digest=sha256:8bc80cd49898538adb5927311a77424a05f52add269c1c8445cb986c319bbf36

Observation 45c94ae9-fc16-402b-8211-7254586c5c2d · outbound

This paper cites OneDiff: A Generalist Model for Image Difference Captioning.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models OneDiff: A Generalist Model for Image Difference Captioning

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T04:40:10.928996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:40:10.928996Z digest=sha256:c1b0f191d26e76a52d5c097e8284abc4834d774fddb7625fa4715498a780890f

Observation 5b022a48-c6e7-4b77-9dd2-bb0a1b0408fb · outbound

This paper cites Summers, and Yingying Zhu.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Summers, and Yingying Zhu

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:15.277133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:10.937360Z digest=sha256:12a3fe366c412093345c3ab46873519d73689805dc55b0551c7f4b9016633891

Observation 8937f13c-a5ff-45a9-8869-fd583805823e · outbound

This paper cites Learning to describe differences between pairs of similar images.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Learning to describe differences between pairs of similar images

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:15.242389Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:10.945699Z digest=sha256:f516df9f36f8a1c5a6906fe208b2d7339a47190dd7cadefc5d56600e73dcf2f1

Observation 2ef8ce1a-999d-44fb-bdec-d0ef11d773ad · outbound

This paper cites Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Img-Diff: Contrastive Data Synthesis for Multimodal Large Language Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T04:40:10.953813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:40:10.953813Z digest=sha256:7e1f47ec7d2389780d27371d975cf032dcc2a928a78eba283b2a4f7f9d43a8a9

Observation d64106d8-b0f3-4233-834b-4fe7c4677643 · outbound

This paper cites Now You See Me (CME): Concept-based Model Extraction.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Now You See Me (CME): Concept-based Model Extraction

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T04:40:10.963613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:40:10.963613Z digest=sha256:b367b9664e0648f5df14ef2c08c9634936f9745b3c5c793d48d6c4481766b2d1

Observation 7e1137a7-d478-4a28-8324-a5f2b80514da · outbound

This paper cites Vilt: Vision- and-language transformer without convolution or region su- pervision.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Vilt: Vision- and-language transformer without convolution or region su- pervision

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T04:40:10.995995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:40:10.995995Z digest=sha256:92e7576b6e9856fde7014b5799ad5aa27fd7cb5019d27c0951016a22e8968254

Observation acb038ef-8d9f-4e25-aa68-f565ace05d34 · outbound

This paper cites Concept bottleneck models.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Concept bottleneck models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:15.164075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.008316Z digest=sha256:f4336d592f67f9faa9979cb659ab11b4e7966a8a545606e8c41d4b22ff7b7d78

Observation b78ace51-0a8e-4d00-9630-b6f806d8f6b6 · outbound

This paper cites The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models The Open Images Dataset V4: Unified image classification, object detection, and visual relationship detection at scale

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T04:40:11.015862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:40:11.015862Z digest=sha256:de0cd792f148bd83a771a86f578b78a75c4b70b44bae54d503f18891a490dd71

Observation eb05402c-8232-4edd-9bf2-f24441ffc57b · outbound

This paper cites Dynamic graph enhanced contrastive learning for chest x-ray report generation.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Dynamic graph enhanced contrastive learning for chest x-ray report generation

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:15.096854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.034423Z digest=sha256:7b0c2f5303259ec75b14c08455bd36d4c6177898757a6d2676bd902418b680f4

Observation 70d59d8e-abcb-4198-b3eb-721d097a2876 · outbound

This paper cites Liang, C.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Liang, C

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:15.031847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.044375Z digest=sha256:d29dda2728d986f752da210a066afcfc7c393688a891d03e08f92baed391fbd6

Observation 257477fb-1220-4192-835b-30b774815ffb · outbound

This paper cites ROUGE: A package for automatic evaluation of summaries.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models ROUGE: A package for automatic evaluation of summaries

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:14.975150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.057886Z digest=sha256:941514c3491da89dbe7a9ef5bb6f1617b63e844e71dedd27856d98ebf0a86744

Observation 5d62f66c-5431-4e21-9d37-08600d125242 · outbound

This paper cites Improved baselines with visual instruction tuning.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Improved baselines with visual instruction tuning

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:14.927937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.066135Z digest=sha256:e703a18c50a0e07f69b1f86e4d3c6cbbc8e8c2b7ff40e1491de5f1ff589a8218

Observation 53a7c641-d834-4a9e-89ca-f778bfba4cf4 · outbound

This paper cites Interpretability Beyond Classification Output: Semantic Bottleneck Networks.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Interpretability Beyond Classification Output: Semantic Bottleneck Networks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T04:40:11.077011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:40:11.077011Z digest=sha256:9c92326aece37b642f802d454948ea5d2b3414dded61aaefdabf4ed1d0feb5c8

Observation a1cb517a-7d4d-4927-85b6-aa45e5449c18 · outbound

This paper cites SGDR: Stochastic gradi- ent descent with warm restarts.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models SGDR: Stochastic gradi- ent descent with warm restarts

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T04:40:11.084296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:40:11.084296Z digest=sha256:9a425a12c41a0464513eb5396a07fbc6ba9aa62e40e0ec5782078e8637a530dd

Observation 121a2486-b8b7-4d04-996d-0c0d557d5cd8 · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:14.841861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.093962Z digest=sha256:8e9a67b92e51c314550a1fa91d71dd0a3b3e91d46235d76d46d8e499077d8c54

Observation a788c71f-1001-40b8-913b-e3b3c820acae · outbound

This paper cites Spot the difference: Difference visual ques- tion answering with residual alignment.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Spot the difference: Difference visual ques- tion answering with residual alignment

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:14.791939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.100490Z digest=sha256:89d41299dec3a8479854b600116d51f746cdc1878d143d4cc8dd3e3020785c7c

Observation 020238cf-4847-4ad4-96e5-1f5fe5cbd9a5 · outbound

This paper cites Locating and editing factual associations in gpt.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Locating and editing factual associations in gpt

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:14.755785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.107821Z digest=sha256:8e70675d3c153cdc9f1b28988e1c6fa014d23c3183bb2e66b7ac58fdcea4b079

Observation dc535ae8-4efb-4fea-b512-aa420ea07542 · outbound

This paper cites Memory-based model editing at scale.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Memory-based model editing at scale

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:14.708439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.113373Z digest=sha256:1dea4db0c2336fef8e83fe328d33bc41304df665377fc3709624f3886d7b8aa4

Observation 0af686c2-2903-48ef-947b-f6676e1905a1 · outbound

This paper cites The ef- fectiveness of feature attribution methods and its correlation 11 with automatic evaluation scores.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models The ef- fectiveness of feature attribution methods and its correlation 11 with automatic evaluation scores

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:14.661838Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.126841Z digest=sha256:c0de0fc78b86c247017bfb92a5a8342be0777d9e63cd967052b4628a6f2c2abe

Observation f602c512-5c6d-4331-bf93-54a9fb05c6a3 · outbound

This paper cites Improving change detection by incor- porating correspondence information.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Improving change detection by incor- porating correspondence information

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:14.607287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.134529Z digest=sha256:6e6eb316dfe0364d056ace60bf838e13c89ecc4548cfff3b0e4fe88c6985ecb5

Observation 006609de-5b9d-41b5-855c-282ccaf6df6e · outbound

This paper cites How explainable are adversarially-robust CNNs?.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models How explainable are adversarially-robust CNNs?

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T04:40:11.141972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:40:11.141972Z digest=sha256:276b803fa17f127aaf5c2f5013637ae7dea2ff15853d30f4229e7883c8657744

Observation 04f704ac-538f-45c5-913c-a5ce1b336319 · outbound

This paper cites an unresolved cited work.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:40:14.567033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.150674Z digest=sha256:a440214d14ae94936a69b30d06a954037bd760856f154aac5fe0cd1801c06de6

Observation a2cc38e9-c70d-4f94-9025-937ff13cd304 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Bleu: a method for automatic evaluation of machine translation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:14.511660Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.158240Z digest=sha256:675b3419477f4ed6c72759e588d38be10280b3e5de5db74fe5cc6076b3e005d9

Observation 69ca58c8-6fa3-436f-a35f-e3c9c5b7f636 · outbound

This paper cites Robust change captioning.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Robust change captioning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T04:40:11.175648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:40:11.175648Z digest=sha256:9ac3fb6835766d201860ce1ea321ce596c0b40ff880c5b7c51ae5fd197623894

Observation 3369db3a-15af-4f37-8cbd-017abb361e0c · outbound

This paper cites PEEB: Part-based image clas- sifiers with an explainable and editable language bottleneck.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models PEEB: Part-based image clas- sifiers with an explainable and editable language bottleneck

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:14.334751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.197195Z digest=sha256:47107d14562ff4e0917bef52266a3e7365dd4570cabae6cb82649cf3955df36c

Observation c651e62f-0f01-40e2-aaf2-328d54eda638 · outbound

This paper cites Deepface-emd: Re-ranking using patch-wise earth mover’s distance improves out-of- distribution face identification.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Deepface-emd: Re-ranking using patch-wise earth mover’s distance improves out-of- distribution face identification

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:14.237969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.228326Z digest=sha256:ffbbf0548a6e6c4d1862cc733640b62a09acd35769f70d6e2183d1b8f020ec1f

Observation fcc54815-77c0-48e6-9886-1897cb300b12 · outbound

This paper cites Fast and interpretable face identification for out-of- distribution data using vision transformers.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Fast and interpretable face identification for out-of- distribution data using vision transformers

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:14.201490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.249141Z digest=sha256:1ef7e770f14888edf4369225139e23a600c775e05dc28c1ccdfecc6ef6150c95

Observation d099e3ea-c21e-4983-a3de-738ec096d427 · outbound

This paper cites Describing and localizing multiple changes with transform- ers.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Describing and localizing multiple changes with transform- ers

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:14.144746Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.262805Z digest=sha256:14b463a8212ed67fc90e46c0a89216c61cacf014573356d577ae7e773806fd1b

Observation a575f53f-daeb-4c8f-a819-4e6801a15dcd · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Learning transferable visual models from natural language supervi- sion

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-11T04:40:11.268842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:40:11.268842Z digest=sha256:3bb82643328d6c9da3cc6060a71cb0f3f6311b867333b9f5359c84474110bd78

Observation 455668d7-94eb-4261-be5f-09a846fc08ae · outbound

This paper cites The new lawsuit that shows facial recog- nition is officially a civil rights issue.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models The new lawsuit that shows facial recog- nition is officially a civil rights issue

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:14.035972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.281773Z digest=sha256:7e95975bea5f1d91a24b178267e147c2542106807823773b2433318efb20ff4a

Observation 6ffd5800-7f6c-449c-96e5-b67d3897a2f4 · outbound

This paper cites The change you want to see.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models The change you want to see

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-11T04:40:11.292494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:40:11.292494Z digest=sha256:fa314b05189ac2f7abafb31b5634f7566df7b98ddfb61722e29aa27ff50c8ab7

Observation c99281f6-7dd6-4ec6-b84d-d9af97f4911b · outbound

This paper cites Explaining deep neural networks and beyond: A review of methods and applications.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Explaining deep neural networks and beyond: A review of methods and applications

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:13.970533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.324751Z digest=sha256:e14d2139b9a5ecbb26e24aa96d1e9f94f687e9a50542e92eb1b4e4379d06ea14

Observation 7767697d-892a-4b99-8615-ab0876f80a16 · outbound

This paper cites The pandemic is testing the limits of face recogni- tion.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models The pandemic is testing the limits of face recogni- tion

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:13.917424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.338720Z digest=sha256:a687ce7273253df48b97b5a7d1767facc7063bbab82260b91dfe6ecd1e122f20

Observation 4fd33787-fdc1-4522-9c73-43ca0576e1c4 · outbound

This paper cites Reclip: A strong zero-shot baseline for referring expression compre- hension.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Reclip: A strong zero-shot baseline for referring expression compre- hension

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:13.874300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.356976Z digest=sha256:6151f2321a1a64a5208bc1236c4e552ed1a92e3b20b81d9caffe35c141288662

Observation 25100f87-6b2a-4528-a6f4-e83acbbe0c03 · outbound

This paper cites The stvchrono dataset: Towards continuous change recognition in time.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models The stvchrono dataset: Towards continuous change recognition in time

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:13.804341Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.394755Z digest=sha256:e2609df181ca59e6406b3f055d67ce8b8ee63be09b8ea72b3531acd5e9c17fa4

Observation 4a825cc5-9115-48bf-8852-d3b040e7d297 · outbound

This paper cites Resolution-robust large mask inpainting with fourier convolutions.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Resolution-robust large mask inpainting with fourier convolutions

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:13.740485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.408220Z digest=sha256:44e5380e82c714fde03fb0d026da000f93be7fbf4ff128be5de85b9574d598d3

Observation 8c1f87f8-280c-4bdc-81c4-fe8e5ed133f8 · outbound

This paper cites Expressing visual relationships via language.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Expressing visual relationships via language

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:13.678333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.424751Z digest=sha256:be932b795d610101f51faa904f2c1001d7dc8d628d321f4d04ff61fa0b8d9326

Observation 3aaf7735-ca9e-435b-836a-048e6344c70e · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:13.623663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.567674Z digest=sha256:7c338a2cb37eb9f86f72b4afc8c394a4aa4815d4271ba5830571d244e2c7c2a6

Observation 0f2f387e-adac-44f4-8843-d400cb5a0b6d · outbound

This paper cites Attention is all you need.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Attention is all you need

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:13.455191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.674750Z digest=sha256:752618f24ba8a780bada436f4f183532756267093d21d8406a1a84e65ba61f5d

Observation 8335ec1d-9c70-4ca6-a125-62c81907d01a · outbound

This paper cites Cider: Consensus-based image description evalua- tion.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Cider: Consensus-based image description evalua- tion

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:13.294760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.768687Z digest=sha256:b299cb04b0983e63f6f36015930d68e680a25f4a2513770f8d13bc07045dbee1

Observation e397a496-d379-4258-8df5-f49192b60329 · outbound

This paper cites Learning bottleneck concepts in image classification.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Learning bottleneck concepts in image classification

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:13.214750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.794749Z digest=sha256:17b3edc69665ea8487926884512cdd2801899e7adb376bc4b0261d5eff903fca

Observation a3dd02bc-8e3e-41ef-909c-d3d422e36bf5 · outbound

This paper cites Co-attention for conditioned image matching.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Co-attention for conditioned image matching

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:13.124760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.810533Z digest=sha256:446c5e854768d0893e9ffc8c65d78613499f2e81f511d10fea6b992a39fe5510

Observation d5f7cd9f-531a-4181-9fb7-a3ae3bfa65dc · outbound

This paper cites L2C: Describing visual differences needs semantic under- 12 standing of individuals.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models L2C: Describing visual differences needs semantic under- 12 standing of individuals

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:13.033000Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.824755Z digest=sha256:66aac1d09ce9ccdb7d92425bd70a267f8ba55c0c59a8ac60de3a5670a4c2efe7

Observation e72e372e-33ec-4ddf-9756-728d21cf47d5 · outbound

This paper cites Image difference captioning with pre-training and contrastive learning.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Image difference captioning with pre-training and contrastive learning

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:12.974961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.842641Z digest=sha256:935b96fa1fb47ba2232cc5b461cde0670db0972c342679f8d6fe2a54ec0fb1b6

Observation 380f9dac-d355-4e39-a4b4-e42913e5fef6 · outbound

This paper cites Scaling vision transformers.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Scaling vision transformers

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-11T04:40:11.850919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:40:11.850919Z digest=sha256:78fbea7f24fd7322bee3ce8187e59509af2e3cbf8f2904f41e0ac057edf5b13b

Observation fd6d6b9c-1bba-4c43-aac2-e3c8e67302c6 · outbound

This paper cites Top-down neu- ral attention by excitation backprop.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Top-down neu- ral attention by excitation backprop

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:12.891960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.881388Z digest=sha256:8085ebf9ef4de78123f3328712a23bb134e305d162432f117a04b28e0ecc128d

Observation 7c33fa14-d549-4911-9c57-9ed1856c1fe4 · outbound

This paper cites BERTScore: Evaluating Text Generation with BERT.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models BERTScore: Evaluating Text Generation with BERT

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-11T04:40:11.898842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:40:11.898842Z digest=sha256:a9d1740ac4542770183778cbe2aa6a3393bba5e692b9d662f32bef4bfdcb9f64

Observation 6fb05bac-2db6-426f-8c98-0f79bd294d30 · outbound

This paper cites Learning deep features for discrimina- tive localization.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Learning deep features for discrimina- tive localization

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-11T04:40:11.912157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:40:11.912157Z digest=sha256:7c01d645217d8bf71a0052b9697c29869885f132bd1b8d675a3ff03eddb2b92d

Observation 06b2ccbf-3274-4bc7-bc81-c6a38ecabb64 · outbound

This paper cites Implementation details A.1.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Implementation details A.1

Reference 2016

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:40:12.784863Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.919275Z digest=sha256:82fd1c897169fbcd469e9c56aa864590899805f4cf340ae7d8666349cb9e5404

Observation 5a9f7c76-530c-4161-89d7-1f81053b20cc · outbound

This paper cites an unresolved cited work.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Unresolved cited work

Reference 2019

Resolution
parse uncertain
raw_fallback, observed 2026-08-11T04:40:14.404282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T04:40:11.187726Z digest=sha256:74d7d65dd01fa8fa7e60025dc9cf662d0bd488bd38018dc6344fdfa44ec7dfbc

Observation 95fc221d-0d70-49d9-b8ec-01e9cebb77ec · outbound

This paper cites an unresolved cited work.

TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models Unresolved cited work

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-11T04:40:11.214913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:40:11.214913Z digest=sha256:82572eb3949ac75ab95f5821561f9883ba25d0eb25e4875f7934db39b2605d6f

Pith citing papers

Observation f5079130-d1f2-4118-a6e1-627332b5c157 · inbound

Legible-by-Construction: Attention and End-to-End Transformers cites this paper.

Legible-by-Construction: Attention and End-to-End Transformers TAB: Transformer Attention Bottlenecks enable User Intervention and Debugging in Vision-Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-11T20:06:55.102341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-11T20:06:55.102341Z digest=sha256:d77176554c09fea940dd5e25b260b30a74837fbf38cbfb58c4aa271e520ad53c