Pith. sign in

Paper Citation Record · LEDGER

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features

As of 9 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 12 inbound Pith citation observations for arXiv:2502.04320.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.04320 v2

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-08T22:51:30.517584Z

measured 64 of 64 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:41:21.200157Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

52 of 52 outbound references displayed

  • verified exact5
  • verified fuzzy0
  • unresolved46
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation 8c7919ed-079b-461e-8aa8-f182a905f229 · outbound

This paper cites Quantifying Attention Flow in Transformers.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Quantifying Attention Flow in Transformers

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.333740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.333740Z digest=sha256:f04048268f0f4f1627d0b8b9d47ff9d824b5cc43df86febb13ecb348ea5d9e19

Observation a83e9e0a-cadd-41c2-881e-f837daf8404b · outbound

This paper cites SegDiff: Image Segmentation with Diffusion Probabilistic Models.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features SegDiff: Image Segmentation with Diffusion Probabilistic Models

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.338470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.338470Z digest=sha256:701608cb0d74a8a56a45ea28e809d55eaf5852eb8f63981b206d15ec6a381ecb

Observation 770be66a-b762-4493-8a3d-cb35658fbf49 · outbound

This paper cites Label-Efficient Semantic Segmentation with Diffusion Models.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Label-Efficient Semantic Segmentation with Diffusion Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.342260Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.342260Z digest=sha256:d99970d191f833443868b85c78488dbcc9239921b23ae71afa78af52f5a2762f

Observation 1e2458e1-311e-4e71-bfce-0e29cbd9e825 · outbound

This paper cites Layer-wise Relevance Propagation for Neural Networks with Local Renormalization Layers.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Layer-wise Relevance Propagation for Neural Networks with Local Renormalization Layers

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.346022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.346022Z digest=sha256:ce30a6258b2973557cd8a48f2192a4c1b8bb3a187342157648951ca6553bde2f

Observation 609d7847-06fe-45fd-81f0-198bdca9e3db · outbound

This paper cites InstructPix2Pix: Learning to Follow Image Editing Instructions.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features InstructPix2Pix: Learning to Follow Image Editing Instructions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.349907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.349907Z digest=sha256:5e35664f1cdb096fe26b8723631ecdfe29a0d3b0255209b710824427533dc016

Observation 9c2b28d1-55b2-4abb-a2b3-0ba950d533b6 · outbound

This paper cites Extracting Training Data from Diffusion Models.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Extracting Training Data from Diffusion Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.353679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.353679Z digest=sha256:c4665a8bf7cd05c9dfcc029de251cb26c31b12fed4df989bfb74217e53429a8f

Observation 72881085-ab39-4e6d-88f9-d7d1d4a4d515 · outbound

This paper cites Emerging Properties in Self-Supervised Vision Transformers.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Emerging Properties in Self-Supervised Vision Transformers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.357789Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.357789Z digest=sha256:16692be7c40671ddd1c95e234a4918a05a2d976e375b73ccc18042036bfbcfdd

Observation f2c4b612-51a8-4d5b-9205-10bc526d8363 · outbound

This paper cites Transformer Interpretability Beyond Attention Visualization.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Transformer Interpretability Beyond Attention Visualization

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.361375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.361375Z digest=sha256:c4f5d1a013d0c98cb71005164c66d0570e75c943ebf848f7e20b594c565b557e

Observation 738c633f-5567-49ac-84ac-e2c21dfce76b · outbound

This paper cites Attend-and- Excite : Attention - Based Semantic Guidance for Text -to- Image Diffusion Models.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Attend-and- Excite : Attention - Based Semantic Guidance for Text -to- Image Diffusion Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.365200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.365200Z digest=sha256:d271158e382eb97ca8a80565b91f1ee3bb0cab560b4c63bc0881d153213429e0

Observation d2a8940a-b9c6-40bf-b148-c1de63cc00c1 · outbound

This paper cites Training-Free Layout Control with Cross-Attention Guidance.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Training-Free Layout Control with Cross-Attention Guidance

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.368808Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.368808Z digest=sha256:27050d35d50bc9f2472c09b3cc8c6b032b6b7810375742ff1be6e85d8ce34abc

Observation 9aaf1493-47d9-4307-b7be-1fcfaa83e78e · outbound

This paper cites PiCIE: Unsupervised Semantic Segmentation using Invariance and Equivariance in Clustering.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features PiCIE: Unsupervised Semantic Segmentation using Invariance and Equivariance in Clustering

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-08-08T22:51:31.022701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:51:30.372406Z digest=sha256:d6b3390a3d39fef7d3f4dd6b446f6d92ada4b7f1185ef555416d7011bc946b8b

Observation 4457faec-e223-4efd-b871-6533c22a2f1e · outbound

This paper cites FluxSpace: Disentangled Semantic Editing in Rectified Flow Transformers.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features FluxSpace: Disentangled Semantic Editing in Rectified Flow Transformers

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.375916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.375916Z digest=sha256:acc26919de66b7458e085e3f58ba8b84beac03277ed7c866568dc0f9b4cc96b6

Observation dd4542d8-af62-4912-92a8-5df459594db5 · outbound

This paper cites Vision Transformers Need Registers.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Vision Transformers Need Registers

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.379390Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.379390Z digest=sha256:0a6c8b1a918d43058a235aa400f69eb67bdb5148ec1ba418bcd34de4cd58fcc4

Observation d7e74634-b778-4be1-b99f-f4391ec3d053 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.383170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.383170Z digest=sha256:7b3177c86201e6400aac5007afd573704e5d776dd73fe22fc35c66222b844584

Observation 0a07c957-12c5-4a2d-9ef3-c1161b607ab5 · outbound

This paper cites Diffusion Self-Guidance for Controllable Image Generation.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Diffusion Self-Guidance for Controllable Image Generation

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-08T22:51:30.982597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:51:30.386761Z digest=sha256:713385f8d80c6375191f8dc8379b764ad93c58db475beb74bea2a477ce9516e0

Observation 40f0e65d-fc3d-4050-bbba-995415c4e134 · outbound

This paper cites Scaling Rectified Flow Transformers for High-Resolution Image Synthesis.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Scaling Rectified Flow Transformers for High-Resolution Image Synthesis

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.390325Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.390325Z digest=sha256:0f049bdf4d6b90d2bfac508f5ed467e437f9c5aec23e723e093060eba6d8c1f0

Observation e950b0dd-ccd9-4dbf-a0a4-cff94d7573cb · outbound

This paper cites an unresolved cited work.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.394346Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.394346Z digest=sha256:afb17d26bd91ccca1763688b11fabd249464fefa545c88882f93a163e5cc2c44

Observation defb6842-ee33-495f-bcea-0fa5f2286cda · outbound

This paper cites Interpreting CLIP's Image Representation via Text-Based Decomposition.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Interpreting CLIP's Image Representation via Text-Based Decomposition

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.398000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.398000Z digest=sha256:e4516c4df3f920e4466a6ba623421ece237d1c0dc6de309c39feb4ecd4a739d4

Observation 13b9f209-b716-4b72-b900-0eef2cc62d94 · outbound

This paper cites ImageNet Auto - Annotation with Segmentation Propagation.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features ImageNet Auto - Annotation with Segmentation Propagation

Reference 19

Resolution
verified exact
doi, observed 2026-08-08T22:51:30.558819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:51:30.401686Z digest=sha256:4547a9c2d186f4f4903e23a8a52cef8827aea0bdd6631e5aec8d1e09119a7637

Observation 856f6a24-3d80-4fa8-a1ca-a9bdc88335d6 · outbound

This paper cites Pre-trained Text-to-Image Diffusion Models Are Versatile Representation Learners for Control.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Pre-trained Text-to-Image Diffusion Models Are Versatile Representation Learners for Control

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.404867Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.404867Z digest=sha256:c4e370176dad5947084d0c78cf4d2f7fd27d3ffca1e5b3635e99c0872c9087fb

Observation ca5d20cf-7f76-4232-b4b3-32e816d04992 · outbound

This paper cites Unsupervised Semantic Segmentation by Distilling Feature Correspondences.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Unsupervised Semantic Segmentation by Distilling Feature Correspondences

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.408783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.408783Z digest=sha256:5b9ea077443a3d7e293c24e987c6a4c7879b9f63713e59759668f7a981b6b779

Observation 775f2ebe-a2fd-41ae-9611-4341a5bf3584 · outbound

This paper cites Prompt-to-Prompt Image Editing with Cross Attention Control.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Prompt-to-Prompt Image Editing with Cross Attention Control

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.412279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.412279Z digest=sha256:c9495ba5d23c53fcad11735b22a2cca660246482b672f3e3d51bc9f8b21ef312

Observation a29b1f14-7036-4d01-9d5f-ed42a0faf9e5 · outbound

This paper cites Generalization in diffusion models arises from geometry-adaptive harmonic representations.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Generalization in diffusion models arises from geometry-adaptive harmonic representations

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.416092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.416092Z digest=sha256:c5b62991c6b04153d937e2281cf3ca653d31aa7803686b3ef072f583bf0489af

Observation a5dc0093-a2ad-4a58-a1d8-449fc0adde92 · outbound

This paper cites Diffusion Models for Open-Vocabulary Segmentation.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Diffusion Models for Open-Vocabulary Segmentation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.419474Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.419474Z digest=sha256:a481b1a9dfe819022018ea2f00ed928bd40853473cc6ed15eab02812e1dd5184

Observation 8c4fd63f-497c-4c19-8ecf-7a7679d91aa7 · outbound

This paper cites Segment Anything.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Segment Anything

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.423270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.423270Z digest=sha256:62cc8dcf3e57cf1a456b7a477da099e39b6c01e3f687cba4102ba59f3799f41c

Observation 645504f3-664f-4a1d-b8bb-024d27065370 · outbound

This paper cites an unresolved cited work.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.426778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.426778Z digest=sha256:0ec321988e678c1d8350c203b7b8134f6f5f1d891ced9246294a72140ac4a79b

Observation 5a2c4880-61f3-4e11-96be-4b05bf4f886c · outbound

This paper cites an unresolved cited work.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-08T22:51:31.129147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:51:30.430207Z digest=sha256:cd8570b18cb6555d8e315b9445425b058b1c467320259306b50fc5250c238182

Observation 0f6928a4-2f4a-452e-a257-6eebbbc4badb · outbound

This paper cites Your Diffusion Model is Secretly a Zero-Shot Classifier.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Your Diffusion Model is Secretly a Zero-Shot Classifier

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.433654Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.433654Z digest=sha256:606c2e1ce75362bbb220139c39d411331b2ed370fac05b75b1bc71cb9211cf70

Observation 0c117fd9-4b46-4f8f-922a-4e389e6440d6 · outbound

This paper cites Open-vocabulary Object Segmentation with Diffusion Models.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Open-vocabulary Object Segmentation with Diffusion Models

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-08T22:51:30.776111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:51:30.437041Z digest=sha256:40ca00471105ff526f1132c808c3d58db4756256a062ec601bbafc62e0f0cc3c

Observation 3295af62-2c4f-4b70-a679-68a0e565a74b · outbound

This paper cites Flow Matching for Generative Modeling.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Flow Matching for Generative Modeling

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.440588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.440588Z digest=sha256:72caf1cd1b7697bfd000bbbc63096b3f662163a1eeabe238fc23a189738ae942

Observation 78f599a0-771b-449e-bcc4-4f8fb1a48f46 · outbound

This paper cites Towards Understanding Cross and Self-Attention in Stable Diffusion for Text-Guided Image Editing.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Towards Understanding Cross and Self-Attention in Stable Diffusion for Text-Guided Image Editing

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.443951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.443951Z digest=sha256:2066519dc5b4bae9c557e324a3f147b781c1e29be626a066251d19fd71dbbfcc

Observation 4b6825cf-7455-447f-b22b-a9eebb838bee · outbound

This paper cites Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Flow Straight and Fast: Learning to Generate and Transfer Data with Rectified Flow

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.447239Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.447239Z digest=sha256:9af2bad837164029ed856f8a13413b085de941698664a11fef1fd6111a643fd2

Observation 5e4174c8-2611-4510-bd2f-5d3f579344b5 · outbound

This paper cites Open-Vocabulary Attention Maps with Token Optimization for Semantic Segmentation in Diffusion Models.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Open-Vocabulary Attention Maps with Token Optimization for Semantic Segmentation in Diffusion Models

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-08T22:51:30.736499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:51:30.450874Z digest=sha256:3771d4e7ca0e8b3f3d94b6250360691df63c39b0fe77c6ffda15fe7d3df34906

Observation 6210da07-c566-4922-b569-c40f2a9bfddf · outbound

This paper cites an unresolved cited work.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-08T22:51:31.118945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:51:30.454246Z digest=sha256:221bdfa167d739d98dfc6655ab504dd2848f0c7c87941899ef536362d5b634c7

Observation 9392d588-5623-4982-a831-ff44d36e8ef6 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features DINOv2: Learning Robust Visual Features without Supervision

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.457401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.457401Z digest=sha256:87cd4795f3f6fa03dc9541f23a6883528db5016b4cfd6ed7bc0babcfdf5653a8

Observation dca7f9e3-9348-49bf-b6b1-22f37c4020c2 · outbound

This paper cites PyTorch: An Imperative Style, High-Performance Deep Learning Library.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features PyTorch: An Imperative Style, High-Performance Deep Learning Library

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.460316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.460316Z digest=sha256:cb379ad850f75921e6c0bf6a5f21770cfba138b08052f9f8e554ecaba5a0a2cf

Observation a225cf19-2489-4d88-96a7-706d4a8b8889 · outbound

This paper cites Scalable Diffusion Models with Transformers.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Scalable Diffusion Models with Transformers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.463663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.463663Z digest=sha256:604decbba8eaceb196a5e12df762d985f0ab33a2eee3f3b2f2e712a7097a7b58

Observation 1d7beaa9-8877-496a-9198-0f0a0de79938 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.467066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.467066Z digest=sha256:cadee0e7bc710ed4229dbcf6d202f4a7a9b3417bbd90d74441c529d1ce9a39f2

Observation 48e90fb0-27ef-49c6-9ee5-f05c1543f8a8 · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Learning Transferable Visual Models From Natural Language Supervision

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.470288Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.470288Z digest=sha256:41073ac570db41e1ccce818d139efe9e8cede4e8b6fb4d45c20a67ae3f5e161f

Observation c2c5a252-5706-4faa-b942-8c70d3f1b984 · outbound

This paper cites Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.473563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.473563Z digest=sha256:6968fcb09fe468f8d1410d4c9602ade7ca43cbf505d067753fdf40ce8f6bdb5f

Observation a25eb93c-c981-427d-bd0d-f44de1c09b83 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features SAM 2: Segment Anything in Images and Videos

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.477146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.477146Z digest=sha256:de4fc54b9584adc99c33db865286f69c2278a854cf7b4b0b66ccbcd06df8fe41

Observation fee83678-a421-4231-aa04-1c2c38c2ee7f · outbound

This paper cites High-Resolution Image Synthesis with Latent Diffusion Models.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features High-Resolution Image Synthesis with Latent Diffusion Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.480460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.480460Z digest=sha256:22730c1b989102f6210109583eca4a23f39af7dff38d0b71fcd511877c7f63fb

Observation eaad2de4-2aa0-4853-9908-81c32c144cb7 · outbound

This paper cites U-Net: Convolutional Networks for Biomedical Image Segmentation.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features U-Net: Convolutional Networks for Biomedical Image Segmentation

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.483665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.483665Z digest=sha256:98d71f145c5ba93b12051d111be79fab07b823d714943aa5de54a15b5f0a48ab

Observation 0d980ec8-17a9-4a44-8cce-1d2e12b4889a · outbound

This paper cites Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.490389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.490389Z digest=sha256:0dc7e8da3e0d0883695596fc0bf3648cb456f04d2ebf4bad8cd6abe609cc068d

Observation d0b2afb8-8fe5-417b-9f19-cb62552bbb8a · outbound

This paper cites CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features CLIP as RNN: Segment Countless Visual Concepts without Training Endeavor

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-08-08T22:51:30.636567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-08-08T22:51:30.493433Z digest=sha256:a26e3973b22906635e5e17bfb31c7a637287349308ec2f571d0d73e7fb56c63b

Observation 7a90582c-664f-43de-ad20-1304c543abfa · outbound

This paper cites What the DAAM: Interpreting Stable Diffusion Using Cross Attention.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features What the DAAM: Interpreting Stable Diffusion Using Cross Attention

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.496805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.496805Z digest=sha256:3404231ef899969fbe07ccea4bd9bceb6139f0fff7811a3c767e5c20e35ab79f

Observation 927927ff-6866-4d31-b25a-ca8ed11d3295 · outbound

This paper cites Diffuse, Attend, and Segment: Unsupervised Zero-Shot Segmentation using Stable Diffusion.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Diffuse, Attend, and Segment: Unsupervised Zero-Shot Segmentation using Stable Diffusion

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.500093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.500093Z digest=sha256:ea7359044ebc57ee18c0c7f0fed1821d56529365a2b89d560976caba80e0450f

Observation 18f6e557-2c90-4cb2-a9cc-9172a5c63a5c · outbound

This paper cites Breaking the Curse of Dimensionality: Diffusion Models Efficiently Learn Low-Dimensional Distributions.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Breaking the Curse of Dimensionality: Diffusion Models Efficiently Learn Low-Dimensional Distributions

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.503757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.503757Z digest=sha256:609eb54557e20c2d237eb6d7dadcfaff74dd580d2e6ac52cbc3cebc1f6eeb4b6

Observation 3f68ee24-48f7-4099-a447-b85ad5b40753 · outbound

This paper cites Understanding and Improving Layer Normalization.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features Understanding and Improving Layer Normalization

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.507227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.507227Z digest=sha256:fa4f92c5edd9e5817fcad16edbcd37c0fa103a0651659ce4ad766a5702715932

Observation 90877743-4f2b-46af-b2b3-28f9ab064aa4 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.510651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.510651Z digest=sha256:af9c0a0926d03684d17b35c813291c6c019835b55060eda709b4bd3094910668

Observation e2fc2cdf-b61b-4a91-bdb2-6cc16926b791 · outbound

This paper cites iBOT: Image BERT Pre-Training with Online Tokenizer.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features iBOT: Image BERT Pre-Training with Online Tokenizer

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.514096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.514096Z digest=sha256:a1c156572b4d9f6295dc251285dfe2bc9e1d832a26454d50ab284adcb5ba7223

Observation 8a00a806-1da0-4a55-98cd-a82587431609 · outbound

This paper cites write newline.

ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features write newline

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-08T22:51:30.517584Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-08T22:51:30.517584Z digest=sha256:7fb0f1c2f7cc2e2cd3a05b3b76449ac5083ed63c6e433dd7ad30417d0591fb64

Pith citing papers

Observation 233f2121-ba18-4444-9534-6a742e222961 · inbound

LoRAShop: Training-Free Multi-Concept Image Generation and Editing with Rectified Flow Transformers cites this paper.

LoRAShop: Training-Free Multi-Concept Image Generation and Editing with Rectified Flow Transformers ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T12:41:21.200157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T12:41:21.200157Z digest=sha256:c6c9c322fd202bd1956cc3ab236f21e8d716d5921e75db8d862b8614595950e2

Observation e5d4b584-b4d5-4ec3-9d37-32aeef6aa1bb · inbound

From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models cites this paper.

From Generation to Generalization: Emergent Few-Shot Learning in Video Diffusion Models ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:31.024475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:31.024475Z digest=sha256:42825028ad026aca4eff2aefdf5e706f79dcf8c8c20e0579989b6c8decc9859b

Observation e373986a-7abc-4b74-a3b0-20d237a9dc01 · inbound

The Cow of Rembrandt - Analyzing Artistic Prompt Interpretation in Text-to-Image Models cites this paper.

The Cow of Rembrandt - Analyzing Artistic Prompt Interpretation in Text-to-Image Models ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-21T23:44:26.945529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-21T23:41:50.534434Z digest=sha256:ab5a16dab9a67f41d48155a8a413fd89a8b499a470d516bc6d8c72d34633be81

Observation 8e151f58-cd7c-4cf5-925e-a8d75f068994 · inbound

Attention of a Kiss: Exploring Attention Maps in Video Diffusion for XAIxArts cites this paper.

Attention of a Kiss: Exploring Attention Maps in Video Diffusion for XAIxArts ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-05T13:27:41.569139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T13:27:41.569139Z digest=sha256:5748268c705905c3d6174275755e3f7b5974a48f5c17e81eaa767c2b2a4b9d7b

Observation 286c8497-c10b-40f7-8d9c-569fb466ffb7 · inbound

S3OD: Towards Generalizable Salient Object Detection with Synthetic Data cites this paper.

S3OD: Towards Generalizable Salient Object Detection with Synthetic Data ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-04T08:19:59.521627Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T08:19:59.521627Z digest=sha256:fc37aec49a441800f0cf4e1af95d643fdc2e1424c9d89e0816d577b79e6d65ed

Observation 014dd85b-2817-4cc2-be1c-e2a8fccbdcff · inbound

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation cites this paper.

TS-Attn: Temporal-wise Separable Attention for Multi-Event Video Generation ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-10T02:53:29.908134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T02:45:10.577070Z digest=sha256:710c3f3e12b0edca7d0eba3f9d3c7770cdfe7a0b5ef80166efd02f073ea6d191

Observation fe15ad0f-f369-4514-aa9c-f95023a90daa · inbound

FAIR_XAI: Improving Multimodal Foundation Model Fairness via Explainability for Wellbeing Assessment cites this paper.

FAIR_XAI: Improving Multimodal Foundation Model Fairness via Explainability for Wellbeing Assessment ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:16:12.253353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T06:16:59.725843Z digest=sha256:95ae7baaf0563533fa288758dd9b248f829dda97f07c656f27682b17ff907333

Observation 26316507-ecd0-4158-ae13-55a10d372608 · inbound

FAIR_XAI: Improving Multimodal Foundation Model Fairness via Explainability for Wellbeing Assessment cites this paper.

FAIR_XAI: Improving Multimodal Foundation Model Fairness via Explainability for Wellbeing Assessment ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-02T15:32:58.066814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T15:32:58.066814Z digest=sha256:c671f98de908a1d578e0620dfbf89453aee4b291d61ca043f9dec5003ef625db

Observation 1e0e87f6-6282-4654-a044-369e679d2a9a · inbound

Consistency Regularised Gradient Flows for Inverse Problems cites this paper.

Consistency Regularised Gradient Flows for Inverse Problems ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features

Reference 66

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T03:25:55.059096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-11T03:21:35.082352Z digest=sha256:3a104bfed96f105a26b951793c4e2154c374bb769c5577a09366bf26c2782d92

Observation b53677f1-09e8-4e1f-b252-47ec5ad2bebb · inbound

What Concepts Lie Within? Detecting and Suppressing Risky Content in Diffusion Transformers cites this paper.

What Concepts Lie Within? Detecting and Suppressing Risky Content in Diffusion Transformers ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-12T06:06:26.433353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T04:35:39.370969Z digest=sha256:dc0a862a26e2af5fac9ff871d76edca219bbc54199dc256cee4a2633923340dc

Observation 794bcc62-c9f1-4579-80ab-8a00799ac7e8 · inbound

Unified Safe In-context Image Generation in Multimodal Diffusion Transformers via Restricting Unsafe Information Flows cites this paper.

Unified Safe In-context Image Generation in Multimodal Diffusion Transformers via Restricting Unsafe Information Flows ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-07-02T16:07:09.083615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T22:58:44.478077Z digest=sha256:38f8f6a3a84b5d51b39e35920df0edb0a8cbbc80135d63633532f8b1b1621ba1

Observation 1c2eb649-ccb4-42b1-a613-5a0adcff069e · inbound

Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers cites this paper.

Text Template Tokens Are Implicit Semantic Registers in Diffusion Transformers ConceptAttention: Diffusion Transformers Learn Highly Interpretable Features

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T13:25:54.283706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T13:25:54.283706Z digest=sha256:8201d5e033d931b9badba2a8af1c2c7f56ffe68da55cddfc34240a160416143e