Pith. sign in

Paper Citation Record · LEDGER

Unified Local and Global Attention Interaction Modeling for Vision Transformers

As of 14 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2412.18778.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.18778 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:34:02.333517Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact4
  • verified fuzzy2
  • unresolved48
  • parse uncertain0
  • malformed identifier3
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation dfca2a14-6f44-4723-92be-7475724edc22 · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:34:03.622632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T04:34:02.075673Z digest=sha256:8bb3cff6e02579123750d8a592d665a418d52ac7553e520dfec8791700d2d21e

Observation d5f60d4d-aebd-49e7-a3c1-2f604f23040e · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.080829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.080829Z digest=sha256:df3fec5d4fdd86170eb746b40bf21443d837f8544451ff5c139df36bae819b69

Observation 3979f508-145e-4e9d-bb31-df8be4b80109 · outbound

This paper cites End-to-End Object Detection with Transformers.

Unified Local and Global Attention Interaction Modeling for Vision Transformers End-to-End Object Detection with Transformers

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.085201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.085201Z digest=sha256:9090931433b17f66defb8e459ba0aeff3de6cb371a85176f28afe3a0075f1b64

Observation 0a088262-0ecd-4887-b30e-553124069d4f · outbound

This paper cites SAM2-Adapter: Evaluating & Adapting Segment Anything 2 in Downstream Tasks: Camouflage, Shadow, Medical Image Segmentation, and More.

Unified Local and Global Attention Interaction Modeling for Vision Transformers SAM2-Adapter: Evaluating & Adapting Segment Anything 2 in Downstream Tasks: Camouflage, Shadow, Medical Image Segmentation, and More

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.090460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.090460Z digest=sha256:df551c0e21aaaca5ac84151bd34741bebb2b51c9baab488dffad685d23c7bea5

Observation 48f98527-4342-4056-9503-cfcd32e7f617 · outbound

This paper cites SAM Fails to Segment Anything? -- SAM-Adapter: Adapting SAM in Underperformed Scenes: Camouflage, Shadow, Medical Image Segmentation, and More.

Unified Local and Global Attention Interaction Modeling for Vision Transformers SAM Fails to Segment Anything? -- SAM-Adapter: Adapting SAM in Underperformed Scenes: Camouflage, Shadow, Medical Image Segmentation, and More

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.095509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.095509Z digest=sha256:a1ace59a9c6fa7c293be5e27595b1a11a3894f3277dec4d75b84223b3e6be833

Observation d920fe8d-ec60-44e5-a3db-d792211ba781 · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 6

Resolution
malformed identifier
no resolver link, observed 2026-08-11T04:34:02.100667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.100667Z digest=sha256:6e85f6ec52b95de23d3c31f96798fbd811861e8a588587771f154739c2d7f0e9

Observation 068a3ac1-b34c-4d50-8f9f-8bc81fd6bf04 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Unified Local and Global Attention Interaction Modeling for Vision Transformers An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.105878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.105878Z digest=sha256:4082c6649a93d18896dff3504239fc9595b4992e1c17ee4eb9176575dbdf57e5

Observation 68fb8423-98e0-41b3-813a-0818bcd50c8d · outbound

This paper cites Concealed Object Detection.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Concealed Object Detection

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:34:03.274848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T04:34:02.110714Z digest=sha256:4e6e0dda118a67bf7d06c1187b7346783e2868b02cbb7b2fb5f30da116bbda9f

Observation 1a0bf6ae-f741-46b5-8b39-7b5e7eeff726 · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.115665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.115665Z digest=sha256:c783dbd77929124c0d2cbf1ee6d8abfe0b087ce098a4ebe1fa24884a5af86507

Observation 73b23bca-a75f-4879-b3ec-1d13cfe03512 · outbound

This paper cites Fast R-CNN.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Fast R-CNN

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.120229Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.120229Z digest=sha256:4206fa0b53897fdeb96aa36ef9223b5bad4a5747dd4f9f0d2716116afc4f27e4

Observation f47e55f7-bccf-4bcf-9040-b5bcb3d395d2 · outbound

This paper cites Rich feature hierarchies for accurate object detection and semantic segmentation.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Rich feature hierarchies for accurate object detection and semantic segmentation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.124878Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.124878Z digest=sha256:981d52b20ab330fa64574aa6368144e5ccd259fb21e8d959b18b3319c44c4705

Observation 105fa00f-3743-44de-af76-3b0dce7e19c5 · outbound

This paper cites Goldberger, Luis A.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Goldberger, Luis A

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.129567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.129567Z digest=sha256:fa70ec2e34320aaf3121e99a43dafec1b05bbede1b272e7effee26a3231d4b3f

Observation 7b199c26-bcd3-467d-bd0b-566ff6fd9edb · outbound

This paper cites CMT: Convolutional Neural Networks Meet Vision Transformers.

Unified Local and Global Attention Interaction Modeling for Vision Transformers CMT: Convolutional Neural Networks Meet Vision Transformers

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:34:03.149792Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T04:34:02.133793Z digest=sha256:b0c90394cee9ace488a6c8c724a023ed117cc0157b0504f4f6b33b2e196a8bd6

Observation 4d20afe7-52e0-44dd-9c4f-b818083f4852 · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:34:03.608646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T04:34:02.138382Z digest=sha256:b16243da362f69ebb699eab1c1735fa4fd1629e778aa4f5837612aa768eceed0

Observation 3f033212-3b81-420b-b429-cefa7a0520a1 · outbound

This paper cites Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Spatial Pyramid Pooling in Deep Convolutional Networks for Visual Recognition

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.146803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.146803Z digest=sha256:3c5db706113994501bbc0ad1b0b3120bb6bf12c95d8805192d0eff06a3b28331

Observation 4e8854c7-6b64-4e60-a66a-f5fe6cbfe725 · outbound

This paper cites Deep Residual Learning for Image Recognition.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Deep Residual Learning for Image Recognition

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.151323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.151323Z digest=sha256:1d14acf60aef68e7216b1dd756a89c470a42ec0f678e533ab7f536469a68bcfb

Observation 18bac6bf-cb98-427e-b1bd-b8855302fea2 · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:34:03.595134Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T04:34:02.156743Z digest=sha256:2f7c2be40767e68f5972f607dc51c7bb2e5b7e43e6187df35a96c29b15c96032

Observation 5e0380c1-9d67-4cdf-8679-c1428be32ecf · outbound

This paper cites Berg, Wan- Yen Lo, Piotr DollÃąr, and Ross Girshick.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Berg, Wan- Yen Lo, Piotr DollÃąr, and Ross Girshick

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.166020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.166020Z digest=sha256:e86167ee18d5dc5b20eed0c3f11a14529514cda55101c0742e454ccde510603e

Observation 41fab339-7f56-48a0-9a2f-c9fb48f9f650 · outbound

This paper cites Similarity of Neural Network Representations Revisited.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Similarity of Neural Network Representations Revisited

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.170332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.170332Z digest=sha256:2b69d7e986869a804fbe43c9bdbd488d5d51374baf931357a39826422e5fc593

Observation 0fbd79ba-4e3a-41c8-86c9-8fcbffbf3bae · outbound

This paper cites CornerNet: Detecting Objects as Paired Keypoints.

Unified Local and Global Attention Interaction Modeling for Vision Transformers CornerNet: Detecting Objects as Paired Keypoints

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.174795Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.174795Z digest=sha256:8af8038433af58e98ca27e0109b27445ba07c207fc2c4a941cc07ca5cafea5ec

Observation 2a5c7759-8818-4780-8eff-913e32be5057 · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 21

Resolution
verified exact
doi, observed 2026-08-11T04:34:02.413966Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T04:34:02.178742Z digest=sha256:d553f7a320fa9e18b703165f345338b81e3624dca03be7e350794597ccf03138

Observation 1e348a27-9edb-4e0d-91db-1b652d3c0096 · outbound

This paper cites Feature Pyramid Networks for Object Detection.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Feature Pyramid Networks for Object Detection

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.182700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.182700Z digest=sha256:f5917ac8f118435483a3cbd7e56a7a854c309ea492227291509ede7b31662ee2

Observation 25fce831-bc8a-4da2-abaf-37d25035ce0f · outbound

This paper cites Girshick, Kaiming He , and Piotr Dollár.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Girshick, Kaiming He , and Piotr Dollár

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:03.581831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T04:34:02.186774Z digest=sha256:a8b0d21555d91c13de902b1505de0313b475953f046dc655eeaabab8d91ab27f

Observation 49fa9b09-8bcd-46f6-851d-962f147d5bd9 · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:34:03.567191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T04:34:02.195040Z digest=sha256:f0990563c968daa8c4bda7898f3bf2a76bce6f250f4d1a336e000538e82df6df

Observation c840ed8a-a38c-41c4-8cec-df60ba1f1fd0 · outbound

This paper cites SSD: Single Shot MultiBox Detector.

Unified Local and Global Attention Interaction Modeling for Vision Transformers SSD: Single Shot MultiBox Detector

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.204149Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.204149Z digest=sha256:a4cf985d397f38172faa8c8d7bab92e37f54411201f9468a056c848f040ebcf9

Observation e5939e69-769e-4628-bc2f-7753a423bb68 · outbound

This paper cites Swin Transformer V2: Scaling Up Capacity and Resolution.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Swin Transformer V2: Scaling Up Capacity and Resolution

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.208712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.208712Z digest=sha256:b3973f7bb3ea3aa50f3c052664a4d4030a78b4213e3b60e1f85832b2d63de83e

Observation 0d9511f5-4150-4826-96ba-6f00a2c07783 · outbound

This paper cites RTMDet: An Empirical Study of Designing Real-Time Object Detectors.

Unified Local and Global Attention Interaction Modeling for Vision Transformers RTMDet: An Empirical Study of Designing Real-Time Object Detectors

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.217206Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.217206Z digest=sha256:05ca085a52ebb5b02620a7b3c0d4809869beabbdf8f0fbf1884c282ee60c2b1a

Observation c1c495be-6c38-42eb-b4fb-2fd689ef2ef1 · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:34:03.540028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T04:34:02.221711Z digest=sha256:0dd34414524d3ed3f287d523f7a576d652ae766f07afa3808ec26c2e58eb817a

Observation d6eb5e06-7cb7-4b5f-ac19-fbf15ba2712c · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:34:03.526104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T04:34:02.230494Z digest=sha256:b9901a439769a586c0afdf6e92425dbff3736e1b6d965fb3b1d4b47d220c1a47

Observation 87c790a6-80dd-4393-9a0f-9a4c6f5a99d6 · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.234743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.234743Z digest=sha256:1bf5415347f30f895db7193ddd33885fe90522ef3318fea489224790f82c07ad

Observation c6c93072-9be3-4538-a6fb-5a78a6b4ac23 · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:34:03.511260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T04:34:02.239108Z digest=sha256:98f97fd7038988a5a90795e92f68058ad76af84fa012624298c02dfe2cacd983

Observation 834b2e6d-3e01-47f3-9800-b6ba83bd9411 · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:34:03.495960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T04:34:02.243315Z digest=sha256:ddbf8adb9ba0c7a5edcf14a9469c0efb7a0bb4bb4fd6c12827a3e1c87cb9d9a5

Observation aa493100-75dd-4e97-98e5-8a631db8c28d · outbound

This paper cites In Computer Vision âĂŞ ECCV 2022 Workshops: Tel A viv, Israel, October 23âĂŞ27, 202 2, Proceedings, Part VII (Tel Aviv, Israel).

Unified Local and Global Attention Interaction Modeling for Vision Transformers In Computer Vision âĂŞ ECCV 2022 Workshops: Tel A viv, Israel, October 23âĂŞ27, 202 2, Proceedings, Part VII (Tel Aviv, Israel)

Reference 34

Resolution
verified exact
doi, observed 2026-08-11T04:34:02.399237Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T04:34:02.226254Z digest=sha256:5702de1b94ffb7c49d7c08011d59a1c433b15b5364320f51a265aed46b6d5fa9

Observation 3608339d-eaf3-479d-b709-927492599dfb · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.252092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.252092Z digest=sha256:71a3a8e948fa8370ecbb6a44c413c690cbd35114bf9d2045207e774cb8408442

Observation 6349ec22-1f98-4cb8-852a-556dc607d144 · outbound

This paper cites You Only Look Once: Unified, Real-Time Object Detection.

Unified Local and Global Attention Interaction Modeling for Vision Transformers You Only Look Once: Unified, Real-Time Object Detection

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.256193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.256193Z digest=sha256:2dd1e1a7dc76c22f248e0c97a472b3a846bb19b7024c4ba1787ee7206b57aead

Observation 92ad948b-d174-40e3-a31e-ad24d503f29b · outbound

This paper cites Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.260971Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.260971Z digest=sha256:874798d18f13c5eff15b875809a1c308ce19b4cda57b083a12b8b54312c20f80

Observation 6e0e3ab7-ea8d-4460-8a77-87b671302cff · outbound

This paper cites Wu, Safwan S.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Wu, Safwan S

Reference 38

Resolution
malformed identifier
no resolver link, observed 2026-08-11T04:34:02.265565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.265565Z digest=sha256:9c2891382c7934779d3a60a3c65fa3f5fd665b7c82c1311b4f4f9f597a506e36

Observation 3c005129-e5c7-4535-b446-c23dfc39da68 · outbound

This paper cites Designing Network Design Spaces.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Designing Network Design Spaces

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.247765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.247765Z digest=sha256:fd5c00c8a3dfa25c02eba3f47f97660f614a9f3d806b707ea7e71201a089b25c

Observation ba13177f-9bf0-471e-9195-d98182a76eef · outbound

This paper cites EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks.

Unified Local and Global Attention Interaction Modeling for Vision Transformers EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.274332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.274332Z digest=sha256:af1bf9a2ccd74d0e74aeb87d3f82c16c7d8bd4a1a25377267078f3caa715378c

Observation 16d5f83c-c569-413c-9e76-288417e3fba6 · outbound

This paper cites Attention Is All You Need.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Attention Is All You Need

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.278900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.278900Z digest=sha256:ba203f85df21d348ece911954b84a3e6f8f5aac7ce87f1bbff05cf5313f85a1a

Observation 777a7d41-9015-43dd-b854-7f2413e41bec · outbound

This paper cites Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Pyramid Vision Transformer: A Versatile Backbone for Dense Prediction without Convolutions

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.283296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.283296Z digest=sha256:a218b195d4378bd608f2f7b0a2bb88fb83e9d9d8b15050ec33f83f894c82545c

Observation 3e7a63b4-e830-4d54-8eb5-e6acd6f21135 · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.287811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.287811Z digest=sha256:8e291888c75e717df79edb712821d078eda0232b273a8dcdd2b05118b4458e48

Observation 7a4333fc-7eb7-43bd-b467-216f5ccae013 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.269929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.269929Z digest=sha256:aeb7a6b15eecf83a6824343f62933917a837f4633b04de49841eb63468fadcf2

Observation a5c2843f-2825-4012-a6e8-b5efc510faab · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:34:03.482056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T04:34:02.296177Z digest=sha256:a09b9da947d7582ece473837b69cbd4432c1780bfb8fb041860de8401a25efbb

Observation c18b2d64-e096-4035-98aa-d7ef4360257b · outbound

This paper cites CvT: Introducing Convolutions to Vision Transformers.

Unified Local and Global Attention Interaction Modeling for Vision Transformers CvT: Introducing Convolutions to Vision Transformers

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.303884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.303884Z digest=sha256:2bb14c324c47824a83e9a3ab4ffb6f04284734f6ba042e9123ce7e8b1601bedc

Observation 7e4fd844-f2b2-405a-bf3a-98dd106fe840 · outbound

This paper cites Vision Transformer with Deformable Attention.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Vision Transformer with Deformable Attention

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.307775Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.307775Z digest=sha256:ed16eae257c6d16295478186ad0cb56fde0a6d3d86271ce8dd4e08a464c71a72

Observation 0cd8eb6a-d2a0-466b-a740-54b8b849facf · outbound

This paper cites DAT++: Spatially Dynamic Vision Transformer with Deformable Attention.

Unified Local and Global Attention Interaction Modeling for Vision Transformers DAT++: Spatially Dynamic Vision Transformer with Deformable Attention

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.312181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.312181Z digest=sha256:228861b8b177f3c6dfb528f5d7df311b8423b3fb55ab8346eec51371871490bc

Observation 111a8269-85dd-4803-87e2-86f69389b322 · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.292082Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.292082Z digest=sha256:10d66d1a512f7c4c3c34d394676633910b593974db7821cd3a1e7705fe7ac005

Observation 02108eb6-d90c-429b-9cdb-33eb7c7c405f · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:34:03.439246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T04:34:02.320053Z digest=sha256:253be33315b18f8bd2d45dc9999a0366dfd6d52509abb7033ac2c9378864c1e3

Observation b92fc2b6-7f21-40f3-bdb9-c60dbe31c84b · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 51

Resolution
malformed identifier
raw_fallback, observed 2026-08-11T04:34:03.468138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T04:34:02.300079Z digest=sha256:c77ced3e983e3d4ca1b79a5f28c2e94367014a461a520b0529fca3a754cececc

Observation e8ab321e-0a2c-4bc5-9805-592a55e3392e · outbound

This paper cites Detecting Twenty-thousand Classes using Image-level Supervision.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Detecting Twenty-thousand Classes using Image-level Supervision

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.329007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.329007Z digest=sha256:5c1583bb5a3d3c9fe7bfee3bcd8757a94929cba1762bcbe2a9a1fdd88773ef41

Observation a20076af-da48-43d9-98a2-2ffe6485f391 · outbound

This paper cites Deformable DETR: Deformable Transformers for End-to-End Object Detection.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Deformable DETR: Deformable Transformers for End-to-End Object Detection

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.333517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.333517Z digest=sha256:8103ef6ba9c7a6600ca6a992cd481701e4b42af3564947c88b8f770ad93c14f5

Observation 425a0832-06ff-48f9-aa15-a3124d76c745 · outbound

This paper cites an unresolved cited work.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-11T04:34:03.453889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T04:34:02.316240Z digest=sha256:ad59d80ad438a8fb8e4d2210de58c03ec20ab2ecd5587d8e30ff8447d1cc5ea9

Observation 2d71f19f-a4f0-471d-915b-09ef6c0ecb1f · outbound

This paper cites Refiner: Refining Self-attention for Vision Transformers.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Refiner: Refining Self-attention for Vision Transformers

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.324295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.324295Z digest=sha256:9f251ed011041eeeda2e9b81f703a8bd4178bf0d0f6b4ac85f6147511dab0add

Observation 8a833448-8f08-4bad-9398-4f50009ee08c · outbound

This paper cites Focal Loss for Dense Object Detection.

Unified Local and Global Attention Interaction Modeling for Vision Transformers Focal Loss for Dense Object Detection

Reference 2017

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.190793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.190793Z digest=sha256:4a62974c08c6114f60ebc0ffafbadd45f9b4e3d923c192962d3ede19902ca1d6

Observation 37315f81-30d0-4d19-b9a3-87a0d21739dc · outbound

This paper cites In European Conference on Computer Vision.

Unified Local and Global Attention Interaction Modeling for Vision Transformers In European Conference on Computer Vision

Reference 2021

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T04:34:03.553975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T04:34:02.199377Z digest=sha256:c5805b3968298af653260c7432d23f72d48091f866420c74a2e5f0f59d320bb1

Observation 2abe69fa-245a-45ba-b72f-dc6f285dc221 · outbound

This paper cites In 2023 IEEE/CVF Con- ference on Computer Vision and Pattern Recognition (CVPR).

Unified Local and Global Attention Interaction Modeling for Vision Transformers In 2023 IEEE/CVF Con- ference on Computer Vision and Pattern Recognition (CVPR)

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-11T04:34:02.142594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:34:02.142594Z digest=sha256:3986945aa872865a395c76dbcc1bb0bfcc0137733f52eb9427aa227f33cfddc9

Observation c0795a57-27f5-48de-a036-56cbc66feb09 · outbound

This paper cites SurANet: Surrounding-Aware Network for Concealed Object Detection via Highly-Efficient Interactive Contrastive Learning Strategy.

Unified Local and Global Attention Interaction Modeling for Vision Transformers SurANet: Surrounding-Aware Network for Concealed Object Detection via Highly-Efficient Interactive Contrastive Learning Strategy

Reference 2024

Resolution
metadata mismatch
local_arxiv, observed 2026-08-11T04:34:03.027113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-11T04:34:02.161232Z digest=sha256:0311ca5a73c82728c1b0a06c2fd4f8bd0b24ec7b103779d1a240b29e19901bf1

Pith citing papers

No inbound Pith citation observations are available.