Pith. sign in

Paper Citation Record · LEDGER

Native Segmentation Vision Transformers

As of 9 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 1 inbound Pith citation observation for arXiv:2505.16993.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16993 v1

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:55:40.776768Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-13T06:02:40.158866Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T06:07:22.539904Z

Reference resolution

74 of 74 outbound references displayed

  • verified exact0
  • verified fuzzy53
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2460445c-9575-432c-a64a-35e294bcf3ca · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Native Segmentation Vision Transformers Swin transformer: Hierarchical vision transformer using shifted windows

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:34.491663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:34.491663Z digest=sha256:31df2b3675b2847edccde141c1aa0b72d19d8bd08848b1c4b799141b54a66f0e

Observation 39342ace-32b9-4815-9690-66756b5c6ab3 · outbound

This paper cites Convnext v2: Co-designing and scaling convnets with masked autoencoders.

Native Segmentation Vision Transformers Convnext v2: Co-designing and scaling convnets with masked autoencoders

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:50.558282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:34.575933Z digest=sha256:f63b9dd331a374649f3c7806b38220c5eb16613a7a8d0bb95a0208b7d96ae43d

Observation f518c955-48b5-4dc5-a695-53331bd64673 · outbound

This paper cites Neighborhood attention transformer.

Native Segmentation Vision Transformers Neighborhood attention transformer

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:50.429716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:34.685306Z digest=sha256:8b1834336d337fe44a0c7acdd8d2dbab1859476fe3ebf682650a2bc8237b015b

Observation 969585b4-5aa4-4515-985d-81316a64135c · outbound

This paper cites Backpropagation applied to handwritten zip code recognition.

Native Segmentation Vision Transformers Backpropagation applied to handwritten zip code recognition

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:34.815187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:34.815187Z digest=sha256:c071ad639d5e5e389533c073a52d4fa294bd6c59be65313fcbee23e8e7fb07a7

Observation a9843379-0ad2-44e0-a301-20e7a7d3e85d · outbound

This paper cites an unresolved cited work.

Native Segmentation Vision Transformers Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:34.922357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:34.922357Z digest=sha256:6a736e8a5445b3a292fed8122409e1644efd915431abb4b11b878b88b515f568

Observation da6f51d4-8ab3-42fd-99f6-aed459b366ac · outbound

This paper cites Schwing, Alexander Kirillov, and Rohit Girdhar.

Native Segmentation Vision Transformers Schwing, Alexander Kirillov, and Rohit Girdhar

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:35.010276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:35.010276Z digest=sha256:c6845bd3510a8662094df6a27175dc3739c386368903c5b76abdb2773d87873b

Observation cdb95e00-622d-4ef2-8539-659d45743bd9 · outbound

This paper cites Feature pyramid networks for object detection.

Native Segmentation Vision Transformers Feature pyramid networks for object detection

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:35.085000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:35.085000Z digest=sha256:d108afe799bd96d0ad7087f8ce5a1a00db3a3a630325aa1f1c99ab365dce64ba

Observation 991767dd-4efc-4e69-9305-9957da71bd7f · outbound

This paper cites FaPN: Feature-aligned pyramid network for dense image prediction.

Native Segmentation Vision Transformers FaPN: Feature-aligned pyramid network for dense image prediction

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:50.262110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:35.224616Z digest=sha256:80a30bc535d04eb42eeb029d10949ae6f622e00c92b0eae9e376e28ea65890aa

Observation ee282015-9ab6-4dc9-a550-1d84fc6a92af · outbound

This paper cites an unresolved cited work.

Native Segmentation Vision Transformers Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:55:50.132482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:35.321961Z digest=sha256:46518b842b17b387b8d1524e555bd3bbd29b0c36910addf896d6c07b31313a85

Observation 3e5f1493-744d-47ed-934f-63537f47f13b · outbound

This paper cites Clusterformer: Clustering as a universal visual learner.

Native Segmentation Vision Transformers Clusterformer: Clustering as a universal visual learner

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:49.967714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:35.479711Z digest=sha256:cdb90e9f0243b1a5e824320797adc8a37491971a13f9722cafa74a6086c3553c

Observation fa701606-54c5-4881-b075-b7c7e1865346 · outbound

This paper cites Learning hierarchical image segmentation for recognition and by recognition.

Native Segmentation Vision Transformers Learning hierarchical image segmentation for recognition and by recognition

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:49.870902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:35.562953Z digest=sha256:dadf942556a4e81733929aaf2cd87d34ff31ff3cb9e4f94a5bc0e27b4def35fd

Observation 15597a3d-7613-4f33-8943-a46cfd1b6687 · outbound

This paper cites Tcformer: Visual recognition via token clustering transformer.

Native Segmentation Vision Transformers Tcformer: Visual recognition via token clustering transformer

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:49.742996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:35.643011Z digest=sha256:a9f5c12ab5720b5bb9924c7380b693e876e77f4e6f73e96c023f19bc06cc319a

Observation 2d765e73-8eea-41de-9bc5-055e8bb36032 · outbound

This paper cites Schwing, and Alexander Kirillov.

Native Segmentation Vision Transformers Schwing, and Alexander Kirillov

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:49.604488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:35.735251Z digest=sha256:4704fd15d244a884b643bba12fe8d52f5e967d634981ebbf93695efdb32dca28

Observation 6f38e803-886f-4711-aeef-f7b547b01593 · outbound

This paper cites Slic superpixels compared to state-of-the-art superpixel methods.

Native Segmentation Vision Transformers Slic superpixels compared to state-of-the-art superpixel methods

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:49.402810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:35.837086Z digest=sha256:5fdc898ad0a8522bb78fa88a3bdc2defe4d058b2c0d1872a78735e5492cf6deb

Observation 8bd76a18-d8a2-42f5-999d-27a3bca78879 · outbound

This paper cites Object-centric learning with slot attention.

Native Segmentation Vision Transformers Object-centric learning with slot attention

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:49.233922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:35.935254Z digest=sha256:2aa1e0c8b31c2fef1379cca105f68d7a363bbf2c78f89d3e0dcd45c987383570

Observation 2fa49496-b243-4745-a451-4f5bb7571d4b · outbound

This paper cites Mean shift: A robust approach toward feature space analysis.

Native Segmentation Vision Transformers Mean shift: A robust approach toward feature space analysis

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:49.016911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:36.074370Z digest=sha256:62a1a0264a12f69fdb231f40a2192179f6927d0843cd7254e2930ca3209d43a8

Observation f1abbd73-3554-4b10-82bc-659c6ddb72ae · outbound

This paper cites Felzenszwalb and Daniel P.

Native Segmentation Vision Transformers Felzenszwalb and Daniel P

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:48.853816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:36.168668Z digest=sha256:94961ca5404ad12a0dacf53a0b763e12514456feb4a7a80c6f55cf57e753ea46

Observation b6cc0892-2632-4d79-a46b-dea5645c87d8 · outbound

This paper cites Contour detection and hierarchical image segmentation.

Native Segmentation Vision Transformers Contour detection and hierarchical image segmentation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:48.710783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:36.279399Z digest=sha256:31a57e04ebb19a6cccec673e112d1105563bc84baed4b9beebefe8f939e3cb57

Observation 6909a4cd-4635-4a9c-b02d-8368068fba64 · outbound

This paper cites Seeds: Superpixels extracted via energy-driven sampling.

Native Segmentation Vision Transformers Seeds: Superpixels extracted via energy-driven sampling

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:48.536144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:36.403076Z digest=sha256:c469fb5c4a9251dfca4b4806c455ad4d810c9e94e5224b7ee36fb358e81c9c43

Observation 6f8390d9-e7de-427f-ba63-7e7c1684f633 · outbound

This paper cites Semantic understanding of scenes through the ade20k dataset.

Native Segmentation Vision Transformers Semantic understanding of scenes through the ade20k dataset

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:48.395546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:36.473035Z digest=sha256:85988ebb3189b472ada5cac47ccfb91bdc86118aaf06ddcc277437b0ffd1e1e8

Observation 22559590-2a4f-410d-bdfa-744dd9271f2f · outbound

This paper cites Panoptic segmentation.

Native Segmentation Vision Transformers Panoptic segmentation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:48.202978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:36.569923Z digest=sha256:c91f1b0a806c36b232c5bd5f404802523f0e245f213e2c363184a7d6c8d72a07

Observation cc44dad6-1121-4f2f-a77c-7b88b2bded88 · outbound

This paper cites Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position.

Native Segmentation Vision Transformers Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:36.693822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:36.693822Z digest=sha256:1cd134b73413c73add4e40d9f7ffdcf80cf00665031d741a52fe66c422374284

Observation 60d9cf58-f5f8-4f77-9266-976677bb21c0 · outbound

This paper cites Gradient-based learning applied to document recognition.

Native Segmentation Vision Transformers Gradient-based learning applied to document recognition

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:48.038344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:36.746616Z digest=sha256:cbda56b6218d06dcaf3edc2e28e4e4f0798f295821c18aa617b3a1dc18b91b66

Observation 11364c9b-1f78-438f-add2-63c1c7732b99 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Native Segmentation Vision Transformers An image is worth 16x16 words: Transformers for image recognition at scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:36.825137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:36.825137Z digest=sha256:4d02092f9bf066b83a1bb260895c987af36319bf536e10c526a2c0c8ec753061

Observation edb1c32d-38b4-47cb-b971-d50f5cdb5fe7 · outbound

This paper cites A convnet for the 2020s.

Native Segmentation Vision Transformers A convnet for the 2020s

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:36.922349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:36.922349Z digest=sha256:5b327742c171423425fdb83a0620f1177fea697d83bbd8450ef1c5fbfb46e58b

Observation 0350cb82-c86e-4800-bd90-f7e8a3f9b611 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Native Segmentation Vision Transformers SAM 2: Segment Anything in Images and Videos

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:37.004068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:37.004068Z digest=sha256:004590246257bd0fd939985ba0df7b7692ff8b4728d0a661e684a41f296c7806

Observation 9851423e-6728-49b0-a99d-6c901950ccb1 · outbound

This paper cites Fully convolutional networks for semantic segmentation.

Native Segmentation Vision Transformers Fully convolutional networks for semantic segmentation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:37.099662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:37.099662Z digest=sha256:72e727094c2689f83ffa766d71d88ffb3fe6ce4302a92cd468c07ac6bcf3e3bd

Observation 8e70d7bb-005f-436d-adb4-ffaa0f5b9655 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

Native Segmentation Vision Transformers U-net: Convolutional networks for biomedical image segmentation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:47.850044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:37.172521Z digest=sha256:1ab08a545588b5f36cde253093271b1d70bd6832db6decbdb908640ed8b7f903

Observation 7296edd4-bfdd-41fa-b45b-d1ef3d797e07 · outbound

This paper cites Rich feature hierarchies for accurate object detection and semantic segmentation.

Native Segmentation Vision Transformers Rich feature hierarchies for accurate object detection and semantic segmentation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:47.634897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:37.259829Z digest=sha256:6ff076049cecf1882e1a91597fef68fa6afdbe26fcc89da4650d03d1625dbc22

Observation 4617a109-394d-4a8d-b17d-2a6863660fa7 · outbound

This paper cites Fast r-cnn.

Native Segmentation Vision Transformers Fast r-cnn

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:37.360904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:37.360904Z digest=sha256:33f9b172c1b9b7eaa57fc1a108c57b780b58ca6128c12f99752331f61c2aec95

Observation abbb1278-5402-40c8-b8a9-a796dda25b68 · outbound

This paper cites End-to-end object detection with transformers.

Native Segmentation Vision Transformers End-to-end object detection with transformers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:37.466195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:37.466195Z digest=sha256:baba5f1ccaaffbd60696c40302255816cb93386090b1b56f2a5c73bd180c7f55

Observation e95186b4-c75a-4946-8c30-eb2aae10353d · outbound

This paper cites Normalized cuts and image segmentation.

Native Segmentation Vision Transformers Normalized cuts and image segmentation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:47.416823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:37.548532Z digest=sha256:6b8831bbcfc50b1c2c2f01bd97cb14c6790cf8ed6331ac64cc5f4855806980ac

Observation 2a1bbe62-2422-4b2f-9320-d17f37cbd3cb · outbound

This paper cites Multiscale combinatorial grouping.

Native Segmentation Vision Transformers Multiscale combinatorial grouping

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:47.198109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:37.643606Z digest=sha256:1c80cd8667412ac8e1065a81acd53e22e42443db440a7dcac3cab56934c9cce3

Observation 46dc9746-5865-474f-b668-83260ca4fe38 · outbound

This paper cites Superpixel samping networks.

Native Segmentation Vision Transformers Superpixel samping networks

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:47.016623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:37.702010Z digest=sha256:805a6dc7348442958c0dc4568b839e67032e08aecdba0ead9c0b29d6dec8b2b8

Observation abda8ddf-d305-4097-8a72-bf103965d096 · outbound

This paper cites an unresolved cited work.

Native Segmentation Vision Transformers Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:55:46.802461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:37.819714Z digest=sha256:a726fe285df3c3e54f79840c04b6867da01bd62694345782e64e3767a5edf059

Observation db5cc85c-a4f2-48a8-b46e-064f879dc9a9 · outbound

This paper cites an unresolved cited work.

Native Segmentation Vision Transformers Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:55:46.557896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:37.930389Z digest=sha256:cdc6181a65b4d7566a904a97355a77134d595a34b0e8687c97a391d5b089015f

Observation 9eca3a2d-053d-4021-b61b-b205ef653898 · outbound

This paper cites Unified perceptual parsing for scene understanding.

Native Segmentation Vision Transformers Unified perceptual parsing for scene understanding

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:46.374569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:38.018959Z digest=sha256:ca1a81b4ac206de9079a0f8ad9b76243b39374e1818fbef653f6d09e0d5f5612

Observation 58de59c6-c80f-492b-a683-8e12dfa77e14 · outbound

This paper cites Berg, and Li Fei-Fei.

Native Segmentation Vision Transformers Berg, and Li Fei-Fei

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:46.201278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:38.112464Z digest=sha256:895cf39d9d5f605b116dce209672518c01d5c5f03f957343b1bc11b1d79dad20

Observation e4bb2b64-dfeb-44b8-8a9f-32b9a46018a4 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Native Segmentation Vision Transformers Learning transferable visual models from natural language supervision

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:38.182246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:38.182246Z digest=sha256:22d8b85b841d1385214fe29581b78037226c46d27b76626ab202ae97c9c9d31b

Observation 051e48ba-726e-45b9-b577-5958b6803ef4 · outbound

This paper cites Le, Yun- Hsuan Sung, Zhen Li, and Tom Duerig.

Native Segmentation Vision Transformers Le, Yun- Hsuan Sung, Zhen Li, and Tom Duerig

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:46.027620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:38.254570Z digest=sha256:5a1d55834813feb15ecac1cb9604f85de35c3f7255c9cd6d2c1597d9a84ccd39

Observation 4c5dfd3e-7978-414e-93c7-3a0c2cfde865 · outbound

This paper cites A simple framework for text-supervised semantic segmentation.

Native Segmentation Vision Transformers A simple framework for text-supervised semantic segmentation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:45.829111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:38.330391Z digest=sha256:18e469586bcaf3b5888bbbd5a707ba07219fed70ede939f18bc2f0f740c22d16

Observation 522dc38d-724f-4e3d-bd7d-5badc2072f75 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

Native Segmentation Vision Transformers Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:45.691037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:38.424956Z digest=sha256:3e2333ae12c66f1180f868cfeb2fa065e22af7179d3ff1f109936c0b00718338

Observation 4b1d9c74-8136-437f-94e8-736a37685071 · outbound

This paper cites Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts.

Native Segmentation Vision Transformers Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:38.530639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:38.530639Z digest=sha256:ad6cc146651af881526e1cf8d8b6335ecd16b3b4440aa108fa8b74b6b937ae3d

Observation 7663898a-73b5-447d-912e-d66af4100685 · outbound

This paper cites Redcaps: Web-crawled image-text data created by the people, for the people.

Native Segmentation Vision Transformers Redcaps: Web-crawled image-text data created by the people, for the people

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:45.509767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:38.622568Z digest=sha256:02b2ab324938361893cf7fdef8d121e0064751cf6059e1eab0c96f3332a17309

Observation 421a5750-7cb4-45b2-844d-c2068675156c · outbound

This paper cites Learning to generate text-grounded mask for open-world semantic segmentation from only image-text pairs.

Native Segmentation Vision Transformers Learning to generate text-grounded mask for open-world semantic segmentation from only image-text pairs

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:45.310461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:38.698678Z digest=sha256:fbebed86e4e3951f37ee6fc6261b502d1ffadd0390401e2895403522ba0c5ab8

Observation e6c0e57e-0842-4104-93aa-39c716c4ff09 · outbound

This paper cites Everingham, L.

Native Segmentation Vision Transformers Everingham, L

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:45.134327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:38.790705Z digest=sha256:6c6ed0fb51930af3c6ddba9c702de8a91d04cd278541e0b3f66769d60c537823

Observation 72274a1a-d26c-4238-8e16-a879ee05c245 · outbound

This paper cites The role of context for object detection and semantic segmentation in the wild.

Native Segmentation Vision Transformers The role of context for object detection and semantic segmentation in the wild

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:44.877729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:38.870638Z digest=sha256:1812133ad56040cf09ae52c0fe4f6ad2748dfa2b7a40e865b21ac234192a9315

Observation 12ff2374-85bd-42dc-a2a2-3e24c9307cc2 · outbound

This paper cites an unresolved cited work.

Native Segmentation Vision Transformers Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:55:44.748497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:38.933552Z digest=sha256:14a8cd4a03482cb841d33409a976c14f2f8ca34227eeeecc89f0aacce5b0fd42

Observation 0863d82c-10c4-4e67-9774-4d4e6d0519a9 · outbound

This paper cites Coco-stuff: Thing and stuff classes in context.

Native Segmentation Vision Transformers Coco-stuff: Thing and stuff classes in context

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:44.550601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:39.012943Z digest=sha256:74ab5c01ac7a842df2b6685056cb5b10035674847ef8b3fd4b12d00ede7a0ebf

Observation 45573f47-e3d0-44e4-acab-46e74982d378 · outbound

This paper cites Cordts, M.

Native Segmentation Vision Transformers Cordts, M

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:39.058264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:39.058264Z digest=sha256:6812857de9838d94ca77214497f39f38fce5946741239feef4e1479ed0c3d910

Observation 971d4e1d-0360-4075-8853-572bf21cf678 · outbound

This paper cites Open- world semantic segmentation via contrasting and clustering vision-language embedding.

Native Segmentation Vision Transformers Open- world semantic segmentation via contrasting and clustering vision-language embedding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:44.340556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:39.128618Z digest=sha256:7a0439dd953fe2e9aa9a074837772e03f0d0447e321856318cb0b58015ac1a08

Observation c5829d93-cf5c-4493-a113-a528c55e49b4 · outbound

This paper cites Segclip: Patch aggregation with learnable centers for open-vocabulary semantic segmentation.

Native Segmentation Vision Transformers Segclip: Patch aggregation with learnable centers for open-vocabulary semantic segmentation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:44.151077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:39.200859Z digest=sha256:8b5ac22e13892096b195f370069fccf5d65770c6a248e7461335bc15f7f144f8

Observation e922457d-570b-402d-829f-1ed192d4d53b · outbound

This paper cites Image-text co-decomposition for text-supervised semantic segmentation.

Native Segmentation Vision Transformers Image-text co-decomposition for text-supervised semantic segmentation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:43.994656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:39.283083Z digest=sha256:1c9c35eac1b6ee5456a38967816c2261450118a3eb0992efdc39404a311be128

Observation fab32581-2094-4638-8719-105837d4e6f9 · outbound

This paper cites Viewco: Discovering text-supervised segmentation masks via multi-view semantic consistency.

Native Segmentation Vision Transformers Viewco: Discovering text-supervised segmentation masks via multi-view semantic consistency

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:43.799254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:39.361768Z digest=sha256:6992f2179cd5bdb5fc8ee44ab8cd8af4778f5a569ddbdefcf93a5c165fe630b2

Observation 8fc88444-f9a1-41e5-9db4-72413e568c66 · outbound

This paper cites Rewrite caption semantics: Bridging semantic gaps for language-supervised semantic segmentation.

Native Segmentation Vision Transformers Rewrite caption semantics: Bridging semantic gaps for language-supervised semantic segmentation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:43.625993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:39.421485Z digest=sha256:82655ee66c2da0e028f736deeeba8d636a69c05111527834b742c56bc0af8b14

Observation fc8b093b-d0cd-47b1-a17d-aabcc22eb2f9 · outbound

This paper cites Uncovering prototypical knowledge for weakly open-vocabulary semantic segmentation.

Native Segmentation Vision Transformers Uncovering prototypical knowledge for weakly open-vocabulary semantic segmentation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:43.458392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:39.494191Z digest=sha256:be27aea47084337a06de464b1242353d9f32dedb5adc8aed8116a1374766b7c2

Observation 6ca380dd-fdbd-4b8c-ad63-df2fcdf1a52a · outbound

This paper cites Efficient inference in fully connected crfs with gaussian edge potentials.

Native Segmentation Vision Transformers Efficient inference in fully connected crfs with gaussian edge potentials

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:43.313744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:39.625717Z digest=sha256:623bc3ab541e52bbc17578e43cbd198851cef21935f6603b208d2d9d5c0b9fa6

Observation 10d3c55d-06db-4062-bd14-7ef921ee45ea · outbound

This paper cites Single-stage semantic segmentation from image labels.

Native Segmentation Vision Transformers Single-stage semantic segmentation from image labels

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:43.172338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:39.659409Z digest=sha256:e9b5b891ef878995c12a18da4f8a2427a54569b818625275fefa73aee9360f1e

Observation c961662f-da5d-4207-8db6-5a18fcc46496 · outbound

This paper cites Vmamba: Visual state space model.

Native Segmentation Vision Transformers Vmamba: Visual state space model

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:43.049455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:39.699866Z digest=sha256:3634365bad4cd2001c4b19a57f3d5fcdb502819f94d68da6c3a34e2fa94d914c

Observation dec6cc42-45af-4d31-ac73-3f8f589fe156 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Native Segmentation Vision Transformers Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:42.871618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:39.776032Z digest=sha256:8db36e10ae3377113eade5e521637cf62e06bf88aaba622706ffbbff898a0723

Observation 67f73f93-4e98-4907-9f6a-301595fefb4f · outbound

This paper cites Panoptic feature pyramid networks.

Native Segmentation Vision Transformers Panoptic feature pyramid networks

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:42.750426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:39.859502Z digest=sha256:5cde51cbfed3eb33853405d9f5bd6985eba93073a17ff64bfc7fb8eb420ce422

Observation d62ce440-cb23-407d-a27d-786cc8396158 · outbound

This paper cites Segmenter: Transformer for semantic segmentation.

Native Segmentation Vision Transformers Segmenter: Transformer for semantic segmentation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:42.618951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:39.938874Z digest=sha256:3ae317ba3b52c6c6b13b2e6203a2e1b740620fbaa4c457cce6447d31c803a4ae

Observation 40dc5a72-0625-40f4-8cc7-38b435006774 · outbound

This paper cites an unresolved cited work.

Native Segmentation Vision Transformers Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:55:42.424557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:39.973272Z digest=sha256:d45e467b1fd3238a358ac60f2f4838c1e362666070ff0657d1a568eef5310a3f

Observation a3bbe0f1-d2f1-4079-87cf-a4b174ad51be · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

Native Segmentation Vision Transformers Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:42.334089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:40.084480Z digest=sha256:2858611ca51a864476690fa04c821f3366dd533da40b3b503a6582d7fd623ab1

Observation 1f78a695-9c15-4dbe-9b6a-30f1560e0474 · outbound

This paper cites Self-attention with relative position rep- resentations.

Native Segmentation Vision Transformers Self-attention with relative position rep- resentations

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:42.192849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:40.167748Z digest=sha256:2e651f01c068e999e7bf6a5794e6b284617994b2e16906d6735c57c9f36c122b

Observation 8c28d6ac-aa0c-4cb7-813d-7735c818d99a · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

Native Segmentation Vision Transformers Roformer: Enhanced transformer with rotary position embedding

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:42.054391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:40.218252Z digest=sha256:d706ad5b4a1153a81d5e76719520966e7bcad6f77fe7dc2ee26368763d35d7c6

Observation 437ba9f9-51cf-4272-8968-749e3242de6f · outbound

This paper cites Rotary position embedding for vision transformer.

Native Segmentation Vision Transformers Rotary position embedding for vision transformer

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:41.934310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:40.261863Z digest=sha256:3afc586f5afcde4ec48055c342484535d7d2bdc0da41237f8bf8e0d4d1603994

Observation bdfafc2b-bed4-4ea6-a779-5aaf230de942 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

Native Segmentation Vision Transformers Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:41.770603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:40.305574Z digest=sha256:7128f84e9aed6264efcbc745e85d656693271d93962a015532c7654147509b20

Observation 3f1e26f5-c66c-415e-bb6b-d6518d192254 · outbound

This paper cites Faster neighborhood attention: Reducing the o(n^2) cost of self attention at the threadblock level.

Native Segmentation Vision Transformers Faster neighborhood attention: Reducing the o(n^2) cost of self attention at the threadblock level

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:41.619633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:40.399294Z digest=sha256:0b261346d8af1fa30b6d13befc17a33b7d5201a763927f014d38fa0e172ea86c

Observation 14e5460f-1f9e-4ca8-8d8e-684f2faea1de · outbound

This paper cites Training data-efficient image transformers; distillation through attention.

Native Segmentation Vision Transformers Training data-efficient image transformers; distillation through attention

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:41.448153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:40.479785Z digest=sha256:c8e101da8deb74ba9e0b0f90458a9203562583fd4cd33fc2d63276925770d6bf

Observation 8a207c97-2d2a-41fc-b4af-80f60be67ba7 · outbound

This paper cites Weinberger.

Native Segmentation Vision Transformers Weinberger

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:41.352789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:40.578280Z digest=sha256:88b3f3de7421ccf82191961682d29fb23f14b965f783b0630ed77ed7ddc6db3e

Observation bfa2a9c9-10fd-4b9b-a662-2997cd397c59 · outbound

This paper cites Empirically, we observe a negligible decrease of downstream classification and segmentation performance, and a significant increase in speed.

Native Segmentation Vision Transformers Empirically, we observe a negligible decrease of downstream classification and segmentation performance, and a significant increase in speed

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:41.181944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:40.651143Z digest=sha256:3fa281c65446fd0b60e298d12072328f4c27b6111a3da5fb178e87a2fb1a3ebd

Observation 92b5b608-2b35-4808-89b9-bab8545c1252 · outbound

This paper cites An image of {CLASS }.

Native Segmentation Vision Transformers An image of {CLASS }

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:41.063393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:40.725910Z digest=sha256:94562616205809a1010893bce2df75175edf9903a9abfcbada70e2114535b69f

Observation e32f542a-07bc-4bab-b9e8-6397532cc881 · outbound

This paper cites an unresolved cited work.

Native Segmentation Vision Transformers Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:55:40.952338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-07T14:55:40.776768Z digest=sha256:352a331a5a930777c2d412b46877c337532f97c65f261568422054da28a2c339

Pith citing papers

Observation ade4aadf-4816-44a2-9198-566b2ec09f15 · inbound

Elastic Attention Cores for Scalable Vision Transformers cites this paper.

Elastic Attention Cores for Scalable Vision Transformers Native Segmentation Vision Transformers

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:07:22.541489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-13T06:02:40.158866Z digest=sha256:9602c684d607af0ad0d3b4b8c4153f4710953e9d99e3781cad68317e08552838