Pith. sign in

Paper Citation Record · LEDGER

Native Segmentation Vision Transformers

As of 10 August 2026, this Paper Citation Record lists 74 of 74 outbound references and 1 inbound Pith citation observation for arXiv:2505.16993.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16993 v1

Coverage vector

measured 74 of 74 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:55:40.776768Z

measured 75 of 75 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-13T06:02:40.158866Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-13T06:07:22.539904Z

Reference resolution

74 of 74 outbound references displayed

  • verified exact0
  • verified fuzzy53
  • unresolved21
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2460445c-9575-432c-a64a-35e294bcf3ca · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Native Segmentation Vision Transformers Swin transformer: Hierarchical vision transformer using shifted windows

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:34.491663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:34.491663Z digest=sha256:31df2b3675b2847edccde141c1aa0b72d19d8bd08848b1c4b799141b54a66f0e

Observation 39342ace-32b9-4815-9690-66756b5c6ab3 · outbound

This paper cites Convnext v2: Co-designing and scaling convnets with masked autoencoders.

Native Segmentation Vision Transformers Convnext v2: Co-designing and scaling convnets with masked autoencoders

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:50.558282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:34.575933Z digest=sha256:6f7269adccbfabfe990695de40d128cc86bc8a37650f8b87909bed4a83cffca8

Observation f518c955-48b5-4dc5-a695-53331bd64673 · outbound

This paper cites Neighborhood attention transformer.

Native Segmentation Vision Transformers Neighborhood attention transformer

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:50.429716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:34.685306Z digest=sha256:3b0255b6147f15f5b267a3b3c81ace0689d484bb828a630d5b53793583b48646

Observation 969585b4-5aa4-4515-985d-81316a64135c · outbound

This paper cites Backpropagation applied to handwritten zip code recognition.

Native Segmentation Vision Transformers Backpropagation applied to handwritten zip code recognition

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:34.815187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:34.815187Z digest=sha256:c071ad639d5e5e389533c073a52d4fa294bd6c59be65313fcbee23e8e7fb07a7

Observation a9843379-0ad2-44e0-a301-20e7a7d3e85d · outbound

This paper cites an unresolved cited work.

Native Segmentation Vision Transformers Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:34.922357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:34.922357Z digest=sha256:6a736e8a5445b3a292fed8122409e1644efd915431abb4b11b878b88b515f568

Observation da6f51d4-8ab3-42fd-99f6-aed459b366ac · outbound

This paper cites Schwing, Alexander Kirillov, and Rohit Girdhar.

Native Segmentation Vision Transformers Schwing, Alexander Kirillov, and Rohit Girdhar

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:35.010276Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:35.010276Z digest=sha256:c6845bd3510a8662094df6a27175dc3739c386368903c5b76abdb2773d87873b

Observation cdb95e00-622d-4ef2-8539-659d45743bd9 · outbound

This paper cites Feature pyramid networks for object detection.

Native Segmentation Vision Transformers Feature pyramid networks for object detection

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:35.085000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:35.085000Z digest=sha256:d108afe799bd96d0ad7087f8ce5a1a00db3a3a630325aa1f1c99ab365dce64ba

Observation 991767dd-4efc-4e69-9305-9957da71bd7f · outbound

This paper cites FaPN: Feature-aligned pyramid network for dense image prediction.

Native Segmentation Vision Transformers FaPN: Feature-aligned pyramid network for dense image prediction

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:50.262110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:35.224616Z digest=sha256:249bd452881dba8211d2208fe2fe8bc0928c0d76eb6154110753b1c04dcb2410

Observation ee282015-9ab6-4dc9-a550-1d84fc6a92af · outbound

This paper cites an unresolved cited work.

Native Segmentation Vision Transformers Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:55:50.132482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:35.321961Z digest=sha256:13e0207eff3c7ec4a35e1ded150b1baf9d443d562b84455bf420059214103283

Observation 3e5f1493-744d-47ed-934f-63537f47f13b · outbound

This paper cites Clusterformer: Clustering as a universal visual learner.

Native Segmentation Vision Transformers Clusterformer: Clustering as a universal visual learner

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:49.967714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:35.479711Z digest=sha256:ba0339c2914df32498edbf3d278dc9da37c4c09b9a8d0ec4ba53cb3919b292d6

Observation fa701606-54c5-4881-b075-b7c7e1865346 · outbound

This paper cites Learning hierarchical image segmentation for recognition and by recognition.

Native Segmentation Vision Transformers Learning hierarchical image segmentation for recognition and by recognition

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:49.870902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:35.562953Z digest=sha256:1e8ac51a8c27532b7a250673e9ae37fb95510adf807667c9d1b8c53b9e6896b6

Observation 15597a3d-7613-4f33-8943-a46cfd1b6687 · outbound

This paper cites Tcformer: Visual recognition via token clustering transformer.

Native Segmentation Vision Transformers Tcformer: Visual recognition via token clustering transformer

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:49.742996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:35.643011Z digest=sha256:0d2ec4556276456c7c4d8bed40255a48dcdae8b2c1a3677556a6b7d1dcc42e0d

Observation 2d765e73-8eea-41de-9bc5-055e8bb36032 · outbound

This paper cites Schwing, and Alexander Kirillov.

Native Segmentation Vision Transformers Schwing, and Alexander Kirillov

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:49.604488Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:35.735251Z digest=sha256:de211e101fe8b805b9011eee189535609808ab78f02d424b8edc4e13e77862fa

Observation 6f38e803-886f-4711-aeef-f7b547b01593 · outbound

This paper cites Slic superpixels compared to state-of-the-art superpixel methods.

Native Segmentation Vision Transformers Slic superpixels compared to state-of-the-art superpixel methods

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:49.402810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:35.837086Z digest=sha256:162d3094f7c70f06f0c32a098ad8122557735b7574290ed43a99d0e6822a05b0

Observation 8bd76a18-d8a2-42f5-999d-27a3bca78879 · outbound

This paper cites Object-centric learning with slot attention.

Native Segmentation Vision Transformers Object-centric learning with slot attention

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:49.233922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:35.935254Z digest=sha256:72a5c74f5aa6bdc2fe80db3b7af008e3db3f89ed4a4249299feba5511548dac1

Observation 2fa49496-b243-4745-a451-4f5bb7571d4b · outbound

This paper cites Mean shift: A robust approach toward feature space analysis.

Native Segmentation Vision Transformers Mean shift: A robust approach toward feature space analysis

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:49.016911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:36.074370Z digest=sha256:34a07d7a11ced10116d617abc209d264d0dcd50c9154dea627a57aee60ff7dd4

Observation f1abbd73-3554-4b10-82bc-659c6ddb72ae · outbound

This paper cites Felzenszwalb and Daniel P.

Native Segmentation Vision Transformers Felzenszwalb and Daniel P

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:48.853816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:36.168668Z digest=sha256:0e92a3eaccc7e8a5718794d09add40f9bc4ac9c0a6fed18d99c616f141174c19

Observation b6cc0892-2632-4d79-a46b-dea5645c87d8 · outbound

This paper cites Contour detection and hierarchical image segmentation.

Native Segmentation Vision Transformers Contour detection and hierarchical image segmentation

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:48.710783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:36.279399Z digest=sha256:d2d6d33e71aa4bbff14b55d64ac335abf3a420e9eff6f213345b0bb62a8a5557

Observation 6909a4cd-4635-4a9c-b02d-8368068fba64 · outbound

This paper cites Seeds: Superpixels extracted via energy-driven sampling.

Native Segmentation Vision Transformers Seeds: Superpixels extracted via energy-driven sampling

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:48.536144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:36.403076Z digest=sha256:25d36692491d465eac153cb5757d3842f4f9aeafffed046070f707740dd1cbf5

Observation 6f8390d9-e7de-427f-ba63-7e7c1684f633 · outbound

This paper cites Semantic understanding of scenes through the ade20k dataset.

Native Segmentation Vision Transformers Semantic understanding of scenes through the ade20k dataset

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:48.395546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:36.473035Z digest=sha256:879e82b75bb253496aa36678cb9a4a4bcf5f077fd1e38455fbfeb1e4590a7ab8

Observation 22559590-2a4f-410d-bdfa-744dd9271f2f · outbound

This paper cites Panoptic segmentation.

Native Segmentation Vision Transformers Panoptic segmentation

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:48.202978Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:36.569923Z digest=sha256:e71b1be55ffaeebdc7dde8cdbb34d942b13a160cd35415673ff6187e5d7e6cb5

Observation cc44dad6-1121-4f2f-a77c-7b88b2bded88 · outbound

This paper cites Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position.

Native Segmentation Vision Transformers Neocognitron: A self-organizing neural network model for a mechanism of pattern recognition unaffected by shift in position

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:36.693822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:36.693822Z digest=sha256:1cd134b73413c73add4e40d9f7ffdcf80cf00665031d741a52fe66c422374284

Observation 60d9cf58-f5f8-4f77-9266-976677bb21c0 · outbound

This paper cites Gradient-based learning applied to document recognition.

Native Segmentation Vision Transformers Gradient-based learning applied to document recognition

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:48.038344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:36.746616Z digest=sha256:b9044639cce0d97182911822fd47c59825b3532eb9062d6955948c37df64a45e

Observation 11364c9b-1f78-438f-add2-63c1c7732b99 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Native Segmentation Vision Transformers An image is worth 16x16 words: Transformers for image recognition at scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:36.825137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:36.825137Z digest=sha256:4d02092f9bf066b83a1bb260895c987af36319bf536e10c526a2c0c8ec753061

Observation edb1c32d-38b4-47cb-b971-d50f5cdb5fe7 · outbound

This paper cites A convnet for the 2020s.

Native Segmentation Vision Transformers A convnet for the 2020s

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:36.922349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:36.922349Z digest=sha256:5b327742c171423425fdb83a0620f1177fea697d83bbd8450ef1c5fbfb46e58b

Observation 0350cb82-c86e-4800-bd90-f7e8a3f9b611 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

Native Segmentation Vision Transformers SAM 2: Segment Anything in Images and Videos

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:37.004068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:37.004068Z digest=sha256:004590246257bd0fd939985ba0df7b7692ff8b4728d0a661e684a41f296c7806

Observation 9851423e-6728-49b0-a99d-6c901950ccb1 · outbound

This paper cites Fully convolutional networks for semantic segmentation.

Native Segmentation Vision Transformers Fully convolutional networks for semantic segmentation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:37.099662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:37.099662Z digest=sha256:72e727094c2689f83ffa766d71d88ffb3fe6ce4302a92cd468c07ac6bcf3e3bd

Observation 8e70d7bb-005f-436d-adb4-ffaa0f5b9655 · outbound

This paper cites U-net: Convolutional networks for biomedical image segmentation.

Native Segmentation Vision Transformers U-net: Convolutional networks for biomedical image segmentation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:47.850044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:37.172521Z digest=sha256:366cc5e435fb1e8e118acb467190492b581cf180170c76e1b4607a6e0e1a45b8

Observation 7296edd4-bfdd-41fa-b45b-d1ef3d797e07 · outbound

This paper cites Rich feature hierarchies for accurate object detection and semantic segmentation.

Native Segmentation Vision Transformers Rich feature hierarchies for accurate object detection and semantic segmentation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:47.634897Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:37.259829Z digest=sha256:f2a844e7dd9710bcbce819109ac8da1331f38f013a9eb51f2a2b6ae819bbbaf3

Observation 4617a109-394d-4a8d-b17d-2a6863660fa7 · outbound

This paper cites Fast r-cnn.

Native Segmentation Vision Transformers Fast r-cnn

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:37.360904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:37.360904Z digest=sha256:33f9b172c1b9b7eaa57fc1a108c57b780b58ca6128c12f99752331f61c2aec95

Observation abbb1278-5402-40c8-b8a9-a796dda25b68 · outbound

This paper cites End-to-end object detection with transformers.

Native Segmentation Vision Transformers End-to-end object detection with transformers

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:37.466195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:37.466195Z digest=sha256:baba5f1ccaaffbd60696c40302255816cb93386090b1b56f2a5c73bd180c7f55

Observation e95186b4-c75a-4946-8c30-eb2aae10353d · outbound

This paper cites Normalized cuts and image segmentation.

Native Segmentation Vision Transformers Normalized cuts and image segmentation

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:47.416823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:37.548532Z digest=sha256:b66699b73d7c0d0f6ecc9ae30db26a53894b5a8cadde320861d6a33bc8bbcc57

Observation 2a1bbe62-2422-4b2f-9320-d17f37cbd3cb · outbound

This paper cites Multiscale combinatorial grouping.

Native Segmentation Vision Transformers Multiscale combinatorial grouping

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:47.198109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:37.643606Z digest=sha256:379451f35316740fc8d76bfbd1eb62dfcd30e7e659398e296b1b0c2f0e9ee1d8

Observation 46dc9746-5865-474f-b668-83260ca4fe38 · outbound

This paper cites Superpixel samping networks.

Native Segmentation Vision Transformers Superpixel samping networks

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:47.016623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:37.702010Z digest=sha256:71a15f10624f23c95124c0674ef602add1bc56115098335770e9ed1acf5266b5

Observation abda8ddf-d305-4097-8a72-bf103965d096 · outbound

This paper cites an unresolved cited work.

Native Segmentation Vision Transformers Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:55:46.802461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:37.819714Z digest=sha256:56aa617276c179afce028940c3b3daee9390d00727efadd40d8cbe7749bf74f5

Observation db5cc85c-a4f2-48a8-b46e-064f879dc9a9 · outbound

This paper cites an unresolved cited work.

Native Segmentation Vision Transformers Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:55:46.557896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:37.930389Z digest=sha256:491161f3a24f473c2be63292180998e42a3d7a5602741b7eca5f5f26b3805889

Observation 9eca3a2d-053d-4021-b61b-b205ef653898 · outbound

This paper cites Unified perceptual parsing for scene understanding.

Native Segmentation Vision Transformers Unified perceptual parsing for scene understanding

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:46.374569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:38.018959Z digest=sha256:67ddcd11c04d8129a386f70028a2214a69b7775492071fc80db0a2fa2ad7a604

Observation 58de59c6-c80f-492b-a683-8e12dfa77e14 · outbound

This paper cites Berg, and Li Fei-Fei.

Native Segmentation Vision Transformers Berg, and Li Fei-Fei

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:46.201278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:38.112464Z digest=sha256:0db5aef2f4412bf675cb35d6f3355323fa4b2bbad1488b8a4549a84117dfdbed

Observation e4bb2b64-dfeb-44b8-8a9f-32b9a46018a4 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Native Segmentation Vision Transformers Learning transferable visual models from natural language supervision

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:38.182246Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:38.182246Z digest=sha256:22d8b85b841d1385214fe29581b78037226c46d27b76626ab202ae97c9c9d31b

Observation 051e48ba-726e-45b9-b577-5958b6803ef4 · outbound

This paper cites Le, Yun- Hsuan Sung, Zhen Li, and Tom Duerig.

Native Segmentation Vision Transformers Le, Yun- Hsuan Sung, Zhen Li, and Tom Duerig

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:46.027620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:38.254570Z digest=sha256:c12dccfdd85a829835d46bd5b9d4ed47700a5dcea04d3929c9897708d65a6a51

Observation 4c5dfd3e-7978-414e-93c7-3a0c2cfde865 · outbound

This paper cites A simple framework for text-supervised semantic segmentation.

Native Segmentation Vision Transformers A simple framework for text-supervised semantic segmentation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:45.829111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:38.330391Z digest=sha256:a1827c060518dde8813ac05738e74d89540ce596200b716821e0f8c7f25ab096

Observation 522dc38d-724f-4e3d-bd7d-5badc2072f75 · outbound

This paper cites Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning.

Native Segmentation Vision Transformers Conceptual captions: A cleaned, hypernymed, image alt-text dataset for automatic image captioning

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:45.691037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:38.424956Z digest=sha256:9d60a5bf4a58243ad6f68259650ef5b2b390ce7d0ee1e3e43d6b81689adef892

Observation 4b1d9c74-8136-437f-94e8-736a37685071 · outbound

This paper cites Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts.

Native Segmentation Vision Transformers Conceptual 12M: Pushing Web-Scale Image-Text Pre-Training To Recognize Long-Tail Visual Concepts

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:38.530639Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:38.530639Z digest=sha256:ad6cc146651af881526e1cf8d8b6335ecd16b3b4440aa108fa8b74b6b937ae3d

Observation 7663898a-73b5-447d-912e-d66af4100685 · outbound

This paper cites Redcaps: Web-crawled image-text data created by the people, for the people.

Native Segmentation Vision Transformers Redcaps: Web-crawled image-text data created by the people, for the people

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:45.509767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:38.622568Z digest=sha256:f3ea41aa572abe9135d14b08d2c38eb07ab39705f38b5bb1c798e5cda6371674

Observation 421a5750-7cb4-45b2-844d-c2068675156c · outbound

This paper cites Learning to generate text-grounded mask for open-world semantic segmentation from only image-text pairs.

Native Segmentation Vision Transformers Learning to generate text-grounded mask for open-world semantic segmentation from only image-text pairs

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:45.310461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:38.698678Z digest=sha256:f552c70ede2ddfc03170a0fc4bdbeef9052c7f66d6929435039a3adc564c781c

Observation e6c0e57e-0842-4104-93aa-39c716c4ff09 · outbound

This paper cites Everingham, L.

Native Segmentation Vision Transformers Everingham, L

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:45.134327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:38.790705Z digest=sha256:9787a762fd1f462a5018e5e9fb18ac586c701d280a4d91a74aeabfab6a13c79d

Observation 72274a1a-d26c-4238-8e16-a879ee05c245 · outbound

This paper cites The role of context for object detection and semantic segmentation in the wild.

Native Segmentation Vision Transformers The role of context for object detection and semantic segmentation in the wild

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:44.877729Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:38.870638Z digest=sha256:6fb4547a81102ddf9864b4e69d3950a0bf57e9a4dc742278600dc9a2c2132e04

Observation 12ff2374-85bd-42dc-a2a2-3e24c9307cc2 · outbound

This paper cites an unresolved cited work.

Native Segmentation Vision Transformers Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:55:44.748497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:38.933552Z digest=sha256:5fb70eeb64920d9c85f4f22383ec4a310506bb6cd903d4a788674015ffa7948b

Observation 0863d82c-10c4-4e67-9774-4d4e6d0519a9 · outbound

This paper cites Coco-stuff: Thing and stuff classes in context.

Native Segmentation Vision Transformers Coco-stuff: Thing and stuff classes in context

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:44.550601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:39.012943Z digest=sha256:37f2b0e8b5dc3dbf0e6212900f3ad56ff366a75fd4867957b05ec3acee760bfd

Observation 45573f47-e3d0-44e4-acab-46e74982d378 · outbound

This paper cites Cordts, M.

Native Segmentation Vision Transformers Cordts, M

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T14:55:39.058264Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:55:39.058264Z digest=sha256:6812857de9838d94ca77214497f39f38fce5946741239feef4e1479ed0c3d910

Observation 971d4e1d-0360-4075-8853-572bf21cf678 · outbound

This paper cites Open- world semantic segmentation via contrasting and clustering vision-language embedding.

Native Segmentation Vision Transformers Open- world semantic segmentation via contrasting and clustering vision-language embedding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:44.340556Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:39.128618Z digest=sha256:1247b7683f72d3b396a8be2bb90d3c2d6306e6209f426204312b32ca87e8a3dd

Observation c5829d93-cf5c-4493-a113-a528c55e49b4 · outbound

This paper cites Segclip: Patch aggregation with learnable centers for open-vocabulary semantic segmentation.

Native Segmentation Vision Transformers Segclip: Patch aggregation with learnable centers for open-vocabulary semantic segmentation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:44.151077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:39.200859Z digest=sha256:33a5256fcca8405a190f768076d776d30354e7ace8a23e74fa6630d3784ebfb2

Observation e922457d-570b-402d-829f-1ed192d4d53b · outbound

This paper cites Image-text co-decomposition for text-supervised semantic segmentation.

Native Segmentation Vision Transformers Image-text co-decomposition for text-supervised semantic segmentation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:43.994656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:39.283083Z digest=sha256:3868409bdfc6d9be4be92e50d425e847255e195f0979c5558bf407004d54ab28

Observation fab32581-2094-4638-8719-105837d4e6f9 · outbound

This paper cites Viewco: Discovering text-supervised segmentation masks via multi-view semantic consistency.

Native Segmentation Vision Transformers Viewco: Discovering text-supervised segmentation masks via multi-view semantic consistency

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:43.799254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:39.361768Z digest=sha256:d74bd6d2b69d360b2542cdfabb9b21c1e53fa9611135695815d638dd0a5d66b3

Observation 8fc88444-f9a1-41e5-9db4-72413e568c66 · outbound

This paper cites Rewrite caption semantics: Bridging semantic gaps for language-supervised semantic segmentation.

Native Segmentation Vision Transformers Rewrite caption semantics: Bridging semantic gaps for language-supervised semantic segmentation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:43.625993Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:39.421485Z digest=sha256:081f63ec626d589f7a76c4fd514e0f1b5ff0073851c9873f852fcee384deb56b

Observation fc8b093b-d0cd-47b1-a17d-aabcc22eb2f9 · outbound

This paper cites Uncovering prototypical knowledge for weakly open-vocabulary semantic segmentation.

Native Segmentation Vision Transformers Uncovering prototypical knowledge for weakly open-vocabulary semantic segmentation

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:43.458392Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:39.494191Z digest=sha256:fdc72280a5923fb06ae747e600b9e295ac6567a52e12856e73415ccec9cc9371

Observation 6ca380dd-fdbd-4b8c-ad63-df2fcdf1a52a · outbound

This paper cites Efficient inference in fully connected crfs with gaussian edge potentials.

Native Segmentation Vision Transformers Efficient inference in fully connected crfs with gaussian edge potentials

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:43.313744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:39.625717Z digest=sha256:17f2f983d300e2935eab83ed01dedb340c134ebff279dba7629ddf616c6f5d12

Observation 10d3c55d-06db-4062-bd14-7ef921ee45ea · outbound

This paper cites Single-stage semantic segmentation from image labels.

Native Segmentation Vision Transformers Single-stage semantic segmentation from image labels

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:43.172338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:39.659409Z digest=sha256:8662381c80838cda09ce90f3a8fcded2b3a23c2efcb01e3fef2c4dc8d291ffa4

Observation c961662f-da5d-4207-8db6-5a18fcc46496 · outbound

This paper cites Vmamba: Visual state space model.

Native Segmentation Vision Transformers Vmamba: Visual state space model

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:43.049455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:39.699866Z digest=sha256:aa24ee1a542b3ae81610fcf01ed80a5676a1b481f54b4a826c778315abdc7521

Observation dec6cc42-45af-4d31-ac73-3f8f589fe156 · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Native Segmentation Vision Transformers Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:42.871618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:39.776032Z digest=sha256:be26baa111864a309574f896383c9e50237dd0a503312252d634d3149e8fe234

Observation 67f73f93-4e98-4907-9f6a-301595fefb4f · outbound

This paper cites Panoptic feature pyramid networks.

Native Segmentation Vision Transformers Panoptic feature pyramid networks

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:42.750426Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:39.859502Z digest=sha256:9a71571c7968d7590f3b0d9fdfff75d79a656b2e0ca588b3c7aded4f5b840be4

Observation d62ce440-cb23-407d-a27d-786cc8396158 · outbound

This paper cites Segmenter: Transformer for semantic segmentation.

Native Segmentation Vision Transformers Segmenter: Transformer for semantic segmentation

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:42.618951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:39.938874Z digest=sha256:f8e47578c3e3b1ec6be9063659398810c447b2e5472badf7f3f00e7b2e1b4b93

Observation 40dc5a72-0625-40f4-8cc7-38b435006774 · outbound

This paper cites an unresolved cited work.

Native Segmentation Vision Transformers Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:55:42.424557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:39.973272Z digest=sha256:94b35a5a614fd06f2548f26eee32f4e5310c161a9659dcbc068d4ab651a6a600

Observation a3bbe0f1-d2f1-4079-87cf-a4b174ad51be · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

Native Segmentation Vision Transformers Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:42.334089Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:40.084480Z digest=sha256:2d03124ccfabf8e700659017756f0d334c13a549ea4a95056fd6305416227f32

Observation 1f78a695-9c15-4dbe-9b6a-30f1560e0474 · outbound

This paper cites Self-attention with relative position rep- resentations.

Native Segmentation Vision Transformers Self-attention with relative position rep- resentations

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:42.192849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:40.167748Z digest=sha256:06c8888f011ab815f63d0fe5c846b32ae322438c829892ee35b1846b479876bc

Observation 8c28d6ac-aa0c-4cb7-813d-7735c818d99a · outbound

This paper cites Roformer: Enhanced transformer with rotary position embedding.

Native Segmentation Vision Transformers Roformer: Enhanced transformer with rotary position embedding

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:42.054391Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:40.218252Z digest=sha256:7bdddba534a1b1bcdf04c1b6d05efaac242f39757f2801e1e7ee763238e7d444

Observation 437ba9f9-51cf-4272-8968-749e3242de6f · outbound

This paper cites Rotary position embedding for vision transformer.

Native Segmentation Vision Transformers Rotary position embedding for vision transformer

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:41.934310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:40.261863Z digest=sha256:85bb4c693efb9c1be46450a2f6eb7ab1c7f9bfb734b773829ac69dc0830ec074

Observation bdfafc2b-bed4-4ea6-a779-5aaf230de942 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

Native Segmentation Vision Transformers Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:41.770603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:40.305574Z digest=sha256:a9ac07f70586fca7638e203585d9fb653e337c47dc5b909fe9c8b35363e31d41

Observation 3f1e26f5-c66c-415e-bb6b-d6518d192254 · outbound

This paper cites Faster neighborhood attention: Reducing the o(n^2) cost of self attention at the threadblock level.

Native Segmentation Vision Transformers Faster neighborhood attention: Reducing the o(n^2) cost of self attention at the threadblock level

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:41.619633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:40.399294Z digest=sha256:9c1ac7902f5d970bc34f6226c27c1d0cbd6f35db2c9099b840ebc95c530d48c6

Observation 14e5460f-1f9e-4ca8-8d8e-684f2faea1de · outbound

This paper cites Training data-efficient image transformers; distillation through attention.

Native Segmentation Vision Transformers Training data-efficient image transformers; distillation through attention

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:41.448153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:40.479785Z digest=sha256:0bac635be61b4b1f98118d845d1295ac6a88c84b20a0fe481807f24996956412

Observation 8a207c97-2d2a-41fc-b4af-80f60be67ba7 · outbound

This paper cites Weinberger.

Native Segmentation Vision Transformers Weinberger

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:41.352789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:40.578280Z digest=sha256:de7e298b7f424bae580688a68dbf67a27e4b6a369e4d8de9a6106ec1dd2466a4

Observation bfa2a9c9-10fd-4b9b-a662-2997cd397c59 · outbound

This paper cites Empirically, we observe a negligible decrease of downstream classification and segmentation performance, and a significant increase in speed.

Native Segmentation Vision Transformers Empirically, we observe a negligible decrease of downstream classification and segmentation performance, and a significant increase in speed

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:41.181944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:40.651143Z digest=sha256:02fc7f2704a0ec20e4a2d67e0ec9ffb0c8fa460310eed583e442096fce5d0dfa

Observation 92b5b608-2b35-4808-89b9-bab8545c1252 · outbound

This paper cites An image of {CLASS }.

Native Segmentation Vision Transformers An image of {CLASS }

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T14:55:41.063393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:40.725910Z digest=sha256:4eda2cc6b8b8e8b0db35a29d365a1adff321f1416fe5777a6db40a2374782b74

Observation e32f542a-07bc-4bab-b9e8-6397532cc881 · outbound

This paper cites an unresolved cited work.

Native Segmentation Vision Transformers Unresolved cited work

Reference 74

Resolution
unresolved
raw_fallback, observed 2026-08-07T14:55:40.952338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T14:55:40.776768Z digest=sha256:9d259b1fdcce6c40c333f476be80ecbb3d80b683b9667d47b87914910b9b3239

Pith citing papers

Observation ade4aadf-4816-44a2-9198-566b2ec09f15 · inbound

Elastic Attention Cores for Scalable Vision Transformers cites this paper.

Elastic Attention Cores for Scalable Vision Transformers Native Segmentation Vision Transformers

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-13T06:07:22.541489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-13T06:02:40.158866Z digest=sha256:08e8255e0615ca883b0766b58762b91d23c540f43aaa6c184a187046c48bea28