Pith. sign in

Paper Citation Record · LEDGER

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation

As of 12 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2412.14145.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14145 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:28:54.692620Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact2
  • verified fuzzy23
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b5427c4b-9860-4c57-bace-aff9c1397599 · outbound

This paper cites Learning transferable visual models from natural lan- guage supervision,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Learning transferable visual models from natural lan- guage supervision,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.173886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:28:54.531565Z digest=sha256:6785531a36a49cd8d2df89628bf723a92904e839fcdaa2bd59afa5af7be8e349

Observation 50b60ac0-0715-4811-b6fd-3f5bee94b0bd · outbound

This paper cites Scaling open- vocabulary image segmentation with image-level labels,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Scaling open- vocabulary image segmentation with image-level labels,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.163182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:28:54.536071Z digest=sha256:eb8bee1b2d7441e0ffc218352eb35e45c4cd859e2014ea79f6f8f7bc1a0bf7cb

Observation cc893f0c-a7f4-4c98-823d-1eb76eee6c6f · outbound

This paper cites Maskclip: Masked self-distillation advances contrastive language- image pretraining,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Maskclip: Masked self-distillation advances contrastive language- image pretraining,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.153290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:28:54.539928Z digest=sha256:897f88c348e1003cda58ccb898e7fdc6317f7dad2b969f66c35d2ff2f3aed558

Observation 972b8691-b4cd-449e-883a-67f6f17c0bbf · outbound

This paper cites A simple framework for text-supervised semantic seg- mentation,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation A simple framework for text-supervised semantic seg- mentation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.143201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:28:54.543564Z digest=sha256:d93468d3caba64596bac82c9bb36b508048dc9eaf6fc8f0dc76576da5473dbb4

Observation 9c82c757-c7b6-4429-8528-db57f157d503 · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Open-vocabulary semantic segmentation with mask-adapted clip,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.132843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:28:54.547221Z digest=sha256:27943dea56d3a871bafe46929bd78ebe6dec602d508c2abd9247a525e7cfe5af

Observation 992b89fe-45a8-4713-aedf-80ea4459adf4 · outbound

This paper cites Segment anything,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Segment anything,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.551301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.551301Z digest=sha256:94199a094ef835879c1ecf11e7bd3bc64264333ee49ab1a66a73ac8cde19fad8

Observation 8ca79fc6-8334-484c-98f6-9a6c41bd07be · outbound

This paper cites Open-Vocabulary SAM: Segment and Recognize Twenty-thousand Classes Interactively.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Open-Vocabulary SAM: Segment and Recognize Twenty-thousand Classes Interactively

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.555338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.555338Z digest=sha256:69a245cf03b36659826e289d459d80a37b9e2b43c257dd0caa986e173a1f4d70

Observation ffd22855-5026-4349-b6c4-3fceed9484d2 · outbound

This paper cites PosSAM: Panoptic Open-vocabulary Segment Anything.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation PosSAM: Panoptic Open-vocabulary Segment Anything

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.559154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.559154Z digest=sha256:275c9165ec10df9bcdccca97da5dffbb07f141c6d704d70bad44923f785e9d7f

Observation df7f4408-53c8-47de-9c8c-8327f1c0a018 · outbound

This paper cites Diffusion Models for Open-Vocabulary Segmentation.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Diffusion Models for Open-Vocabulary Segmentation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.562982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.562982Z digest=sha256:a3faf7b54c7d56924bff6a08cbac84532cc8a60c2ba54756f10bd96b56db4e88

Observation e605d83f-9f1a-4fbd-ba21-b35340394eb6 · outbound

This paper cites Open-vocabulary panoptic segmentation with text-to- image diffusion models,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Open-vocabulary panoptic segmentation with text-to- image diffusion models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.116402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:28:54.566808Z digest=sha256:181fba92ec32fe6fee126f9d77030487d84822d10875f250b2d7cd39fc19e357

Observation d6b4f834-4c8d-40e6-ad00-57a957021889 · outbound

This paper cites CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.570344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.570344Z digest=sha256:f21cb6d1ee8bfcab50936955f872085243a8fe93117260018aa927cdefada92d

Observation 1cdb15f9-1aa4-489c-aa98-652d46c7807e · outbound

This paper cites CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.573992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.573992Z digest=sha256:32a15587fd4a0f81d44c918b80445cddc53e092ab1c6fd53522f338ed4bc025b

Observation bcacc359-537f-40ee-ba76-661dedc8a5a1 · outbound

This paper cites CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.577649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.577649Z digest=sha256:045dacc18b8f60886eb7e3cc81d4c31b7fc840f0c197ff534ae660e75574e6fe

Observation 4e5bdf5b-22c8-4752-8bf7-ef7ad3bce46a · outbound

This paper cites What do Vision Transformers Learn? A Visual Exploration.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation What do Vision Transformers Learn? A Visual Exploration

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.581153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.581153Z digest=sha256:5992d527b76288e00bc525b7aca304830040b85ef86dde5cb07076a7c5dc984b

Observation c4766506-8db1-4ba6-a40d-041431829357 · outbound

This paper cites Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.584693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.584693Z digest=sha256:70fc206588034fa4339edaa9ff98ee14ffb74a7572c1fd79fb1874cb03a9f3ab

Observation 3503b708-5e4a-4eda-981d-76a2732ff5c8 · outbound

This paper cites Visualizing and understanding con- volutional networks,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Visualizing and understanding con- volutional networks,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.105885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:28:54.588193Z digest=sha256:0a4d2b24d3ec0fa000d5b1b17aa33a94986502d7ffa0ccb3f8d42f6d6b28fe86

Observation 83ea2eb9-3c3a-410d-8d99-db8510550720 · outbound

This paper cites Masked-attention mask transformer for universal image segmentation,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Masked-attention mask transformer for universal image segmentation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.095809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:28:54.591700Z digest=sha256:b1e97aa78676deeb570be45cbcda9f22f6e61fc44a0124d157080a4a2be4a6ad

Observation 1c0601f6-8278-40c1-8726-6806869e7b6f · outbound

This paper cites kMaX-DeepLab: k-means Mask Transformer.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation kMaX-DeepLab: k-means Mask Transformer

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-11T12:28:54.835331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:28:54.594970Z digest=sha256:8450bdee5caefc455efc02aa60f6f500ab4cf622d704903e9db8aec61862b0f6

Observation d4249441-34d9-4b6a-841e-fc8b08bb3275 · outbound

This paper cites Mean Shift Mask Transformer for Unseen Object Instance Segmentation.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Mean Shift Mask Transformer for Unseen Object Instance Segmentation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.598543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.598543Z digest=sha256:335bdb6c4e94cd0438d575ea096b130412d997e71c5c90c303d9748ab63e04b2

Observation 049da7c3-a12f-492d-b79e-62c805178002 · outbound

This paper cites Feature pyramid networks for object detection,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Feature pyramid networks for object detection,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.085224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:28:54.602047Z digest=sha256:17cf331910c34acadfec10c629ab3fb446a578b9df78ded8c8e8f62087b9ea16

Observation ff4509e9-a618-4a62-8248-819842c682a9 · outbound

This paper cites U-net: Con- volutional networks for biomedical image segmentation,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation U-net: Con- volutional networks for biomedical image segmentation,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.605146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.605146Z digest=sha256:4ddaa2a247ef2b5534b8dbaa724ea11d46b69d2f33bbf7488181d51487bab54e

Observation 6466be89-36ee-4f16-ac21-cd62f1279f0f · outbound

This paper cites Side adapter network for open-vocabulary semantic segmentation,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Side adapter network for open-vocabulary semantic segmentation,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.069221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:28:54.608476Z digest=sha256:4a60e0dfe8e4840ecfd16362ee627c72a3f4c13030895e656db15fc8624fe7a2

Observation b9136e1f-4ba0-40c0-984a-6e0758938a33 · outbound

This paper cites Convolutions die hard: Open-vocabulary segmentation with single frozen convolutional clip,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Convolutions die hard: Open-vocabulary segmentation with single frozen convolutional clip,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.058960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:28:54.611725Z digest=sha256:acb41d3eff2d45a247157538aabffc09b2e8658d49610d78e17495f760918533

Observation d68e1c9f-488b-4ccc-aacd-814675b4002c · outbound

This paper cites SED: A Simple Encoder-Decoder for Open-Vocabulary Semantic Segmentation.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation SED: A Simple Encoder-Decoder for Open-Vocabulary Semantic Segmentation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.615005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.615005Z digest=sha256:1d38bcd4cc8888148eaba2554babe7c15d6e3220434e2e8007caf82f8aecce68

Observation 4ef04c76-04ba-46fb-9400-41d6bd4e275d · outbound

This paper cites Neural discrete representa- tion learning,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Neural discrete representa- tion learning,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.048805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:28:54.618322Z digest=sha256:7da396a2cdbb118f445346c043877665210004f1ff3a34aff8480179a254f4b2

Observation 2cb0daf9-a28b-45c9-85b6-7763eb10a64d · outbound

This paper cites Taming transformers for high-resolution image synthesis,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Taming transformers for high-resolution image synthesis,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.038366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:28:54.621378Z digest=sha256:a5e9d4c5a268abb26da5fdfa8366e5e52fde8516927b8cd0056a45cab22b41f4

Observation f4a46ed8-aeab-453d-ab11-88ecb5bb7917 · outbound

This paper cites Vector-quantized Image Modeling with Improved VQGAN.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Vector-quantized Image Modeling with Improved VQGAN

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.625040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.625040Z digest=sha256:46e840a62a6486adc616ced28a8535d287e3313e598635027e14f493b35f633b

Observation f47ff8d7-6508-41c2-9ea0-dd5cdb03df43 · outbound

This paper cites High-resolution image synthesis with latent diffu- sion models,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation High-resolution image synthesis with latent diffu- sion models,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.628528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.628528Z digest=sha256:f4a3762c8f38dbde9f42fb453780d07dd446cca054cabdef1f09cdb8228c84c2

Observation fbd8280d-ead5-4518-a2ee-fc4bb9d8c747 · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.631621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.631621Z digest=sha256:a158ebd40b70b1abd58f8292cff24341088e7c2b11fe56306c01c51f0c225403

Observation 85aff2cb-5774-4beb-bb3d-de5bc04b0be0 · outbound

This paper cites Coco-stuff: Thing and stuff classes in context,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Coco-stuff: Thing and stuff classes in context,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.021844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:28:54.635233Z digest=sha256:a4e214010ddf40b836309a739d85e1837cf8e5ab5e0f2f4f50107f0cab6d1959

Observation 765bd182-0e1a-48e3-8356-c39d892e1915 · outbound

This paper cites The role of context for object detection and semantic segmentation in the wild,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation The role of context for object detection and semantic segmentation in the wild,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.011252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:28:54.638523Z digest=sha256:358a308e983ffb1b63bea3c4751cf891812c9f14443d245f6d1a3d942d7fca19

Observation 506a3616-5c65-4581-8a3e-4db13a096757 · outbound

This paper cites Scene parsing through ade20k dataset,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Scene parsing through ade20k dataset,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.000841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:28:54.641740Z digest=sha256:f0595626ff396eec3053d7acf9830c2c1154cc5c085badac0f59a311bdfa439a

Observation 1827dfdb-0312-4da4-8d63-403c67c30525 · outbound

This paper cites PeCo: Perceptual Codebook for BERT Pre-training of Vision Transformers.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation PeCo: Perceptual Codebook for BERT Pre-training of Vision Transformers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.644923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.644923Z digest=sha256:8ae9120ccf039a1c1483b7a438e372eb21d9f5b3e12b6944cadda476462f514f

Observation 3e5307f2-b09c-4d91-8dcf-9906fd5db058 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.648555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.648555Z digest=sha256:7c1c90ca7146a0c0e59e147acef8adbdf626b326b9b7bf6e127689fd7f8cefc6

Observation 90ec1cc5-245c-47dc-8e4c-2ea08a419ff8 · outbound

This paper cites Spae: Semantic pyramid autoencoder for multi- modal generation with frozen llms,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Spae: Semantic pyramid autoencoder for multi- modal generation with frozen llms,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:54.984454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:28:54.651853Z digest=sha256:4d396e5d5bcc83aa674c89db9a68782f0c0831b27e2a755e48739913714555e5

Observation c0cc1548-03b4-4e3c-b387-0305a1973e0d · outbound

This paper cites A simple baseline for open-vocabulary semantic segmentation with pre-trained vision-language model,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation A simple baseline for open-vocabulary semantic segmentation with pre-trained vision-language model,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:54.974485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:28:54.655155Z digest=sha256:f773afea68116c3fdef64fc2b206c944c2e35cec33b761ea90d53dfcd298ac75

Observation 0f38c679-ce8a-41a5-9dbd-e37c28aced91 · outbound

This paper cites Denoising Vision Transformers.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Denoising Vision Transformers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.658445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.658445Z digest=sha256:c7fc226ca5b7de96066d313397ea81b5de5639bc9f9e16207b4ebb78d05c708e

Observation abf9a5df-293d-42dd-9316-3926d3dff4ce · outbound

This paper cites FeatUp: A Model-Agnostic Framework for Features at Any Resolution.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation FeatUp: A Model-Agnostic Framework for Features at Any Resolution

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.661887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.661887Z digest=sha256:20ad1c2dd17645ad3aae5c7c4cebe2cb669fd7d738285060d6726d5f2048a132

Observation 9ab15bd0-f713-4ce1-9c03-f32bfa764ad1 · outbound

This paper cites Learning to upsample by learning to sample,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Learning to upsample by learning to sample,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:54.963282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:28:54.665388Z digest=sha256:129cfb4856972f887ca2c3c17bede0b4e5e5c8c3833e0cbefd1c0e5b6970d417

Observation 120bb874-1ed5-416f-bade-6f6ce7ada825 · outbound

This paper cites Unlocking Pre-trained Image Backbones for Semantic Image Synthesis.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Unlocking Pre-trained Image Backbones for Semantic Image Synthesis

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-11T12:28:54.755964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:28:54.668560Z digest=sha256:7d80bb702cc073f3f0da95cf078717e1dc9d9ef4d2da2b38cea1cfa04a240af6

Observation 53127a81-6610-456e-9c5c-2734a18ba502 · outbound

This paper cites Semantic image synthesis with spatially-adaptive normalization,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Semantic image synthesis with spatially-adaptive normalization,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:54.952918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:28:54.671891Z digest=sha256:6d4a678f83d7ff39f1dc9441c911bd58855c6dd86557e8f901fc43202e420055

Observation ec541cbe-96cf-4935-9d8e-1782438f64e6 · outbound

This paper cites Perceptual losses for real-time style transfer and super-resolution,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Perceptual losses for real-time style transfer and super-resolution,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:54.942402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:28:54.675122Z digest=sha256:96690db609093feef9494eb89cc89c3e65372fe569210b180941e0d529f94e5a

Observation b4c18bf1-7bb6-47f5-b9a3-7b845fbb448e · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.678661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.678661Z digest=sha256:302c35f2ddbc4c87684d06d879f64f9171a89bec89197da48debb0dd2fad1082

Observation f17d4c6b-b584-44cb-b1f1-cec17b07b2b9 · outbound

This paper cites Microsoft coco: Common objects in context,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Microsoft coco: Common objects in context,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:54.931400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:28:54.682331Z digest=sha256:d8b063c17917ccbfa6bf3b2106551cfe0e5702c63e8a37eff7fa801bfc69ca64

Observation 3cf76ff3-1008-4d6d-9bee-78d16881e09a · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.685689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.685689Z digest=sha256:b8a07be7658e17202554ee2d150954f09db23cb1dd23b36d0e7571a6f1cb9c35

Observation ac6e7f16-ce8d-4df1-a5d8-7bb70ea2e82d · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Gans trained by a two time-scale update rule converge to a local nash equilibrium,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:54.920975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-11T12:28:54.689247Z digest=sha256:617b16d8b435bb2f372757a426b49b0877e1d13ebe885fe6ff2373f8b70784d5

Observation cb7be813-d7a4-40df-af8d-21964bf99eef · outbound

This paper cites Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.692620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.692620Z digest=sha256:1e0a460e053d13336bbebf09ef834cadc251ae2f7803688fbdeaae04af54efba

Pith citing papers

No inbound Pith citation observations are available.