Pith. sign in

Paper Citation Record · LEDGER

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation

As of 11 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2412.14145.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.14145 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T12:28:54.692620Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact2
  • verified fuzzy23
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b5427c4b-9860-4c57-bace-aff9c1397599 · outbound

This paper cites Learning transferable visual models from natural lan- guage supervision,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Learning transferable visual models from natural lan- guage supervision,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.173886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T12:28:54.531565Z digest=sha256:a1e62ccc06978ce921da701d5bb4b788b02f38b0ad2f1e5798910e0e7b3b76af

Observation 50b60ac0-0715-4811-b6fd-3f5bee94b0bd · outbound

This paper cites Scaling open- vocabulary image segmentation with image-level labels,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Scaling open- vocabulary image segmentation with image-level labels,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.163182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T12:28:54.536071Z digest=sha256:f15f31966112b00e1835f87fe865cebe7c9852928b6c2e6fd3f376da798adf1c

Observation cc893f0c-a7f4-4c98-823d-1eb76eee6c6f · outbound

This paper cites Maskclip: Masked self-distillation advances contrastive language- image pretraining,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Maskclip: Masked self-distillation advances contrastive language- image pretraining,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.153290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T12:28:54.539928Z digest=sha256:2f8ed832e5df8d2c8e30f5eb14ca09121d2c341292065b0410c9a053984c7d9e

Observation 972b8691-b4cd-449e-883a-67f6f17c0bbf · outbound

This paper cites A simple framework for text-supervised semantic seg- mentation,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation A simple framework for text-supervised semantic seg- mentation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.143201Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T12:28:54.543564Z digest=sha256:24c236c828c1121d6d2846fdb79274de226b9dd02dd603e63d46f3d55eb54d57

Observation 9c82c757-c7b6-4429-8528-db57f157d503 · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Open-vocabulary semantic segmentation with mask-adapted clip,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.132843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T12:28:54.547221Z digest=sha256:a4d3f57cfe2be1e24dad46a448f76b8a05c7c7fbeb866ee7f5aeee68ad9e0078

Observation 992b89fe-45a8-4713-aedf-80ea4459adf4 · outbound

This paper cites Segment anything,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Segment anything,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.551301Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.551301Z digest=sha256:94199a094ef835879c1ecf11e7bd3bc64264333ee49ab1a66a73ac8cde19fad8

Observation 8ca79fc6-8334-484c-98f6-9a6c41bd07be · outbound

This paper cites Open-Vocabulary SAM: Segment and Recognize Twenty-thousand Classes Interactively.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Open-Vocabulary SAM: Segment and Recognize Twenty-thousand Classes Interactively

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.555338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.555338Z digest=sha256:69a245cf03b36659826e289d459d80a37b9e2b43c257dd0caa986e173a1f4d70

Observation ffd22855-5026-4349-b6c4-3fceed9484d2 · outbound

This paper cites PosSAM: Panoptic Open-vocabulary Segment Anything.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation PosSAM: Panoptic Open-vocabulary Segment Anything

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.559154Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.559154Z digest=sha256:275c9165ec10df9bcdccca97da5dffbb07f141c6d704d70bad44923f785e9d7f

Observation df7f4408-53c8-47de-9c8c-8327f1c0a018 · outbound

This paper cites Diffusion Models for Open-Vocabulary Segmentation.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Diffusion Models for Open-Vocabulary Segmentation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.562982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.562982Z digest=sha256:a3faf7b54c7d56924bff6a08cbac84532cc8a60c2ba54756f10bd96b56db4e88

Observation e605d83f-9f1a-4fbd-ba21-b35340394eb6 · outbound

This paper cites Open-vocabulary panoptic segmentation with text-to- image diffusion models,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Open-vocabulary panoptic segmentation with text-to- image diffusion models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.116402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T12:28:54.566808Z digest=sha256:198974a0eee65b7f3981932d760d0a508f21f40b7218d4c38f169f476c27a417

Observation d6b4f834-4c8d-40e6-ad00-57a957021889 · outbound

This paper cites CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.570344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.570344Z digest=sha256:f21cb6d1ee8bfcab50936955f872085243a8fe93117260018aa927cdefada92d

Observation 1cdb15f9-1aa4-489c-aa98-652d46c7807e · outbound

This paper cites CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation CLIPSelf: Vision Transformer Distills Itself for Open-Vocabulary Dense Prediction

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.573992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.573992Z digest=sha256:32a15587fd4a0f81d44c918b80445cddc53e092ab1c6fd53522f338ed4bc025b

Observation bcacc359-537f-40ee-ba76-661dedc8a5a1 · outbound

This paper cites CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation CLIP-DINOiser: Teaching CLIP a few DINO tricks for open-vocabulary semantic segmentation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.577649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.577649Z digest=sha256:045dacc18b8f60886eb7e3cc81d4c31b7fc840f0c197ff534ae660e75574e6fe

Observation 4e5bdf5b-22c8-4752-8bf7-ef7ad3bce46a · outbound

This paper cites What do Vision Transformers Learn? A Visual Exploration.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation What do Vision Transformers Learn? A Visual Exploration

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.581153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.581153Z digest=sha256:5992d527b76288e00bc525b7aca304830040b85ef86dde5cb07076a7c5dc984b

Observation c4766506-8db1-4ba6-a40d-041431829357 · outbound

This paper cites Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Deep Inside Convolutional Networks: Visualising Image Classification Models and Saliency Maps

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.584693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.584693Z digest=sha256:70fc206588034fa4339edaa9ff98ee14ffb74a7572c1fd79fb1874cb03a9f3ab

Observation 3503b708-5e4a-4eda-981d-76a2732ff5c8 · outbound

This paper cites Visualizing and understanding con- volutional networks,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Visualizing and understanding con- volutional networks,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.105885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T12:28:54.588193Z digest=sha256:d48f74ac4a97d68f282aae9bffff051718aa18f915e7ff0a4df994579e43ca90

Observation 83ea2eb9-3c3a-410d-8d99-db8510550720 · outbound

This paper cites Masked-attention mask transformer for universal image segmentation,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Masked-attention mask transformer for universal image segmentation,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.095809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T12:28:54.591700Z digest=sha256:dbbbe63b8083243430847f0b7e425db0cb62b2caf80beb872261b48df0726a29

Observation 1c0601f6-8278-40c1-8726-6806869e7b6f · outbound

This paper cites kMaX-DeepLab: k-means Mask Transformer.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation kMaX-DeepLab: k-means Mask Transformer

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-11T12:28:54.835331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T12:28:54.594970Z digest=sha256:2848e3583c4aab57c2ad1a8b8924274d10a777048c97dd03900a6884c0c5b3f8

Observation d4249441-34d9-4b6a-841e-fc8b08bb3275 · outbound

This paper cites Mean Shift Mask Transformer for Unseen Object Instance Segmentation.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Mean Shift Mask Transformer for Unseen Object Instance Segmentation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.598543Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.598543Z digest=sha256:335bdb6c4e94cd0438d575ea096b130412d997e71c5c90c303d9748ab63e04b2

Observation 049da7c3-a12f-492d-b79e-62c805178002 · outbound

This paper cites Feature pyramid networks for object detection,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Feature pyramid networks for object detection,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.085224Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T12:28:54.602047Z digest=sha256:254e67eff03cf88c48efb9e998bf46b14445d5c58f942c9dccde62b0bcc6da56

Observation ff4509e9-a618-4a62-8248-819842c682a9 · outbound

This paper cites U-net: Con- volutional networks for biomedical image segmentation,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation U-net: Con- volutional networks for biomedical image segmentation,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.605146Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.605146Z digest=sha256:4ddaa2a247ef2b5534b8dbaa724ea11d46b69d2f33bbf7488181d51487bab54e

Observation 6466be89-36ee-4f16-ac21-cd62f1279f0f · outbound

This paper cites Side adapter network for open-vocabulary semantic segmentation,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Side adapter network for open-vocabulary semantic segmentation,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.069221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T12:28:54.608476Z digest=sha256:83f8561741cf440e12bb82166cb11847e8c298fccd5fc463f7ccd2ea6b6e7720

Observation b9136e1f-4ba0-40c0-984a-6e0758938a33 · outbound

This paper cites Convolutions die hard: Open-vocabulary segmentation with single frozen convolutional clip,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Convolutions die hard: Open-vocabulary segmentation with single frozen convolutional clip,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.058960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T12:28:54.611725Z digest=sha256:4061ffd056cebfca31bb72300247c313c5ed64e7c266ec16ae753ab0c1a5f972

Observation d68e1c9f-488b-4ccc-aacd-814675b4002c · outbound

This paper cites SED: A Simple Encoder-Decoder for Open-Vocabulary Semantic Segmentation.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation SED: A Simple Encoder-Decoder for Open-Vocabulary Semantic Segmentation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.615005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.615005Z digest=sha256:559687d31106f87e05353d8025d3951f3a3d7968da10c9c43db4c25849b58fdd

Observation 4ef04c76-04ba-46fb-9400-41d6bd4e275d · outbound

This paper cites Neural discrete representa- tion learning,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Neural discrete representa- tion learning,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.048805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T12:28:54.618322Z digest=sha256:94d9b14e1149d05413bbe388954e33cf3cfff9fc713d687da819d82a4425466c

Observation 2cb0daf9-a28b-45c9-85b6-7763eb10a64d · outbound

This paper cites Taming transformers for high-resolution image synthesis,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Taming transformers for high-resolution image synthesis,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.038366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T12:28:54.621378Z digest=sha256:16707cd725973bbd1334183df95b17b8f1b737a4b564a260380a05d283d35dda

Observation f4a46ed8-aeab-453d-ab11-88ecb5bb7917 · outbound

This paper cites Vector-quantized Image Modeling with Improved VQGAN.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Vector-quantized Image Modeling with Improved VQGAN

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.625040Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.625040Z digest=sha256:46e840a62a6486adc616ced28a8535d287e3313e598635027e14f493b35f633b

Observation f47ff8d7-6508-41c2-9ea0-dd5cdb03df43 · outbound

This paper cites High-resolution image synthesis with latent diffu- sion models,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation High-resolution image synthesis with latent diffu- sion models,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.628528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.628528Z digest=sha256:f4a3762c8f38dbde9f42fb453780d07dd446cca054cabdef1f09cdb8228c84c2

Observation fbd8280d-ead5-4518-a2ee-fc4bb9d8c747 · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.631621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.631621Z digest=sha256:a158ebd40b70b1abd58f8292cff24341088e7c2b11fe56306c01c51f0c225403

Observation 85aff2cb-5774-4beb-bb3d-de5bc04b0be0 · outbound

This paper cites Coco-stuff: Thing and stuff classes in context,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Coco-stuff: Thing and stuff classes in context,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.021844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T12:28:54.635233Z digest=sha256:5787c2b42664c784944224991e386e975818f5f8ae30bf64d4114afd4612d112

Observation 765bd182-0e1a-48e3-8356-c39d892e1915 · outbound

This paper cites The role of context for object detection and semantic segmentation in the wild,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation The role of context for object detection and semantic segmentation in the wild,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.011252Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T12:28:54.638523Z digest=sha256:7f3b6af9ca89ebb0c69d6f5fccdb686c0a52c9216375f1f83dd12604191f9d6c

Observation 506a3616-5c65-4581-8a3e-4db13a096757 · outbound

This paper cites Scene parsing through ade20k dataset,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Scene parsing through ade20k dataset,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:55.000841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T12:28:54.641740Z digest=sha256:bf6bb43a9bedbde07b2bedbd1c7a73425400c6ee1ceaa84d84604a08b843ed0d

Observation 1827dfdb-0312-4da4-8d63-403c67c30525 · outbound

This paper cites PeCo: Perceptual Codebook for BERT Pre-training of Vision Transformers.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation PeCo: Perceptual Codebook for BERT Pre-training of Vision Transformers

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.644923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.644923Z digest=sha256:8ae9120ccf039a1c1483b7a438e372eb21d9f5b3e12b6944cadda476462f514f

Observation 3e5307f2-b09c-4d91-8dcf-9906fd5db058 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.648555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.648555Z digest=sha256:7c1c90ca7146a0c0e59e147acef8adbdf626b326b9b7bf6e127689fd7f8cefc6

Observation 90ec1cc5-245c-47dc-8e4c-2ea08a419ff8 · outbound

This paper cites Spae: Semantic pyramid autoencoder for multi- modal generation with frozen llms,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Spae: Semantic pyramid autoencoder for multi- modal generation with frozen llms,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:54.984454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T12:28:54.651853Z digest=sha256:b550912c8ca6482849f1bedcadd47dcded5fae0327f9c8e74c94433361bf9879

Observation c0cc1548-03b4-4e3c-b387-0305a1973e0d · outbound

This paper cites A simple baseline for open-vocabulary semantic segmentation with pre-trained vision-language model,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation A simple baseline for open-vocabulary semantic segmentation with pre-trained vision-language model,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:54.974485Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T12:28:54.655155Z digest=sha256:9337ec2e739b9a54e449739a2932765281ad1b4ea8068f46c4995b496026f407

Observation 0f38c679-ce8a-41a5-9dbd-e37c28aced91 · outbound

This paper cites Denoising Vision Transformers.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Denoising Vision Transformers

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.658445Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.658445Z digest=sha256:c7fc226ca5b7de96066d313397ea81b5de5639bc9f9e16207b4ebb78d05c708e

Observation abf9a5df-293d-42dd-9316-3926d3dff4ce · outbound

This paper cites FeatUp: A Model-Agnostic Framework for Features at Any Resolution.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation FeatUp: A Model-Agnostic Framework for Features at Any Resolution

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.661887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.661887Z digest=sha256:20ad1c2dd17645ad3aae5c7c4cebe2cb669fd7d738285060d6726d5f2048a132

Observation 9ab15bd0-f713-4ce1-9c03-f32bfa764ad1 · outbound

This paper cites Learning to upsample by learning to sample,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Learning to upsample by learning to sample,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:54.963282Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T12:28:54.665388Z digest=sha256:3aa1de5509f291fe23738edcfe0c6cb265b40804f71b5388fef098d518d7784e

Observation 120bb874-1ed5-416f-bade-6f6ce7ada825 · outbound

This paper cites Unlocking Pre-trained Image Backbones for Semantic Image Synthesis.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Unlocking Pre-trained Image Backbones for Semantic Image Synthesis

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-11T12:28:54.755964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T12:28:54.668560Z digest=sha256:17e0de279d4ec43f05cd0f292f917d705b06686b9fe62c008556854bed19e9fa

Observation 53127a81-6610-456e-9c5c-2734a18ba502 · outbound

This paper cites Semantic image synthesis with spatially-adaptive normalization,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Semantic image synthesis with spatially-adaptive normalization,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:54.952918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T12:28:54.671891Z digest=sha256:a320164ba1771dc5254a1519de2a9f9e5db1238c5556eeb040a9b090e371c1a4

Observation ec541cbe-96cf-4935-9d8e-1782438f64e6 · outbound

This paper cites Perceptual losses for real-time style transfer and super-resolution,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Perceptual losses for real-time style transfer and super-resolution,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:54.942402Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T12:28:54.675122Z digest=sha256:542621cf6855d57acabcc374a6fa1ba4dfdeef239792ab1343b31f5bc6abf904

Observation b4c18bf1-7bb6-47f5-b9a3-7b845fbb448e · outbound

This paper cites Microsoft COCO Captions: Data Collection and Evaluation Server.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Microsoft COCO Captions: Data Collection and Evaluation Server

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.678661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.678661Z digest=sha256:302c35f2ddbc4c87684d06d879f64f9171a89bec89197da48debb0dd2fad1082

Observation f17d4c6b-b584-44cb-b1f1-cec17b07b2b9 · outbound

This paper cites Microsoft coco: Common objects in context,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Microsoft coco: Common objects in context,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:54.931400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T12:28:54.682331Z digest=sha256:bb914e4c9505c2613405fe52764e7931cb6993ad52d0cd4a8f012b9bf8b4263f

Observation 3cf76ff3-1008-4d6d-9bee-78d16881e09a · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.685689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.685689Z digest=sha256:b8a07be7658e17202554ee2d150954f09db23cb1dd23b36d0e7571a6f1cb9c35

Observation ac6e7f16-ce8d-4df1-a5d8-7bb70ea2e82d · outbound

This paper cites Gans trained by a two time-scale update rule converge to a local nash equilibrium,.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Gans trained by a two time-scale update rule converge to a local nash equilibrium,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T12:28:54.920975Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-11T12:28:54.689247Z digest=sha256:10028150024280d91190b8191cc68b3395e72515ec47e435526acf919f11a8d8

Observation cb7be813-d7a4-40df-af8d-21964bf99eef · outbound

This paper cites Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction.

Incorporating Feature Pyramid Tokenization and Open Vocabulary Semantic Segmentation Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-11T12:28:54.692620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T12:28:54.692620Z digest=sha256:1e0a460e053d13336bbebf09ef834cadc251ae2f7803688fbdeaae04af54efba

Pith citing papers

No inbound Pith citation observations are available.