Pith. sign in

Paper Citation Record · LEDGER

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation

As of 11 August 2026, this Paper Citation Record lists 83 of 83 outbound references and 0 inbound Pith citation observations for arXiv:2501.09688.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.09688 v2

Coverage vector

measured 83 of 83 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T19:48:36.016883Z

measured 83 of 83 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

83 of 83 outbound references displayed

  • verified exact3
  • verified fuzzy49
  • unresolved31
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 090ca234-56dd-44bb-9fbb-8c4ed0a49cd8 · outbound

This paper cites Label-embedding for image classification.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Label-embedding for image classification

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:34.551388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:34.551388Z digest=sha256:66996c4de5e394537ac9f8c665d38ff9dd4c220e3f4d7efba3ce262554a5be9e

Observation 3e10791d-adac-4413-a546-65a5294568ef · outbound

This paper cites Evaluation of output embeddings for fine-grained image classification.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Evaluation of output embeddings for fine-grained image classification

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:34.556685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:34.556685Z digest=sha256:58c4c27871037d34bdc08dfa970de014dad20100d46c4b39a3984e32996d6b88

Observation 2f808f93-d0aa-44e9-acd8-285a5bb19532 · outbound

This paper cites Paco: a novel procrustes application to co- phylogenetic analysis.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Paco: a novel procrustes application to co- phylogenetic analysis

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:34.561748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:34.561748Z digest=sha256:ce4945d77d437ac10430e183683a59d124edadc16a1e7579b666f26f5bc20198

Observation 6cbbd2d7-74a0-401f-b102-f9b4038c525f · outbound

This paper cites Zero-shot semantic segmentation.Advances in Neural Information Processing Systems, 32, 2019.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Zero-shot semantic segmentation.Advances in Neural Information Processing Systems, 32, 2019

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:39.564928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:34.567215Z digest=sha256:7c50b0c32bf4e2b8bfb1dd8d831ef9e9fb45c070a8fdac4027a31185e9ccfd5d

Observation 65078026-410d-4c78-a478-67f154958cfa · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Emerg- ing properties in self-supervised vision transformers

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:39.480034Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:34.571723Z digest=sha256:b30ce2c5df037883a96a5028666b7f62fe6865ca0b97de5f0aa00b9a7c17c664

Observation e5294c37-e609-4e46-8a3c-cf8095d8f3e0 · outbound

This paper cites Rethinking Atrous Convolution for Semantic Image Segmentation.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Rethinking Atrous Convolution for Semantic Image Segmentation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:34.576387Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:34.576387Z digest=sha256:f2b0a3893404104215ebcb7c4eddfb39822c08335b0d1d24a412cf88039f0888

Observation 85f6eeff-ba27-4a71-b4bd-8cadb962c36b · outbound

This paper cites CostFormer:Cost Transformer for Cost Aggregation in Multi-view Stereo.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation CostFormer:Cost Transformer for Cost Aggregation in Multi-view Stereo

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:34.582607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:34.582607Z digest=sha256:df7f2eebdb6cea532a02aac6f5e3602dba1ef65359f11aad3041186ef055a69a

Observation b16815cb-f87c-422f-9d19-db49140c377a · outbound

This paper cites Detect what you can: Detecting and representing objects using holistic mod- els and body parts.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Detect what you can: Detecting and representing objects using holistic mod- els and body parts

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:34.587914Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:34.587914Z digest=sha256:362eba31f40c8434d6ba441e4939e588d5e4bc4734eabb92f9c06b94a73c206f

Observation 960bfbdc-2388-4ebf-b73c-ae8aff1459b1 · outbound

This paper cites Per- pixel classification is not all you need for semantic segmen- tation.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Per- pixel classification is not all you need for semantic segmen- tation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:39.454214Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:34.593114Z digest=sha256:8875f97ae13dd0911325f929d061815e5d12c7faa472f7632cb64fdd454718a9

Observation 017410d7-a138-4a38-a0e6-e682bb6ec7a4 · outbound

This paper cites Masked-attention mask transformer for universal image segmentation.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Masked-attention mask transformer for universal image segmentation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:39.420381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:34.597669Z digest=sha256:f99fc99d585976cf0813bd7df700f6277161ba9fe548c1f0eb0041c7fc804d4f

Observation 3c8c3fda-25fd-4659-92ca-def64e05f723 · outbound

This paper cites Cats: Cost aggre- gation transformers for visual correspondence.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Cats: Cost aggre- gation transformers for visual correspondence

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:39.318158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:34.602289Z digest=sha256:04e1893c99c9fa2e31d0614d773e9463b40a06358af8c82624c9d2104475e308

Observation 42b2915f-ac15-49a0-b8f1-90eabb1b2121 · outbound

This paper cites Cats++: Boosting cost aggregation with convolutions and transformers.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Cats++: Boosting cost aggregation with convolutions and transformers

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:39.156086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:34.606902Z digest=sha256:a9926bb2329faf915c80cfaec76d84f2b243eba2260afe6ce151d32b4a74f544

Observation 258a8713-8db5-4b25-a073-76d7ab53602c · outbound

This paper cites CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation CAT-Seg: Cost Aggregation for Open-Vocabulary Semantic Segmentation

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:34.611070Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:34.611070Z digest=sha256:296e23cf7d4c708b8033d6e2df371130dff47c25ec39f20aaf27c292a829abd3

Observation 090ff2d1-46c1-453f-81ae-81c11ef56891 · outbound

This paper cites Understanding Multi-Granularity for Open-Vocabulary Part Segmentation.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Understanding Multi-Granularity for Open-Vocabulary Part Segmentation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-10T19:48:36.359149Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:34.615802Z digest=sha256:3bda4e8c0a513b8b217971db1105fb3beb65d9498e9da0089424f84740445b61

Observation c80a8c29-2b16-4309-95c7-f85850fd7924 · outbound

This paper cites Unsupervised part discovery from con- trastive reconstruction.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Unsupervised part discovery from con- trastive reconstruction

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:39.090221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:34.620404Z digest=sha256:dc5eb363eab5e3ebf7a306fccd3a92c0c0c4d8021cc42d176a9e649c07247a7a

Observation a2df2f78-3101-438a-815d-8fc89ea04d1a · outbound

This paper cites Histograms of oriented gra- dients for human detection.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Histograms of oriented gra- dients for human detection

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:34.625031Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:34.625031Z digest=sha256:78f47aedb3af190cd0f68e1a13734cf3aa4958841c1216a0f801937269f278a5

Observation a8d0c380-4cc9-49ff-83eb-8a7e575ce328 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Imagenet: A large-scale hierarchical image database

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:39.060716Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:34.666573Z digest=sha256:e60d3bb4f41bed00510c187a258b2720c3520cc5cbf849a390cbee865f97050d

Observation 2c478850-7782-4a0c-b1d1-dc5066260f8b · outbound

This paper cites Open-Vocabulary Universal Image Segmentation with MaskCLIP.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Open-Vocabulary Universal Image Segmentation with MaskCLIP

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:34.697077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:34.697077Z digest=sha256:70d53043c65610f2b7cce288980bc338b274f9db2c5cbdaed37f40097ae7866b

Observation 4dc9a23b-7481-4ac6-af53-ceb46f007112 · outbound

This paper cites Generalized jensen- shannon divergence loss for learning with noisy labels.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Generalized jensen- shannon divergence loss for learning with noisy labels

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:39.044496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:34.726678Z digest=sha256:9e7ff7dbe2caee426b7bba04cb627dab042f39bdd9cf812106c27eeaba537c34

Observation 06cadc3c-d79f-4da8-a2ef-29f2059df66a · outbound

This paper cites De- vise: A deep visual-semantic embedding model.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation De- vise: A deep visual-semantic embedding model

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:39.020580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:34.746450Z digest=sha256:0e5b270a773b1a207169b67e25750b31afcc49ecabd44454c9ee000253c91d70

Observation fc469418-4c88-4de4-987b-31bace6db67c · outbound

This paper cites Scal- ing open-vocabulary image segmentation with image-level labels.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Scal- ing open-vocabulary image segmentation with image-level labels

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:38.919498Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:34.769817Z digest=sha256:0fb67945dc0ce9a84e112b4af209e476f0ca3cfaa2b9568530221e8416667610

Observation 6235dc49-b571-43c3-a57b-ae6e211aa201 · outbound

This paper cites Open-vocabulary Object Detection via Vision and Language Knowledge Distillation.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:34.792236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:34.792236Z digest=sha256:141f8b26007270807b6e403d309f4fa50583db190d88d2eb82b074ae7605394a

Observation fdd10945-a73b-4b98-ac2b-1ef077c1edbf · outbound

This paper cites Aˆ 3: Accelerating attention mechanisms in neural networks with approximation.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Aˆ 3: Accelerating attention mechanisms in neural networks with approximation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:38.903395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:34.815372Z digest=sha256:8286e279385bee8c9fe9a14554e22b4e4276f90e994b92039ff1624c8b290c7d

Observation 87b5c4ee-a96b-4463-a708-630252f86d16 · outbound

This paper cites Global knowledge calibration for fast open-vocabulary segmentation.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Global knowledge calibration for fast open-vocabulary segmentation

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:38.886589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:34.839230Z digest=sha256:74a1398e23aac1b4a8053ac4c20ed41b417da7662f09bf0653839705a3dfbce2

Observation a27ce1e4-aef7-4885-b14d-15ea7e063f22 · outbound

This paper cites Partimagenet: A large, high- quality dataset of parts.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Partimagenet: A large, high- quality dataset of parts

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:38.869563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:34.845068Z digest=sha256:5fd544a363eb6c32ab0f4822241a64bff7c47bc9d6f7f8e0e458982bf999e1b6

Observation 36c60660-742d-46e9-a6cc-d3ca80b2c58b · outbound

This paper cites Compositor: Bottom-up clustering and compositing for robust part and object segmentation.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Compositor: Bottom-up clustering and compositing for robust part and object segmentation

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:38.805186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:34.853836Z digest=sha256:bb714e09b1737e952cfb0701c83ff97e9e53dcfb2fd3470ab9ead1321b5f189c

Observation 70df0c52-92af-403e-863d-ad6011ba1586 · outbound

This paper cites Deep residual learning for image recognition.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Deep residual learning for image recognition

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:34.865219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:34.865219Z digest=sha256:b1892e261e15895afa38c2264a441eb1aee099684e08d9ee188f9fd4aaa6d150

Observation 18c90947-ced4-4015-9b1a-bebec4c8ff41 · outbound

This paper cites Mask r-cnn.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Mask r-cnn

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:38.682497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:34.892492Z digest=sha256:941d4c4255f843ec728785079094c1177a12eb46fd2b31d78153b3e30a07a8ef

Observation 7daf053a-b5a3-4517-8205-b2b6b1338d37 · outbound

This paper cites Cost aggregation with 4d convolutional swin transformer for few-shot segmentation.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Cost aggregation with 4d convolutional swin transformer for few-shot segmentation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:38.626851Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:34.906538Z digest=sha256:6ee52f420b87c5534cc8fbefd16f5d41dc24754fac0a069d1b34df1ccdbd6d96

Observation 3d87664e-c3e1-4449-809e-564d669b56c6 · outbound

This paper cites Unifying feature and cost aggregation with transformers for semantic and visual correspondence.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Unifying feature and cost aggregation with transformers for semantic and visual correspondence

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:38.610800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:34.931238Z digest=sha256:f8ed0022a4db03b8c4c1af6ef31e82109752592edbabd0c7576b15679c1324fb

Observation f83b5638-299d-49c2-a4a6-d65f736e96ca · outbound

This paper cites Fast cost-volume filtering for visual correspondence and beyond.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Fast cost-volume filtering for visual correspondence and beyond

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:38.594462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:34.949337Z digest=sha256:2751bbbed6730cc47d98f37f2f6b1b87bddcf84cdf8bb08ab49007b64c9f451d

Observation c9c64778-b3ea-45c8-8585-60d622656e39 · outbound

This paper cites Scops: Self-supervised co-part segmentation.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Scops: Self-supervised co-part segmentation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:34.956760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:34.956760Z digest=sha256:170a7eba63c68db0f491d76997eb8270fce128e351b0c8dce23094b20e12d41c

Observation b46d92f2-81db-469c-bf86-f67cab1766e1 · outbound

This paper cites Scaling up visual and vision-language representa- tion learning with noisy text supervision.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Scaling up visual and vision-language representa- tion learning with noisy text supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:34.968102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:34.968102Z digest=sha256:d9f5da4b1198dbc19be2c8a18dfd05cb3cf7f2b62c11d7abe37bcb5528b70683

Observation 2111dec3-a319-4d3a-bfc2-e79983edec3d · outbound

This paper cites Salad: Part-level latent diffusion for 3d shape gen- eration and manipulation.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Salad: Part-level latent diffusion for 3d shape gen- eration and manipulation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:38.526142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:34.973831Z digest=sha256:3902718e7d7980389cb00365f2b58a8b62ade5a69bf5e4152e8046880f5220cf

Observation fa1371b3-37bf-49f7-8c91-1d4948253fcc · outbound

This paper cites Language-driven Semantic Segmentation.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Language-driven Semantic Segmentation

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:34.980861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:34.980861Z digest=sha256:e1508d1625091c86bcbe02fb9a849be4aff2bd6dc61ccb2bb9de2467c28c9765

Observation dca016f8-153a-4fcf-b9dc-9a88c5d6a78a · outbound

This paper cites Mask dino: Towards a unified transformer-based framework for object detection and segmentation.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Mask dino: Towards a unified transformer-based framework for object detection and segmentation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:38.427123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:34.986851Z digest=sha256:3e1a83b365e7b5a24786d9b0d6cb806264ef0677fab3365dd62a7e897ed1e42b

Observation 84d41189-f081-49c7-894d-9a08336ca5af · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:38.371041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:34.992238Z digest=sha256:f9881ce7d5812e352bb3e600164bb4ecab11158e4e1ea7f41ea7653e92db8aa5

Observation 412c3d69-cee2-4418-aca0-8ab265359378 · outbound

This paper cites PartGLEE: A Foundation Model for Recognizing and Parsing Any Objects.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation PartGLEE: A Foundation Model for Recognizing and Parsing Any Objects

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-08-10T19:48:36.240407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:34.997403Z digest=sha256:95523752c896d036ab88d9edfc6281003d211ff31a7dc93c6351845595e269f7

Observation db745552-0884-4e75-b567-b48044445785 · outbound

This paper cites A Closer Look at the Explainability of Contrastive Language-Image Pre-training.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation A Closer Look at the Explainability of Contrastive Language-Image Pre-training

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:35.003827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:35.003827Z digest=sha256:9d87d503edc8fe93aeb8fd3d23e79a95a4abbb19f875d18bb6c7952df7186497

Observation fd663b1c-fa4d-4d12-a5b1-c3354e95722d · outbound

This paper cites Open-vocabulary semantic segmentation with mask-adapted clip.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Open-vocabulary semantic segmentation with mask-adapted clip

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:35.009952Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:35.009952Z digest=sha256:7ffd37f74437f555d5fd46c1d3e0248002e7c40cd4ef19a1c05744b8ff7abc7d

Observation 211d525b-f2e8-46a8-9a14-e31c3955d583 · outbound

This paper cites Divergence measures based on the shannon en- tropy.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Divergence measures based on the shannon en- tropy

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:35.014541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:35.014541Z digest=sha256:493735955f283f9b77081e143ac0a6b3c319cf46b714ae16fa048f6cf4f503d4

Observation 3fdcd5fb-6d99-41d8-8102-35d80a7dd0ed · outbound

This paper cites Editgan: High-precision semantic image editing.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Editgan: High-precision semantic image editing

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:38.336701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:35.019911Z digest=sha256:3d0114708b78ef1088147268a82478c8190347b27d1ee624476ffa6ae82430aa

Observation 43330f00-7824-4ae3-80a7-b36bf3d4c7fa · outbound

This paper cites Sift flow: Dense correspondence across scenes and its applications.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Sift flow: Dense correspondence across scenes and its applications

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:35.034819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:35.034819Z digest=sha256:52c974dd1b75ac0bdf01699ed313c68e0023da43d2117ab2ea321cddaf19d9ff

Observation ed449e01-2df5-4761-b57c-dbbc53ed5424 · outbound

This paper cites Se- mantic correspondence as an optimal transport problem.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Se- mantic correspondence as an optimal transport problem

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:38.311124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:35.040458Z digest=sha256:9a5911b3178684b6abb4fca76e54087e0ad7bd4d7a4fb24690ec30ca71fe4d91

Observation 202b5526-63d1-4c6e-8ce2-e1a74fc9d079 · outbound

This paper cites Open-Vocabulary Segmentation with Semantic-Assisted Calibration.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Open-Vocabulary Segmentation with Semantic-Assisted Calibration

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-08-10T19:48:36.201592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:35.090187Z digest=sha256:2bc3ebbb6e95fa3444ddf0250fdc8ae3a55af249bb95e39fddebd6c67d5cdb87

Observation 611ca277-4f8d-436c-a108-cde8ad692068 · outbound

This paper cites 3d part guided image editing for fine-grained object understanding.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation 3d part guided image editing for fine-grained object understanding

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:38.295859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:35.168028Z digest=sha256:af5d01e7aec6b0b723bb77b07b108ea0d1abdbbcda2d84aaea605d8309bcdf20

Observation 86fd6094-08d6-4e51-943d-c608a4f07b0c · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Swin transformer: Hierarchical vision transformer using shifted windows

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:38.277161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:35.220364Z digest=sha256:b2bd90f2409ebfc7321ce2dbdaed8aedab0ab57c0fc9839c3a81015e908ac208

Observation 7f0f0b9c-914f-44b0-92ac-9df8af4805c5 · outbound

This paper cites Fully convolutional networks for semantic segmentation.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Fully convolutional networks for semantic segmentation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:38.148560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:35.238221Z digest=sha256:18afd95514c136807bdd7753a3251162f7173956bd379d69deafb2cc83616ab2

Observation 66bdb88f-8352-4ee8-a5aa-949b25c137ce · outbound

This paper cites Decoupled Weight Decay Regularization.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Decoupled Weight Decay Regularization

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:35.270590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:35.270590Z digest=sha256:13e7a18365a6afce13e7684285d65bb4c9b250192373171eb36b8b1a391541d0

Observation f19caff0-bc32-4b04-a66a-11d598250310 · outbound

This paper cites Image segmenta- tion using text and image prompts.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Image segmenta- tion using text and image prompts

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:38.051535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:35.284733Z digest=sha256:5d4f4ee490849752d5040fd7ae806bebf870aa86e1a1a0064a628f483919a2c2

Observation 619768ee-e4cb-4441-a17f-6cdbaa304d53 · outbound

This paper cites Wordnet: a lexical database for english.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Wordnet: a lexical database for english

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:38.036443Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:35.290237Z digest=sha256:b5e912d2b1ef68747b7625c80575ebc64c7990af7119b551fed15a9efffbf0ae

Observation 7e2e7a7b-f640-4ab1-a2b8-119db16410c2 · outbound

This paper cites Hyperpixel flow: Semantic correspondence with multi-layer neural features.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Hyperpixel flow: Semantic correspondence with multi-layer neural features

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:38.019994Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:35.295258Z digest=sha256:6722c593e3fa6cf305f8221b9d654e08152fb0ff0ce203b1baf8fb1c20e41369

Observation 80e081cf-dc03-4f6f-bbbe-9f18dc69bb87 · outbound

This paper cites Learning to compose hypercolumns for visual correspon- dence.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Learning to compose hypercolumns for visual correspon- dence

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:37.992362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:35.300333Z digest=sha256:d0ca4471a34c344f99f782a68616cb35f9c6c72890699c3b7fbbcb4e58fed431

Observation 0e219efb-c350-4ad7-bafa-d0f6cdb3e00d · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation DINOv2: Learning Robust Visual Features without Supervision

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:35.305078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:35.305078Z digest=sha256:15e12cbd9849f87fed9485895a4583cde863c48c5a94878d9ddf729e4cb08f1b

Observation 8b02dd13-5452-43aa-b523-22bd10a27920 · outbound

This paper cites To- wards open-world segmentation of parts.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation To- wards open-world segmentation of parts

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:37.952351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:35.309881Z digest=sha256:d8b32f831138f502e6d9b0367d6e50ceb7d8b17e46afa703a459259950f96675

Observation a1d534f8-1caf-4686-b0c4-6bda92071e15 · outbound

This paper cites Computational optimal transport: With applications to data science.Foundations and Trends® in Machine Learning, 11(5-6):355–607, 2019.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Computational optimal transport: With applications to data science.Foundations and Trends® in Machine Learning, 11(5-6):355–607, 2019

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:37.862486Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:35.314296Z digest=sha256:034b510e4ba023095c0c99aa23df7531094b2dd38d6c267f8de820c85f93ee73

Observation b6521947-2918-498f-92ae-d0dbaac1feb9 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Learning transferable visual models from natural language supervi- sion

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:37.711139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:35.318598Z digest=sha256:5f7827e05a6be4dbe26a4a8985799b8b2c28575ea2258a5f4b3ce4a2305e6372

Observation 6d18e1a8-32d5-4ff2-a73b-61630d19d78e · outbound

This paper cites Neighbourhood con- sensus networks.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Neighbourhood con- sensus networks

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:37.695243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:35.402164Z digest=sha256:6f790084bdbb017a7f5437a68aaae6d2348eb344c805ee6162f315ae2f9d8b45

Observation 6e7561bd-6317-4b1a-a42c-ee91e747402c · outbound

This paper cites U- net: Convolutional networks for biomedical image segmen- tation.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation U- net: Convolutional networks for biomedical image segmen- tation

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:35.475790Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:35.475790Z digest=sha256:fc191ee5aaefca9a5fdaa196d1a3641e23e9f5d3169c57715567cf75e1ce5992

Observation 59452721-16ab-4ab2-aeab-4973fa80a5ce · outbound

This paper cites Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Pwc-net: Cnns for optical flow using pyramid, warping, and cost volume

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:35.521113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:35.521113Z digest=sha256:85c6991b7202c2457b3fe14d1324f3016958b37709f39924ab678fb34a1a151d

Observation caaf3bd3-3989-44d0-a44b-fdc2a81be1ff · outbound

This paper cites Going denser with open-vocabulary part segmentation.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Going denser with open-vocabulary part segmentation

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:37.603310Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:35.528808Z digest=sha256:13327cdcab8411a21fce224cdb34339c679f2b00320a8bdec6a62a3604dbe520

Observation 8f782bcd-1d60-44cd-8084-b5685af8d568 · outbound

This paper cites Parts and wholes in face recognition.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Parts and wholes in face recognition

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:37.407465Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:35.548344Z digest=sha256:343372afc3a3313322ceaf8858ba233b4d1b5596dc34866edf134917a1d41b77

Observation 4a98ab58-959a-4065-a6e0-c904a9e51181 · outbound

This paper cites Glu- net: Global-local universal network for dense flow and corre- spondences.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Glu- net: Global-local universal network for dense flow and corre- spondences

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:35.585172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:35.585172Z digest=sha256:195d5ccef64c89c64beb666f9c8d48f215fff263a97d791a902d06a3831daa81

Observation 6cbf97a8-48c1-45a1-a3f1-739676d60de5 · outbound

This paper cites Pdisconet: Semantically consistent part discovery for fine-grained recognition.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Pdisconet: Semantically consistent part discovery for fine-grained recognition

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:35.643066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:35.643066Z digest=sha256:a20466e7336ff6ed72c605eaa647c045ff1be3a01a519cff5523ef070c86e865

Observation 3725fade-665e-4dcf-b2d1-5b39d852d46c · outbound

This paper cites Attention is all you need.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Attention is all you need

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:35.670739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:35.670739Z digest=sha256:5ceb926e574d5808fa9cde68b6e28dae2c20a4005bbfaa0c4766f95535b6cbcb

Observation 8ad3e1a2-f5b4-4b3d-9645-349a71d6a855 · outbound

This paper cites Optimal transport: old and new.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Optimal transport: old and new

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:37.294110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:35.675236Z digest=sha256:5626911beb7d2c4ed191b0acad7ff50bc9682142e1393a4386771d028e689cb7

Observation 42bf180c-584e-4349-bb11-b2689d7bbf2b · outbound

This paper cites In- structpart: Affordance-based part segmentation from lan- guage instruction.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation In- structpart: Affordance-based part segmentation from lan- guage instruction

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:37.278144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:35.679273Z digest=sha256:0242fa11bd01eac6e33246a1e14150c2b4f875980cd4b17a1ca050a561653e3f

Observation baeead81-f078-4bc4-9c31-1eb858c9192e · outbound

This paper cites Pyramid vision transformer: A versatile backbone for dense prediction without convolutions.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Pyramid vision transformer: A versatile backbone for dense prediction without convolutions

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:37.259446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:35.683582Z digest=sha256:742399d1b908a9f16df97ed476ad5ae4be2a15f2cd356bfc80881bc742012d77

Observation 356ad15a-7c21-48c1-90f3-fc090b51284f · outbound

This paper cites InstanceDiffusion: Instance-level Control for Image Generation.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation InstanceDiffusion: Instance-level Control for Image Generation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:35.687564Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:35.687564Z digest=sha256:623ae4bd7551ab9307d3f5c1a161186b339d4424eed6d9fa2bb73cd01be7fbd9

Observation 0c0411f4-e0f9-42f2-b1ac-38381ad30507 · outbound

This paper cites Ov-parts: Towards open- vocabulary part segmentation.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Ov-parts: Towards open- vocabulary part segmentation

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:37.227174Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:35.691786Z digest=sha256:1b902c865e323b4c4a0cff2d338e554effbcc56e63529b8079fe05534605fff3

Observation da2551fc-7433-4889-bcd0-23e02b1e673b · outbound

This paper cites Semantic projection network for zero-and few-label semantic segmentation.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Semantic projection network for zero-and few-label semantic segmentation

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:37.076600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:35.787678Z digest=sha256:2a46f8582621206fe05e70d22bee55af928c2f8d9f9c8e7101244a2152e5eebb

Observation 73d0a2c2-defd-4361-8643-d1ab7e9d7ffb · outbound

This paper cites SED: A Simple Encoder-Decoder for Open-Vocabulary Semantic Segmentation.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation SED: A Simple Encoder-Decoder for Open-Vocabulary Semantic Segmentation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:35.818590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:35.818590Z digest=sha256:76856417c1b975461b13de609eb1e5736079a188d66a29451f755c2aac7cd8d4

Observation e494ba32-11ca-430c-b152-c5ecf18e8f89 · outbound

This paper cites Open-vocabulary panop- tic segmentation with text-to-image diffusion models.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Open-vocabulary panop- tic segmentation with text-to-image diffusion models

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:35.892353Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:35.892353Z digest=sha256:484479657ce92fc1fff22afca58c77b64c0ceb04f854955fddde130243fed2cc

Observation 7ad2a31a-3111-4423-a6d7-84bc4f46a561 · outbound

This paper cites A simple baseline for open- vocabulary semantic segmentation with pre-trained vision- 11 language model.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation A simple baseline for open- vocabulary semantic segmentation with pre-trained vision- 11 language model

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:36.899779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:35.920269Z digest=sha256:b315bb01a3ba18376aeaac38a7dfbdda7d365a00177133c67e6f54166ad31370

Observation 02c90f8f-228b-48c9-ac0e-3cf840ace960 · outbound

This paper cites HomeRobot: Open-Vocabulary Mobile Manipulation.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation HomeRobot: Open-Vocabulary Mobile Manipulation

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:35.950172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:35.950172Z digest=sha256:fbfa9328d1e8bbcb9932732867cbc5f2e9258bbc7a2ca3e597e4e33bf52ad608

Observation 8f091fff-f0d5-4d06-8878-ec5c753b7a37 · outbound

This paper cites C2fnas: Coarse- to-fine neural architecture search for 3d medical image seg- mentation.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation C2fnas: Coarse- to-fine neural architecture search for 3d medical image seg- mentation

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:36.882710Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:35.984489Z digest=sha256:b7cc68e59c79a5d8243e4f07e55e75428e5326540608e313667de27484788df0

Observation bd5c994d-3bb6-4156-83c4-f5307238b284 · outbound

This paper cites Convolutions die hard: Open-vocabulary seg- mentation with single frozen convolutional clip.Advances in Neural Information Processing Systems, 36, 2024.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Convolutions die hard: Open-vocabulary seg- mentation with single frozen convolutional clip.Advances in Neural Information Processing Systems, 36, 2024

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:36.866331Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:35.989294Z digest=sha256:aa302dc7e3a4c024f714a9b5c348aa0d837e4ad1bf4b5fa2996d8189cc0c736f

Observation bfdf972c-adc4-4cb7-9833-a38f74b2cb8c · outbound

This paper cites Open-vocabulary object detection using captions.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Open-vocabulary object detection using captions

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:36.824494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:35.994288Z digest=sha256:7b1b175c5ece4f98a246e3ff97a8c4dbaec2239f54d6ea1ad9cc64b54fc6cad2

Observation d9dca9e4-dc36-410a-acb5-14709fc7a54f · outbound

This paper cites Open vocabulary scene parsing.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Open vocabulary scene parsing

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:36.729546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:35.999110Z digest=sha256:f2e0b64b82859425c38492e7fb94aa4f20a8ca455f0e276ca06973062788091c

Observation af7668bc-2420-4675-ac2a-d7ae68aca7a6 · outbound

This paper cites Scene parsing through ade20k dataset.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Scene parsing through ade20k dataset

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:36.003829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:36.003829Z digest=sha256:98c6294b5b1835d8918bc0d4c4a5210462d8851c52c6a046a9f53b39da51f209

Observation 77201593-316d-4da1-b232-5debc540a1b0 · outbound

This paper cites Extract free dense labels from clip.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Extract free dense labels from clip

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:36.632220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:36.008198Z digest=sha256:7b109b3e2369ac872ef60cffa073960efff2f24de134a84af2d3a444765aded9

Observation b5d453f8-a3e4-410f-81a9-fa32156e1a75 · outbound

This paper cites Learning to prompt for vision-language models.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation Learning to prompt for vision-language models

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-10T19:48:36.012707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:48:36.012707Z digest=sha256:07dc69a59825d35ce33d77206d974f7415bf2d238b64aff764246cae7ac2ef09

Observation 788bf27f-901c-49ce-9303-eb0147d25a54 · outbound

This paper cites person” (b) “person’s eye.

Fine-Grained Image-Text Correspondence with Cost Aggregation for Open-Vocabulary Part Segmentation person” (b) “person’s eye

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:48:36.543756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-08-10T19:48:36.016883Z digest=sha256:653a367fa159032661505400adc941f06c27f81c079e6e9f20ae2d6d0f0731d0

Pith citing papers

No inbound Pith citation observations are available.