Pith. sign in

Paper Citation Record · LEDGER

EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2406.20076.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.20076 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:49:55.197773Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T08:06:57.772386Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5c63839f-a28c-486d-b7c6-d57edf0e1e06 · inbound

The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding cites this paper.

The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:25:28.430275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-23T07:24:01.527093Z digest=sha256:984599eeda9a186723f91a93003be17d08ab4e49d3e8706346ad7c7ce4bee77a

Observation ba6c847b-5d56-4dda-ab62-a23bebb4c965 · inbound

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos cites this paper.

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:39:22.556638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T11:39:22.340737Z digest=sha256:0da269ef3fe1ae680069395fd46f8340dd4970621df2aa570d8c14a5ab703a79

Observation eea956fa-8b55-47b8-904c-11aabd935bef · inbound

SynRES: Towards Referring Expression Segmentation in the Wild via Synthetic Data cites this paper.

SynRES: Towards Referring Expression Segmentation in the Wild via Synthetic Data EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:26.416689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:26.416689Z digest=sha256:f9b1812444355247d39c962180b9b288e22a41beb6b3eef23d989a7ecec642eb

Observation d7f2f6e5-a0fd-4d05-91d9-d1a3c3424bb8 · inbound

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes cites this paper.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:12.137620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:12.137620Z digest=sha256:6136ea0c4688eb0326b606e05e24b9e8ccc183725960521e190d561d277322e5

Observation 6c734f9e-ea84-4ac8-a2e5-80d9356fc318 · inbound

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping cites this paper.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:29.719733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:29.719733Z digest=sha256:4d5c4827c697acf14d24bb7db83c9814e41f92c92905cb6808d627600dd83eee

Observation cb421d72-2444-4831-b123-426b0b3e7b25 · inbound

Detection of Breast Cancer Lumpectomy Margin with SAM-incorporated Forward-Forward Contrastive Learning cites this paper.

Detection of Breast Cancer Lumpectomy Margin with SAM-incorporated Forward-Forward Contrastive Learning EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T22:40:36.054758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:40:36.054758Z digest=sha256:c5bcc0a73bbcbe3e469921d9dd18d5d5a4852d2afd8125c021c64814781e2fa8

Observation 62c7771d-d80f-4243-ae70-dab1685d75c1 · inbound

CRISP-SAM2: SAM2 with Cross-Modal Interaction and Semantic Prompting for Multi-Organ Segmentation cites this paper.

CRISP-SAM2: SAM2 with Cross-Modal Interaction and Semantic Prompting for Multi-Organ Segmentation EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 126

Resolution
unresolved
no resolver link, observed 2026-08-06T21:54:43.687107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:54:43.687107Z digest=sha256:10e5f68ec6b862ce2f606280112a057d8b23da05cb873e630f82a2c5c83d498f

Observation f18181be-c505-4aa3-bef9-046aac89f670 · inbound

SAM 3: Segment Anything with Concepts cites this paper.

SAM 3: Segment Anything with Concepts EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 161

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.420980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:7f2841eb1c6101b54a6f0c2712a92853ab9a09e8249cf5557c9f4aba4b1fcad1

Observation 276b06cd-f6f2-4a4f-a6f1-143db866ae1c · inbound

Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation cites this paper.

Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:21:01.008627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T18:11:13.376684Z digest=sha256:060b2ffb5cb503542ab159db4f1f84fae4ff5138fb2a700b0beb3b8795418cab

Observation 0aaeb7f0-5d2e-466a-96bd-bdc82a654fc6 · inbound

AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation cites this paper.

AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:43:49.426512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T05:10:44.608959Z digest=sha256:ef120fb89b1d9d8b8ac56b072e03c0a0dc2de3f143d54bf0fddd7574bc659c04

Observation ad041107-2c10-4d15-b05d-cf91ba74fd24 · inbound

SIGMA-ASL: Sensor-Integrated Multimodal Dataset for Sign Language Recognition cites this paper.

SIGMA-ASL: Sensor-Integrated Multimodal Dataset for Sign Language Recognition EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:06:11.352200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-08T07:04:47.169148Z digest=sha256:87e73ed69b31a8222f8a91744db6844ac6db7ed3c2b35c3fc371366f3d7bb36a

Observation fe24e2de-87e2-49c2-b47e-0b30bf539301 · inbound

Vision Harnessing Agent for Open Ad-hoc Segmentation cites this paper.

Vision Harnessing Agent for Open Ad-hoc Segmentation EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:04.383299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T05:52:40.429412Z digest=sha256:5fddfe9321540c43ece3f644999ae4be937c0408c3beeac93d3b4cb89c6631ab

Observation c36a9fb7-f75a-4e05-bf97-86dff3564869 · inbound

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction cites this paper.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:28:05.240709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:6bcb0db97e9ef4e90a8b4c93c7557f74f53324ca2258c570a299d14ce74fc689

Observation 66b2baea-c095-4da5-8163-0da46a97be86 · inbound

IQA-Spider: Unifying Multi-Granularity Image Quality Assessment with Reasoning, Grounding and Referring cites this paper.

IQA-Spider: Unifying Multi-Granularity Image Quality Assessment with Reasoning, Grounding and Referring EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:34:40.563286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T13:28:01.889147Z digest=sha256:2f1fa9af68fd725a723e2b8feb14bddee7c12ab0dab67e61fb1fdc2fdc76743b

Observation 0414d691-0d81-40c8-b5b1-36ef31c4b4b9 · inbound

InstructSAM: Segment Any Instance with Any Instructions cites this paper.

InstructSAM: Segment Any Instance with Any Instructions EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:44:02.029740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T22:34:00.442420Z digest=sha256:3a93b5896c60f84d62c0c648967189e4a3cb3ed9c4134e85f5e22c3026e85cca

Observation 94431ffb-eb37-4d23-8da3-bd26278c199b · inbound

MAOAM: Unified Object and Material Selection with Vision-Language Models cites this paper.

MAOAM: Unified Object and Material Selection with Vision-Language Models EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 121

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:16:26.324793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-06-28T11:08:59.900161Z digest=sha256:c2e22083c2eb5a76246b4c741ede90aa0e9b7966f4028269f2e52432612334e2

Observation 5523097c-b5eb-4508-b58f-4697d7e5016d · inbound

Attribute Retrieving for Open-Vocabulary Endoscopic Compositional Referring Segmentation cites this paper.

Attribute Retrieving for Open-Vocabulary Endoscopic Compositional Referring Segmentation EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-10T08:06:57.773790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-10T08:06:03.952787Z digest=sha256:547db506826f67395ca25c3bd8511a5970cfe8ebb983530167b6f7ced546ad19

Observation 843ff5ba-65e2-4102-a554-a43190bc32e3 · inbound

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO cites this paper.

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-14T04:38:05.237334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:38:05.237334Z digest=sha256:cedcfba82946aca00d5f1cb25b2ee7d28639cbae369dc8ec7dbf2b8f0ece2e15

Observation 77b526f0-3e14-49f8-a679-e6cd4bd18095 · inbound

CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation cites this paper.

CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T00:48:46.730648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:48:46.730648Z digest=sha256:caf201152c1a89ab34b6065b605789f1f89e4be171b73b53029ca8ee40c85311

Observation 2b8b53ac-3859-4730-91a6-9c76df07f9e2 · inbound

CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation cites this paper.

CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-08T00:49:55.197773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:49:55.197773Z digest=sha256:554febecd0640b5c05eda8045b7feb856dcc951b3a4380430865168108369f80