Pith. sign in

Paper Citation Record · LEDGER

EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2406.20076.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2406.20076 v5

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-08T00:49:55.197773Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-10T08:06:57.772386Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5c63839f-a28c-486d-b7c6-d57edf0e1e06 · inbound

The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding cites this paper.

The ART of Composition: Attention-Regularized Training for Compositional Visual Grounding EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 63

Resolution
verified exact
arxiv_id, observed 2026-05-23T07:25:28.430275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-23T07:24:01.527093Z digest=sha256:3e144d8eb5ecbcefafcc4f67d1fb81f0d973fb93fa3d19532d60e2641100b49b

Observation ba6c847b-5d56-4dda-ab62-a23bebb4c965 · inbound

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos cites this paper.

Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 118

Resolution
verified exact
arxiv_id, observed 2026-05-16T11:39:22.556638Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T11:39:22.340737Z digest=sha256:8cb70969565195d0c520aee398cf42fe78a7bad1f91bfb4fdf1444d613c70c6d

Observation eea956fa-8b55-47b8-904c-11aabd935bef · inbound

SynRES: Towards Referring Expression Segmentation in the Wild via Synthetic Data cites this paper.

SynRES: Towards Referring Expression Segmentation in the Wild via Synthetic Data EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T14:47:26.416689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:47:26.416689Z digest=sha256:f51bfa6207d70a30cddf50c62de9e596b38fff251b8e394b6c450a5f28f835aa

Observation d7f2f6e5-a0fd-4d05-91d9-d1a3c3424bb8 · inbound

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes cites this paper.

SAM2-LOVE: Segment Anything Model 2 in Language-aided Audio-Visual Scenes EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T11:44:12.137620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:44:12.137620Z digest=sha256:0ed548490de4ddbcf1b20b90834673dee43e32d59d33bb9758d0264bbe33c9e9

Observation 6c734f9e-ea84-4ac8-a2e5-80d9356fc318 · inbound

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping cites this paper.

MapleGrasp: Mask-guided Feature Pooling for Language-driven Efficient Robotic Grasping EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:59:29.719733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:59:29.719733Z digest=sha256:137a48c2db8e889b385319ed2a024e7b94872541406249777cc257e556c668a9

Observation cb421d72-2444-4831-b123-426b0b3e7b25 · inbound

Detection of Breast Cancer Lumpectomy Margin with SAM-incorporated Forward-Forward Contrastive Learning cites this paper.

Detection of Breast Cancer Lumpectomy Margin with SAM-incorporated Forward-Forward Contrastive Learning EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T22:40:36.054758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:40:36.054758Z digest=sha256:0b5c5ca37b0f91b50030cb88e0a8cd84be28e5fa806fa4e2dd502e3f6c8ff761

Observation 62c7771d-d80f-4243-ae70-dab1685d75c1 · inbound

CRISP-SAM2: SAM2 with Cross-Modal Interaction and Semantic Prompting for Multi-Organ Segmentation cites this paper.

CRISP-SAM2: SAM2 with Cross-Modal Interaction and Semantic Prompting for Multi-Organ Segmentation EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 126

Resolution
unresolved
no resolver link, observed 2026-08-06T21:54:43.687107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:54:43.687107Z digest=sha256:916c154aa94676dd0fc70fdc392e36613316166775e075af11d22dee9a1c6087

Observation f18181be-c505-4aa3-bef9-046aac89f670 · inbound

SAM 3: Segment Anything with Concepts cites this paper.

SAM 3: Segment Anything with Concepts EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 161

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:25:11.420980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-17T20:22:46.220021Z digest=sha256:6a4c4a35bcf8616679dbf7f6b6595df872279a7f6544e95701193d69f76a9592

Observation 276b06cd-f6f2-4a4f-a6f1-143db866ae1c · inbound

Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation cites this paper.

Tarot-SAM3: Training-free SAM3 for Any Referring Expression Segmentation EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 63

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T05:21:01.008627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T18:11:13.376684Z digest=sha256:68ce1989c72b50b7d071c40cefdce650d706b31a5c6a16106b5b12d8f12b5cec

Observation 0aaeb7f0-5d2e-466a-96bd-bdc82a654fc6 · inbound

AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation cites this paper.

AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:43:49.426512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T05:10:44.608959Z digest=sha256:6a97a342829757529ccb2713b3e58d920aefd66038ac93dffd19c205b3943e60

Observation ad041107-2c10-4d15-b05d-cf91ba74fd24 · inbound

SIGMA-ASL: Sensor-Integrated Multimodal Dataset for Sign Language Recognition cites this paper.

SIGMA-ASL: Sensor-Integrated Multimodal Dataset for Sign Language Recognition EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 103

Resolution
verified exact
arxiv_id, observed 2026-05-11T21:06:11.352200Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-08T07:04:47.169148Z digest=sha256:d48868f197285c97491f11aa215a02f9015d47e2b056217bddf8d3a95f1534b3

Observation fe24e2de-87e2-49c2-b47e-0b30bf539301 · inbound

Vision Harnessing Agent for Open Ad-hoc Segmentation cites this paper.

Vision Harnessing Agent for Open Ad-hoc Segmentation EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:04.383299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T05:52:40.429412Z digest=sha256:7d1c4071f7361ca19733f2d98198303c07dfdcde4a98d36bba6549acc83531ea

Observation c36a9fb7-f75a-4e05-bf97-86dff3564869 · inbound

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction cites this paper.

SetCon: Towards Open-Ended Referring Segmentation via Set-Level Concept Prediction EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 68

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:28:05.240709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-20T05:23:32.202297Z digest=sha256:4db0aa000477700ab4dc3c39e9198cec02f9f297037de798b4af26bf4c646ff6

Observation 66b2baea-c095-4da5-8163-0da46a97be86 · inbound

IQA-Spider: Unifying Multi-Granularity Image Quality Assessment with Reasoning, Grounding and Referring cites this paper.

IQA-Spider: Unifying Multi-Granularity Image Quality Assessment with Reasoning, Grounding and Referring EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 13

Resolution
metadata mismatch
arxiv_id, observed 2026-06-30T13:34:40.563286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-30T13:28:01.889147Z digest=sha256:b6dd395cc7f8050a6c46ca423757a6efb225bffe0941cdbd7c437e4746760988

Observation 0414d691-0d81-40c8-b5b1-36ef31c4b4b9 · inbound

InstructSAM: Segment Any Instance with Any Instructions cites this paper.

InstructSAM: Segment Any Instance with Any Instructions EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-06-29T22:44:02.029740Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:34:00.442420Z digest=sha256:c6a899d80dce5cefe311c8d14773c9490e3b4ed2d07752690a32fc4fa71dc2ae

Observation 94431ffb-eb37-4d23-8da3-bd26278c199b · inbound

MAOAM: Unified Object and Material Selection with Vision-Language Models cites this paper.

MAOAM: Unified Object and Material Selection with Vision-Language Models EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 121

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T02:16:26.324793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-28T11:08:59.900161Z digest=sha256:9f89d5d9a8b88f7ca4566f2901aea6058b74c9f5d2494f0e60c47bf4ae6bb752

Observation 5523097c-b5eb-4508-b58f-4697d7e5016d · inbound

Attribute Retrieving for Open-Vocabulary Endoscopic Compositional Referring Segmentation cites this paper.

Attribute Retrieving for Open-Vocabulary Endoscopic Compositional Referring Segmentation EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-07-10T08:06:57.773790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-07-10T08:06:03.952787Z digest=sha256:d7c0ec2e54ae140dc1004b7977194ee976ff8c873ac9754799fa53d4b0e44a5b

Observation 843ff5ba-65e2-4102-a554-a43190bc32e3 · inbound

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO cites this paper.

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 76

Resolution
unresolved
no resolver link, observed 2026-07-14T04:38:05.237334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:38:05.237334Z digest=sha256:cedcfba82946aca00d5f1cb25b2ee7d28639cbae369dc8ec7dbf2b8f0ece2e15

Observation 77b526f0-3e14-49f8-a679-e6cd4bd18095 · inbound

CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation cites this paper.

CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-06T00:48:46.730648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:48:46.730648Z digest=sha256:af309a7a3b0c091fe04d8896ea31d26a824b7927d6ec553b1a15fd08abf3efcb

Observation 2b8b53ac-3859-4730-91a6-9c76df07f9e2 · inbound

CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation cites this paper.

CROSS: Cascaded Distillation and Dual-Constraint Grounding for Remote Sensing Referring Segmentation EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-08T00:49:55.197773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-08T00:49:55.197773Z digest=sha256:10b65ea1f748598368052caf0169312fcb85c01d12dba9a66bf34d996ad80767