Pith. sign in

Paper Citation Record · LEDGER

EPIC: Efficient Prompt Interaction for Text-Image Classification

As of 12 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2507.07415.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07415 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:46:00.186219Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact2
  • verified fuzzy16
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e14bf6c4-7ead-4b3e-97fe-5da5a5c3fc0c · outbound

This paper cites Cma-clip: Cross- modality attention clip for text-image classification,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Cma-clip: Cross- modality attention clip for text-image classification,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.354240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:46:00.137419Z digest=sha256:36682b792319f0059c93e514f8b129207a2f8af475c7e1a758fe0f89e7d11d13

Observation 33fe7a97-2abf-4ce2-99e4-f5d01c59564e · outbound

This paper cites Future-aware diverse trends framework for recommendation,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Future-aware diverse trends framework for recommendation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.348130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:46:00.140117Z digest=sha256:0e07a63ecf0e86f9e2d0fdee7344f566cd3b8d9d8b360090b5e72cc08f1cce5b

Observation 089606f7-3044-4917-b0a3-f1146e4545a7 · outbound

This paper cites Dense fusion network with multimodal residual for sentiment classification,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Dense fusion network with multimodal residual for sentiment classification,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.341633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:46:00.142283Z digest=sha256:6f3165cba7d9700e7276bb1bf957edd8fa8fd49ab3394c5eb885c9370a06767b

Observation 34939d49-d3ad-44b8-923c-2fc5cf032d3a · outbound

This paper cites Tensor Fusion Network for Multimodal Sentiment Analysis.

EPIC: Efficient Prompt Interaction for Text-Image Classification Tensor Fusion Network for Multimodal Sentiment Analysis

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:46:00.144295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:46:00.144295Z digest=sha256:4bea22d92b5c75521a3b19dbc1bd3c516066eabf253d442417c1f31dc2be26da

Observation 56960134-b3a1-4bc5-b46a-bfe3b619497b · outbound

This paper cites Memory fusion network for multi-view sequential learning,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Memory fusion network for multi-view sequential learning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.335901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:46:00.146948Z digest=sha256:8c339befffb34bc2ca9fc595db6c6ff7a88e5eb168847e49abc911e8a3fb8b51

Observation 2e0162a0-262e-4d25-897c-33e4adf5e804 · outbound

This paper cites Misa: Modality-invariant and-specific representations for multimodal sentiment analysis,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Misa: Modality-invariant and-specific representations for multimodal sentiment analysis,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.330209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:46:00.149114Z digest=sha256:73ec5a7e49caa7ec4ec5cf0c288eb93d5f5ab72c02bfe3ee0b00f3e772608a59

Observation c7805295-3929-48e2-84b1-a43373a9f34d · outbound

This paper cites Centralnet: a multilayer approach for multimodal fusion,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Centralnet: a multilayer approach for multimodal fusion,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.324314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:46:00.151389Z digest=sha256:7b0f041b35b2c8c47a1141db58e28839cef40a12446abd47d5ae87dc0cf90cd1

Observation 2830e517-deba-4ae4-b7cf-f4017f2b3250 · outbound

This paper cites Modular and Parameter-Efficient Multimodal Fusion with Prompting.

EPIC: Efficient Prompt Interaction for Text-Image Classification Modular and Parameter-Efficient Multimodal Fusion with Prompting

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:46:00.247864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:46:00.153348Z digest=sha256:c5cd13e6bc469df4f53f63a560ff1d9b1c8e203f17dbf878cbe6d81dbc2767f7

Observation eeb99690-bec2-4b27-98f8-cf993201bc0d · outbound

This paper cites Efficient multimodal fusion via interactive prompting,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Efficient multimodal fusion via interactive prompting,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.317889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:46:00.155562Z digest=sha256:6388cc8319ec61a24420e16fdf58f3b8a22997a606c7f67632cc5e8dee8708d7

Observation ac88227b-92e0-408c-b950-fae682fc0372 · outbound

This paper cites Supervised Multimodal Bitransformers for Classifying Images and Text.

EPIC: Efficient Prompt Interaction for Text-Image Classification Supervised Multimodal Bitransformers for Classifying Images and Text

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:46:00.157567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:46:00.157567Z digest=sha256:acbf0e17680c7d285182c451f428ca354adba0c8915dd261522507e09c92f319

Observation 227179d9-8fa5-4808-8fd0-2c14be0f48ce · outbound

This paper cites Visual prompt tuning,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Visual prompt tuning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.311547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:46:00.159794Z digest=sha256:75cc522c60ffe2fc5292e89bba129f7cdb2c9489cffb203f83fa69823be15155

Observation 98149875-ad99-444a-ade8-d39581cb2e4f · outbound

This paper cites Learning to prompt for vision-language models,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Learning to prompt for vision-language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.305386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:46:00.162217Z digest=sha256:4f88a4fe47c5820ad2034bfe12c51dad8b455b476ddb0357c2867c6453cc484e

Observation 14a89053-f6d6-4a61-b20a-b8f1654001b3 · outbound

This paper cites Con- ditional prompt learning for vision-language models,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Con- ditional prompt learning for vision-language models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.298928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:46:00.164134Z digest=sha256:843b1660b64405b823ee883b2cd7d0d710e4a0b20640585ea63d7b6c39eed6b8

Observation ccb5fc53-3815-46af-a945-cf2dcb5f7c32 · outbound

This paper cites Maple: Multi-modal prompt learning,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Maple: Multi-modal prompt learning,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.292220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:46:00.166005Z digest=sha256:09105cec1c08a05e7f040d1def3718040fd7c4664c1885340607c04a197ac185

Observation 50a382c6-55d6-4255-94f5-63a344d52c98 · outbound

This paper cites HUSE: Hierarchical Universal Semantic Embeddings.

EPIC: Efficient Prompt Interaction for Text-Image Classification HUSE: Hierarchical Universal Semantic Embeddings

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:46:00.232501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:46:00.167906Z digest=sha256:4b226ce6ace9ea450e6535cc0de9ea29e77ef16817435acaf9649efee3267a90

Observation 431cd7f3-ef63-4d69-895f-ff67452f0759 · outbound

This paper cites MultiBench: Multiscale Benchmarks for Multimodal Representation Learning.

EPIC: Efficient Prompt Interaction for Text-Image Classification MultiBench: Multiscale Benchmarks for Multimodal Representation Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:46:00.169992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:46:00.169992Z digest=sha256:1ce9d2b5c5089386eb5d76cfb57a10ef8020a7b8d299dedd8aafc6531d43c13c

Observation 38b259c1-d15f-4dd8-94bc-3bc43a8cf9d4 · outbound

This paper cites Dynamic multimodal fusion,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Dynamic multimodal fusion,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.285857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:46:00.172055Z digest=sha256:96da8023595041a2b6f789011fab536fe298b689a153a010b5742665c6dd4d90

Observation dd0298d8-4358-4f2b-a285-0749ee0c0956 · outbound

This paper cites Unit: Multimodal multitask learning with a unified transformer,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Unit: Multimodal multitask learning with a unified transformer,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.280004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:46:00.173800Z digest=sha256:68d5f75b1fb3a3693fc827b53bea69f152579dc81e13e617cde3d9cb7bd4c71f

Observation 52dfd6b7-6988-475a-9321-9a35ca8260a2 · outbound

This paper cites Vilt: Vision-and-language transformer without convolution or region supervision,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Vilt: Vision-and-language transformer without convolution or region supervision,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.273527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:46:00.175634Z digest=sha256:527f3d3937f863b80c981a6645651c964f0ed38f0b5170e994c0eb489d826b34

Observation f0d32a1e-ec7f-44e2-9a7f-2c7b5cf7eef9 · outbound

This paper cites Recipe recognition with large multimodal food dataset,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Recipe recognition with large multimodal food dataset,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.267354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:46:00.177617Z digest=sha256:6b327690d0b2b3d6aa85a559c30282b13d055ab939432ecc490462f159e20d60

Observation 3b6e58a0-f577-4bed-b1c0-e0d3316eaeaf · outbound

This paper cites Gated Multimodal Units for Information Fusion.

EPIC: Efficient Prompt Interaction for Text-Image Classification Gated Multimodal Units for Information Fusion

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:46:00.179567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:46:00.179567Z digest=sha256:5c720b72e15449b90bc1a4bbc1faba033646a6465a1f4b76db7b01a0ca1fb9a1

Observation e6d8ccb1-0998-40c7-befa-b26073a20ddd · outbound

This paper cites Visual Entailment Task for Visually-Grounded Language Learning.

EPIC: Efficient Prompt Interaction for Text-Image Classification Visual Entailment Task for Visually-Grounded Language Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:46:00.181739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:46:00.181739Z digest=sha256:7ed7062f8edd06efccc8514a64b8e34b42e3c456a67cd9bb5d85e88eb0611ec6

Observation 4d0e251c-11f0-4d21-9a81-699fac385c14 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Learning transferable visual models from natural language supervision,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.260276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-08-06T18:46:00.183804Z digest=sha256:e495cad751589d5eed995e21107be7c03e204c010a0b81eae21be8b8bf63a8eb

Observation fe1d239f-81dd-4190-a66d-455ec0b9a538 · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

EPIC: Efficient Prompt Interaction for Text-Image Classification VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:46:00.186219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:46:00.186219Z digest=sha256:e53776a69bec9a6e3f27bfba67c5dcbcedc81c6592791f1e59b4341d5e311e5d

Pith citing papers

No inbound Pith citation observations are available.