Pith. sign in

Paper Citation Record · LEDGER

EPIC: Efficient Prompt Interaction for Text-Image Classification

As of 10 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2507.07415.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.07415 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:46:00.186219Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact2
  • verified fuzzy16
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e14bf6c4-7ead-4b3e-97fe-5da5a5c3fc0c · outbound

This paper cites Cma-clip: Cross- modality attention clip for text-image classification,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Cma-clip: Cross- modality attention clip for text-image classification,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.354240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:46:00.137419Z digest=sha256:37e19b58d2841bca1e762cae0e6d81ba838ebda2626b19d07bf77d71e3891a64

Observation 33fe7a97-2abf-4ce2-99e4-f5d01c59564e · outbound

This paper cites Future-aware diverse trends framework for recommendation,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Future-aware diverse trends framework for recommendation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.348130Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:46:00.140117Z digest=sha256:cc33b3c31e4b1658ae3b8b7f7d0d4c0adfb852b10827ca9b1180faaa983db41d

Observation 089606f7-3044-4917-b0a3-f1146e4545a7 · outbound

This paper cites Dense fusion network with multimodal residual for sentiment classification,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Dense fusion network with multimodal residual for sentiment classification,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.341633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:46:00.142283Z digest=sha256:faa452631f9cbf29ecc9f4231cb417f66ca20ed211a9fc927ad59fe5039f36fc

Observation 34939d49-d3ad-44b8-923c-2fc5cf032d3a · outbound

This paper cites Tensor Fusion Network for Multimodal Sentiment Analysis.

EPIC: Efficient Prompt Interaction for Text-Image Classification Tensor Fusion Network for Multimodal Sentiment Analysis

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T18:46:00.144295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:46:00.144295Z digest=sha256:4bea22d92b5c75521a3b19dbc1bd3c516066eabf253d442417c1f31dc2be26da

Observation 56960134-b3a1-4bc5-b46a-bfe3b619497b · outbound

This paper cites Memory fusion network for multi-view sequential learning,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Memory fusion network for multi-view sequential learning,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.335901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:46:00.146948Z digest=sha256:50e3f30522c54085eb65050282d30c6da3576b7ae19f3363a2110bbc68ef0879

Observation 2e0162a0-262e-4d25-897c-33e4adf5e804 · outbound

This paper cites Misa: Modality-invariant and-specific representations for multimodal sentiment analysis,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Misa: Modality-invariant and-specific representations for multimodal sentiment analysis,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.330209Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:46:00.149114Z digest=sha256:1588cc5a92aadf33cb0620d466985427ae88e7de01c339e99b29a0cc2363de58

Observation c7805295-3929-48e2-84b1-a43373a9f34d · outbound

This paper cites Centralnet: a multilayer approach for multimodal fusion,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Centralnet: a multilayer approach for multimodal fusion,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.324314Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:46:00.151389Z digest=sha256:a857ddbac9e505bd2655ef4542d567ee1541069bfed6002d583ad8d7113e380e

Observation 2830e517-deba-4ae4-b7cf-f4017f2b3250 · outbound

This paper cites Modular and Parameter-Efficient Multimodal Fusion with Prompting.

EPIC: Efficient Prompt Interaction for Text-Image Classification Modular and Parameter-Efficient Multimodal Fusion with Prompting

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:46:00.247864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:46:00.153348Z digest=sha256:d03ed9d61c2fe332c79599cb9fa7a772c8cd315cb0657d4fd18b7ec8bcdbca5d

Observation eeb99690-bec2-4b27-98f8-cf993201bc0d · outbound

This paper cites Efficient multimodal fusion via interactive prompting,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Efficient multimodal fusion via interactive prompting,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.317889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:46:00.155562Z digest=sha256:2e5bdc7a26a12d7f7aaa14232381dcdab54b259d1fe5a707fab61a4bde2285c1

Observation ac88227b-92e0-408c-b950-fae682fc0372 · outbound

This paper cites Supervised Multimodal Bitransformers for Classifying Images and Text.

EPIC: Efficient Prompt Interaction for Text-Image Classification Supervised Multimodal Bitransformers for Classifying Images and Text

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T18:46:00.157567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:46:00.157567Z digest=sha256:a97b1b89499f46143c130ec2561ed446a3d24f78da2eb4cd272ff4c7ef2f8e0b

Observation 227179d9-8fa5-4808-8fd0-2c14be0f48ce · outbound

This paper cites Visual prompt tuning,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Visual prompt tuning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.311547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:46:00.159794Z digest=sha256:8794975dc34844d8ebcca84a7e00fcded5da2641e81b224b6e45eb2892ccdc4f

Observation 98149875-ad99-444a-ade8-d39581cb2e4f · outbound

This paper cites Learning to prompt for vision-language models,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Learning to prompt for vision-language models,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.305386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:46:00.162217Z digest=sha256:e31191010aace3943f6009124463e974005c4df06a5783a042a9352f5c20b820

Observation 14a89053-f6d6-4a61-b20a-b8f1654001b3 · outbound

This paper cites Con- ditional prompt learning for vision-language models,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Con- ditional prompt learning for vision-language models,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.298928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:46:00.164134Z digest=sha256:3e803497130057bc613fb4e079174a203029bdb87418c846c182bed03a32b27f

Observation ccb5fc53-3815-46af-a945-cf2dcb5f7c32 · outbound

This paper cites Maple: Multi-modal prompt learning,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Maple: Multi-modal prompt learning,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.292220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:46:00.166005Z digest=sha256:c8179baabbb9befda8dad9ff38d5fce0aacf00676b4e9e0d0d33b763659b8af2

Observation 50a382c6-55d6-4255-94f5-63a344d52c98 · outbound

This paper cites HUSE: Hierarchical Universal Semantic Embeddings.

EPIC: Efficient Prompt Interaction for Text-Image Classification HUSE: Hierarchical Universal Semantic Embeddings

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-08-06T18:46:00.232501Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:46:00.167906Z digest=sha256:b9cfa69f653ce0d1b2aaa69cefd3a1fb4043b8ff4f7340b0bffbc378dd664607

Observation 431cd7f3-ef63-4d69-895f-ff67452f0759 · outbound

This paper cites MultiBench: Multiscale Benchmarks for Multimodal Representation Learning.

EPIC: Efficient Prompt Interaction for Text-Image Classification MultiBench: Multiscale Benchmarks for Multimodal Representation Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:46:00.169992Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:46:00.169992Z digest=sha256:eab3d112f3541dc78adad3b4c6c7ee4f3ba006f2554dc5155c4c6ba4bcd3413d

Observation 38b259c1-d15f-4dd8-94bc-3bc43a8cf9d4 · outbound

This paper cites Dynamic multimodal fusion,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Dynamic multimodal fusion,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.285857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:46:00.172055Z digest=sha256:2e70b78f62934618ea3f9e9cb2c65fa7b5557a42696211ea67f35e3d1ce7dfc6

Observation dd0298d8-4358-4f2b-a285-0749ee0c0956 · outbound

This paper cites Unit: Multimodal multitask learning with a unified transformer,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Unit: Multimodal multitask learning with a unified transformer,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.280004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:46:00.173800Z digest=sha256:8b82fe3a53c49e203842dead7249e3fc9c6f9c77831f34dc3b12cfdb431abf54

Observation 52dfd6b7-6988-475a-9321-9a35ca8260a2 · outbound

This paper cites Vilt: Vision-and-language transformer without convolution or region supervision,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Vilt: Vision-and-language transformer without convolution or region supervision,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.273527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:46:00.175634Z digest=sha256:472194d6942f17e63e2342c9726ca0816a57d1154ade6cbd91f0f25fcd9d1fd5

Observation f0d32a1e-ec7f-44e2-9a7f-2c7b5cf7eef9 · outbound

This paper cites Recipe recognition with large multimodal food dataset,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Recipe recognition with large multimodal food dataset,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.267354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:46:00.177617Z digest=sha256:858ecc60bbe72bc3eaad88a327b0fdfe5cc3dd9c3065b4c6e549665de1a59dec

Observation 3b6e58a0-f577-4bed-b1c0-e0d3316eaeaf · outbound

This paper cites Gated Multimodal Units for Information Fusion.

EPIC: Efficient Prompt Interaction for Text-Image Classification Gated Multimodal Units for Information Fusion

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T18:46:00.179567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:46:00.179567Z digest=sha256:5c720b72e15449b90bc1a4bbc1faba033646a6465a1f4b76db7b01a0ca1fb9a1

Observation e6d8ccb1-0998-40c7-befa-b26073a20ddd · outbound

This paper cites Visual Entailment Task for Visually-Grounded Language Learning.

EPIC: Efficient Prompt Interaction for Text-Image Classification Visual Entailment Task for Visually-Grounded Language Learning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:46:00.181739Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:46:00.181739Z digest=sha256:7ed7062f8edd06efccc8514a64b8e34b42e3c456a67cd9bb5d85e88eb0611ec6

Observation 4d0e251c-11f0-4d21-9a81-699fac385c14 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

EPIC: Efficient Prompt Interaction for Text-Image Classification Learning transferable visual models from natural language supervision,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T18:46:00.260276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T18:46:00.183804Z digest=sha256:1c0cbfc635521f6963d23e46853a5a8a6715201a0ad0f3d066a0ddb81036d393

Observation fe1d239f-81dd-4190-a66d-455ec0b9a538 · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

EPIC: Efficient Prompt Interaction for Text-Image Classification VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T18:46:00.186219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:46:00.186219Z digest=sha256:e53776a69bec9a6e3f27bfba67c5dcbcedc81c6592791f1e59b4341d5e311e5d

Pith citing papers

No inbound Pith citation observations are available.