Pith. sign in

Paper Citation Record · LEDGER

Joint Visual and Text Prompting for Improved Object-Centric Perception with Multimodal Large Language Models

As of 14 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2404.04514.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.04514 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:30:00.336670Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.132003Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ec48b512-cfcb-454e-a299-7818f909cf92 · inbound

EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios cites this paper.

EgoPlan-Bench2: A Benchmark for Multimodal Large Language Model Planning in Real-World Scenarios Joint Visual and Text Prompting for Improved Object-Centric Perception with Multimodal Large Language Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-11T21:30:00.336670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:30:00.336670Z digest=sha256:fe81b697bf8671ec635d5ff41c1c44467d36ee6d026b5050cfa6063cc0d9f102

Observation d4f1458c-9434-43a7-96cd-a74cffb507a2 · inbound

HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models cites this paper.

HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models Joint Visual and Text Prompting for Improved Object-Centric Perception with Multimodal Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:00:48.850779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:00:48.850779Z digest=sha256:3c5ad86a90bac3d20dcb77d6a6f534f6d1c5ac19f26de7782c85833d5946ab6c

Observation a06393d7-30f2-4f3a-af02-9e35869be43a · inbound

Fast or Slow? Integrating Fast Intuition and Deliberate Thinking for Enhancing Visual Question Answering cites this paper.

Fast or Slow? Integrating Fast Intuition and Deliberate Thinking for Enhancing Visual Question Answering Joint Visual and Text Prompting for Improved Object-Centric Perception with Multimodal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:10.543405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:10.543405Z digest=sha256:10772cd6f31dd5dc304439fba449e3518ea9bce7dd312b022e270cdd92611559

Observation b698e741-5e69-45fe-b013-9b060e96ccaa · inbound

AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs cites this paper.

AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs Joint Visual and Text Prompting for Improved Object-Centric Perception with Multimodal Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:44.479493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:44.479493Z digest=sha256:b021f1a7a80d34215ddd8174d89d14ce7f52992350662a5b5c0e255e7a6704b9

Observation a1ad24e4-5e56-4f70-b626-6dd19f423b2c · inbound

V2T-CoT: From Vision to Text Chain-of-Thought for Medical Reasoning and Diagnosis cites this paper.

V2T-CoT: From Vision to Text Chain-of-Thought for Medical Reasoning and Diagnosis Joint Visual and Text Prompting for Improved Object-Centric Perception with Multimodal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:46.047700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:46.047700Z digest=sha256:bdc50c0c22de8a20abdfb67d4ebb0c3b74f9bc2c3d292f15a182fe71d5279e6e

Observation d373e16c-9b10-4279-bce9-c2ac6828bcd0 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Joint Visual and Text Prompting for Improved Object-Centric Perception with Multimodal Large Language Models

Reference 125

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.133807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:06833288d65ded2dfe72df7487a022d276d35803988b4829efb329f41df7d3a5