Pith. sign in

Paper Citation Record · LEDGER

ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2405.15738.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.15738 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T04:35:52.235865Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:59:57.250489Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 98d09c32-aef5-431c-8245-a3c823a180fd · inbound

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? cites this paper.

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-16T07:59:32.729673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-16T07:59:32.638758Z digest=sha256:ff5d78ac0f8f165b71b90cf1b3396104107bfc39780bcef6cfefdf22d7713213

Observation 432abb0b-d338-4279-8de6-313248a04d2b · inbound

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion cites this paper.

Enhancing Perception Capabilities of Multimodal LLMs with Training-Free Fusion ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T04:35:52.235865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:35:52.235865Z digest=sha256:9c0fa1507a2933adde3e658aab4aeeea8e48c52339dd4899399fba39e457d28a

Observation 0a835e97-59ee-4fa8-bca9-a090ddfd0bd2 · inbound

CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception cites this paper.

CVSearch: Empowering Multimodal LLMs with Cognitive Visual Search for High-Resolution Image Perception ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:50:21.306292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-25T04:46:52.071071Z digest=sha256:2dd60dfd1d6f6aed89266c6eed39d3435eedc4f7a764bd08b2cd0502af07467a

Observation df9e0a50-bc5e-4a15-a92b-adcae9858e33 · inbound

ActiveScope: Actively Seeking and Correcting Perception for MLLMs cites this paper.

ActiveScope: Actively Seeking and Correcting Perception for MLLMs ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:59:57.252044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-26T01:06:40.481169Z digest=sha256:d430f1efa2a3de89f41fd9a59471d878d0edcc1a4a3a6e88efd1bc902d4c7890

Observation 33c0eaca-233f-4245-a0ec-acf2299cc9d2 · inbound

Towards High-Resolution Visual Perception via Hierarchical Entity Exploration cites this paper.

Towards High-Resolution Visual Perception via Hierarchical Entity Exploration ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-07-02T14:07:02.344333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-02T14:00:31.479090Z digest=sha256:6375747425cc91e39899e4c0713395e92ed98bbcb2754f5dd2a349d3612f8f11

Observation 6da9a312-d3e2-4246-aa58-688f1c359446 · inbound

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception cites this paper.

BVS: Bayesian Visual Search with Multimodal Large Language Model for Fine-grained Perception ConvLLaVA: Hierarchical Backbones as Visual Encoder for Large Multimodal Models

Reference 145

Resolution
unresolved
no resolver link, observed 2026-07-12T04:17:40.198357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T04:17:40.198357Z digest=sha256:dd1eba9ebfc4344ab45e21d8f843fa0b7280075a4ae03a228204317a953a24a0