Pith. sign in

Paper Citation Record · LEDGER

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding

As of 17 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2508.16974.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.16974 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:12:14.964857Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation be8cb9f3-ad98-430a-bf66-d46eb02392a8 · outbound

This paper cites A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T17:12:14.825139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:12:14.825139Z digest=sha256:2725d9b1d231d8490eb4b97285e6dd00b6e5043ce0f424fc1d9435e2663d546e

Observation a51ea89a-8c4f-441e-908b-1b5e83f29dc3 · outbound

This paper cites Visual in-context learning for large vision-language models,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Visual in-context learning for large vision-language models,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T17:12:14.830658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:12:14.830658Z digest=sha256:82d3cffab70b171bde113bb1f6a13d62bb311f952d6481d6152c3ed4b54d8fd5

Observation b619d075-625c-4cb0-8642-dc664a2714c8 · outbound

This paper cites Weak to strong generalization for large language models with multi-capabilities,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Weak to strong generalization for large language models with multi-capabilities,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:12:14.835705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:12:14.835705Z digest=sha256:e6b694704678ffce070fb39059bfa9650f3eee1be1b3e86dfa2f96c8f37b93e6

Observation 46db9d66-cade-4a98-8e46-00e1158efbf6 · outbound

This paper cites Multimodal event transformer for image-guided story ending generation,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Multimodal event transformer for image-guided story ending generation,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T17:12:14.840811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:12:14.840811Z digest=sha256:c58a878fdbce1274b91ab692fc6865eb45bb450901ca346f4b88444f33d5e593

Observation c1c2bb6f-087c-4301-beaf-cf420bbf57f9 · outbound

This paper cites A comprehensive survey of knowledge-based vision question answering systems: The lifecycle of knowledge in visual reasoning task,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding A comprehensive survey of knowledge-based vision question answering systems: The lifecycle of knowledge in visual reasoning task,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.431514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:12:14.846001Z digest=sha256:95759fe76663f09f9e2a4e96c80b38b59a2f4bd680fc02e1c602b03817a437b6

Observation 14f131fd-694d-44db-a0b1-e00ca27cfbfa · outbound

This paper cites Culture and hallucinations: overview and future directions,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Culture and hallucinations: overview and future directions,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.412837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:12:14.851229Z digest=sha256:4e7d4a8ab201f928ad52cd6e834b42d9dae8b9c251b18617d6d57c575b743485

Observation 44b49b13-3a2b-426c-ac1f-ad57371ada57 · outbound

This paper cites Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:12:14.856735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:12:14.856735Z digest=sha256:81848e2b45872d11f3232743e27859e760dfd355db46be9be17bfa2edb9bfaa1

Observation aa513c0e-d795-4ef4-91ad-5bdc1bbde082 · outbound

This paper cites Thread of Thought Unraveling Chaotic Contexts.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Thread of Thought Unraveling Chaotic Contexts

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:12:14.862091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:12:14.862091Z digest=sha256:2a716c368c17bea8d66eefa5e8ee44e4cb89552e38e60f0d3677c054989b9376

Observation 01715967-96e1-405d-a024-b5f9419b7606 · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Flamingo: a visual language model for few-shot learning,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.391633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:12:14.867271Z digest=sha256:887caa77a92902bd601db683532318957b0dd9f15931005fb491c49823c6096a

Observation 01ed2a53-977a-4bf1-a1de-374ea9fe16c9 · outbound

This paper cites BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.372661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:12:14.872031Z digest=sha256:150c8f0635c760bae0df6adee44d70c9f7a5f35053691a1651b2c5743af8ea08

Observation 497e9d34-de45-45ac-b634-cf2d438f6bbc · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Minigpt-4: Enhancing vision-language understanding with advanced large language models,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:12:14.877008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:12:14.877008Z digest=sha256:a5284c6ca93fcec38bbc67a3c3e54f2c1cf7e0120cb8b0a02ba8e597bccf920a

Observation fcb32ac4-8ab4-4e15-855c-ecf7a542b055 · outbound

This paper cites GQA: A new dataset for real- world visual reasoning and compositional question answering,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding GQA: A new dataset for real- world visual reasoning and compositional question answering,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.344129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:12:14.881816Z digest=sha256:8121578540548a43e8f25512bbb115dff8a32ee60d6427f0fcbdd18d20fc7c49

Observation 8a1de3c9-cb99-4917-b2a7-786b127b7691 · outbound

This paper cites A-OKVQA: A benchmark for visual question answering using world knowledge,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding A-OKVQA: A benchmark for visual question answering using world knowledge,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.327589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:12:14.886566Z digest=sha256:282b09e9fd79029f16c8f823713ee74837fa16c311be2a0b3a9b05d3d6ea1074

Observation cfbb57bd-f5e4-4c7a-835a-4f0932060aed · outbound

This paper cites Revisiting referring expression comprehension evaluation in the era of large multimodal models,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Revisiting referring expression comprehension evaluation in the era of large multimodal models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.309390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:12:14.891472Z digest=sha256:c50a1728f2e8c2f1ce2355ee727913d65738dea7b3bc2da295dbd56189c7fd4f

Observation 6e0d7af7-b9de-41ac-abd6-06c3883f30b0 · outbound

This paper cites Debiasing vision-language models for vision tasks: a survey,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Debiasing vision-language models for vision tasks: a survey,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.290672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:12:14.896309Z digest=sha256:1ff04ef89d2114259c9b89185c88dab7dab352671651532017369f0e64705c2c

Observation b49ffc55-3292-4b8b-a6be-24e1bb4de6a4 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Instructblip: Towards general-purpose vision-language models with instruction tuning,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.273132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:12:14.901298Z digest=sha256:a9941c87b9f316cc5e8443804cb1b2f2c811c7b9f6f9532b1387aeaf26318d1c

Observation a32ebee4-87de-4849-8d37-88a7dc31ef0e · outbound

This paper cites Towards multimodal in-context learning for vision and language models,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Towards multimodal in-context learning for vision and language models,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.255867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:12:14.906201Z digest=sha256:bb3ef4e3dda8965fc0a6d0b48827fd897e2af6b253fa0cb780b8f6af9969fb17

Observation a79a8265-01f7-4aee-9753-a14e1fb32ded · outbound

This paper cites VILA: on pre-training for visual language models,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding VILA: on pre-training for visual language models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.236460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:12:14.911242Z digest=sha256:b866e9ae5d646d67b86b05a7b1939c70a120c42aa5d1141267044c20b1dd4a25

Observation 6b522645-74d5-4dcf-8a35-e2b4dfdee998 · outbound

This paper cites MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:12:14.916090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:12:14.916090Z digest=sha256:b0a63d1801d84b5e3b840a7e147119782a2f7ca41ac17d746a6b4f58ff10e6cf

Observation 3eb9e849-3925-4efc-9e9c-d4903767f67e · outbound

This paper cites Deciphering cross-modal alignment in large vision-language models with modality integration rate,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Deciphering cross-modal alignment in large vision-language models with modality integration rate,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.216343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:12:14.921122Z digest=sha256:5ca9b3c680905eed16b3ab16c52768145cc7757df24a98328a6067a2baa52005

Observation c9b93399-892d-4236-8494-030d768cd419 · outbound

This paper cites Quantized prompt for efficient generalization of vision-language models,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Quantized prompt for efficient generalization of vision-language models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.199357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:12:14.925861Z digest=sha256:8f06eb5cff8e755de0e87d0ffffa5e329cae021365b71d79b2e525cdf100c5b7

Observation 9dea9465-5870-4740-82a8-d4f9f4ac9a99 · outbound

This paper cites A closer look at the few-shot adaptation of large vision-language models,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding A closer look at the few-shot adaptation of large vision-language models,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.182253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:12:14.931026Z digest=sha256:9f4b6ec3d75d483278af1eae3e398b7d8b29bddd895c0edc844c21dcf2b78e58

Observation 602b964a-3c95-48ed-a345-2fe974be4eb6 · outbound

This paper cites Analyzing fine-grained alignment and enhancing vision understanding in multimodal language models,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Analyzing fine-grained alignment and enhancing vision understanding in multimodal language models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.165507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:12:14.935957Z digest=sha256:0cff04741e879f68620309f146ef3da6eafc31deab6814b82dbe8c0581980d54

Observation 197454cb-838f-42f4-b2c4-f781031b7724 · outbound

This paper cites Hivg: Hierarchical multimodal fine-grained modulation for visual grounding,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Hivg: Hierarchical multimodal fine-grained modulation for visual grounding,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.148166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:12:14.940962Z digest=sha256:aa07160294567a999f1b89f69e78595e1afb4438461607cadd61dc47fbdcda84

Observation 11e6eae3-c70b-4387-9a59-b2b84f91eb5c · outbound

This paper cites Towards grounded visual spatial reasoning in multi-modal vision language models,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Towards grounded visual spatial reasoning in multi-modal vision language models,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T17:12:14.945803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:12:14.945803Z digest=sha256:0caac5c2fda2fe2c41d760fbc2de9bbec695901410b1b5f0c7124d2f3d9d8abb

Observation 8e1c078d-fdab-44de-a51a-c01d2ce734c0 · outbound

This paper cites Analyzing and boosting the power of fine-grained visual recognition for multi-modal large language models,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Analyzing and boosting the power of fine-grained visual recognition for multi-modal large language models,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.118424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:12:14.950408Z digest=sha256:5135b8cdb9fa5665400fcec86661bf71eddef1ff31edeef7d51e34a2afd2538e

Observation 23b79258-c320-4949-80b8-c5131073abe0 · outbound

This paper cites Insectmamba: State space model with adaptive composite features for insect recognition,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Insectmamba: State space model with adaptive composite features for insect recognition,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.100576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:12:14.955278Z digest=sha256:b68b01bb29028d288b898875785efc0f91eba9cbdb67278ebe6e90c74339d389

Observation 0d18babe-25b9-47a8-bde5-e1a4d08a29ad · outbound

This paper cites Towards visual grounding: A survey,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Towards visual grounding: A survey,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.082407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:12:14.960186Z digest=sha256:15a8fae7b3969dc205b4886a49fc36a43f59eb076e3b70b11b49db608c7de7dd

Observation 4b99eb64-a073-44d3-835b-c14bb9e2902f · outbound

This paper cites Hallucination of multimodal large language models: A survey,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Hallucination of multimodal large language models: A survey,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.064307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T17:12:14.964857Z digest=sha256:df03631e7bab1f09ebda2d8e9d23239702e64fe6cc3977c783458243b0315cad

Pith citing papers

No inbound Pith citation observations are available.