Pith. sign in

Paper Citation Record · LEDGER

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding

As of 18 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2508.16974.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.16974 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T17:12:14.964857Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation be8cb9f3-ad98-430a-bf66-d46eb02392a8 · outbound

This paper cites A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding A survey on large language model (LLM) security and privacy: The good, the bad, and the ugly,

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-15T17:12:14.825139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:12:14.825139Z digest=sha256:2725d9b1d231d8490eb4b97285e6dd00b6e5043ce0f424fc1d9435e2663d546e

Observation a51ea89a-8c4f-441e-908b-1b5e83f29dc3 · outbound

This paper cites Visual in-context learning for large vision-language models,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Visual in-context learning for large vision-language models,

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T17:12:14.830658Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:12:14.830658Z digest=sha256:82d3cffab70b171bde113bb1f6a13d62bb311f952d6481d6152c3ed4b54d8fd5

Observation b619d075-625c-4cb0-8642-dc664a2714c8 · outbound

This paper cites Weak to strong generalization for large language models with multi-capabilities,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Weak to strong generalization for large language models with multi-capabilities,

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T17:12:14.835705Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:12:14.835705Z digest=sha256:e6b694704678ffce070fb39059bfa9650f3eee1be1b3e86dfa2f96c8f37b93e6

Observation 46db9d66-cade-4a98-8e46-00e1158efbf6 · outbound

This paper cites Multimodal event transformer for image-guided story ending generation,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Multimodal event transformer for image-guided story ending generation,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T17:12:14.840811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:12:14.840811Z digest=sha256:c58a878fdbce1274b91ab692fc6865eb45bb450901ca346f4b88444f33d5e593

Observation c1c2bb6f-087c-4301-beaf-cf420bbf57f9 · outbound

This paper cites A comprehensive survey of knowledge-based vision question answering systems: The lifecycle of knowledge in visual reasoning task,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding A comprehensive survey of knowledge-based vision question answering systems: The lifecycle of knowledge in visual reasoning task,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.431514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:12:14.846001Z digest=sha256:d2c2d8a6af84e432d7a074f91bb810544b5a243f684286e8e0bd75ddca5ef386

Observation 14f131fd-694d-44db-a0b1-e00ca27cfbfa · outbound

This paper cites Culture and hallucinations: overview and future directions,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Culture and hallucinations: overview and future directions,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.412837Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:12:14.851229Z digest=sha256:90fc5aaa56e47f7adf488c2f397877230180aa7d77bbe02996d87ed8aeb35895

Observation 44b49b13-3a2b-426c-ac1f-ad57371ada57 · outbound

This paper cites Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Improving Medical Large Vision-Language Models with Abnormal-Aware Feedback

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T17:12:14.856735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:12:14.856735Z digest=sha256:81848e2b45872d11f3232743e27859e760dfd355db46be9be17bfa2edb9bfaa1

Observation aa513c0e-d795-4ef4-91ad-5bdc1bbde082 · outbound

This paper cites Thread of Thought Unraveling Chaotic Contexts.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Thread of Thought Unraveling Chaotic Contexts

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T17:12:14.862091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:12:14.862091Z digest=sha256:2a716c368c17bea8d66eefa5e8ee44e4cb89552e38e60f0d3677c054989b9376

Observation 01715967-96e1-405d-a024-b5f9419b7606 · outbound

This paper cites Flamingo: a visual language model for few-shot learning,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Flamingo: a visual language model for few-shot learning,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.391633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:12:14.867271Z digest=sha256:b574c6d25f7621eac6bbb43db89f6d72da80311c5c1b742b2821dd5d43683c5b

Observation 01ed2a53-977a-4bf1-a1de-374ea9fe16c9 · outbound

This paper cites BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding BLIP-2: bootstrapping language-image pre-training with frozen image encoders and large language models,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.372661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:12:14.872031Z digest=sha256:475f9a96b7f3581b2fbc33cf6d3633e96cff6cfa4c5b8bc864e0c22d4ccf3b77

Observation 497e9d34-de45-45ac-b634-cf2d438f6bbc · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Minigpt-4: Enhancing vision-language understanding with advanced large language models,

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T17:12:14.877008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:12:14.877008Z digest=sha256:a5284c6ca93fcec38bbc67a3c3e54f2c1cf7e0120cb8b0a02ba8e597bccf920a

Observation fcb32ac4-8ab4-4e15-855c-ecf7a542b055 · outbound

This paper cites GQA: A new dataset for real- world visual reasoning and compositional question answering,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding GQA: A new dataset for real- world visual reasoning and compositional question answering,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.344129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:12:14.881816Z digest=sha256:f946eee475c04f8c895681944cb16d36d923c1870d707ca7a9e97817b57efe25

Observation 8a1de3c9-cb99-4917-b2a7-786b127b7691 · outbound

This paper cites A-OKVQA: A benchmark for visual question answering using world knowledge,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding A-OKVQA: A benchmark for visual question answering using world knowledge,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.327589Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:12:14.886566Z digest=sha256:ea5444bf7bac2d2144ae4003d0b37ff592037ba42957f9b27be430859b6e9159

Observation cfbb57bd-f5e4-4c7a-835a-4f0932060aed · outbound

This paper cites Revisiting referring expression comprehension evaluation in the era of large multimodal models,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Revisiting referring expression comprehension evaluation in the era of large multimodal models,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.309390Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:12:14.891472Z digest=sha256:c1893344921c7c25084262ae751db3220333669b4ee621d06c8b0cb2c9564914

Observation 6e0d7af7-b9de-41ac-abd6-06c3883f30b0 · outbound

This paper cites Debiasing vision-language models for vision tasks: a survey,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Debiasing vision-language models for vision tasks: a survey,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.290672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:12:14.896309Z digest=sha256:1ccf26befd1aec9074cfbb39e46981fb84aa01eddcb96ea3cbf880aade678d5a

Observation b49ffc55-3292-4b8b-a6be-24e1bb4de6a4 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Instructblip: Towards general-purpose vision-language models with instruction tuning,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.273132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:12:14.901298Z digest=sha256:abe095cabeb57eea50d58c229bcd76977f15db2a655452f5c33ef2f07dfc3f4d

Observation a32ebee4-87de-4849-8d37-88a7dc31ef0e · outbound

This paper cites Towards multimodal in-context learning for vision and language models,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Towards multimodal in-context learning for vision and language models,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.255867Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:12:14.906201Z digest=sha256:3cbdd404b8085e57e7850c518423146ac919404a68c955ab0eb43e0dbef44861

Observation a79a8265-01f7-4aee-9753-a14e1fb32ded · outbound

This paper cites VILA: on pre-training for visual language models,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding VILA: on pre-training for visual language models,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.236460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:12:14.911242Z digest=sha256:c3e9a795122c3ab1513077e6dfc5427be56069fe8d1a9164557a26045fef4180

Observation 6b522645-74d5-4dcf-8a35-e2b4dfdee998 · outbound

This paper cites MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding MAM: Modular Multi-Agent Framework for Multi-Modal Medical Diagnosis via Role-Specialized Collaboration

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T17:12:14.916090Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:12:14.916090Z digest=sha256:b0a63d1801d84b5e3b840a7e147119782a2f7ca41ac17d746a6b4f58ff10e6cf

Observation 3eb9e849-3925-4efc-9e9c-d4903767f67e · outbound

This paper cites Deciphering cross-modal alignment in large vision-language models with modality integration rate,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Deciphering cross-modal alignment in large vision-language models with modality integration rate,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.216343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:12:14.921122Z digest=sha256:467e7e1f69b8dd737d760a4192687e040323a4723a5ef1f5a4b22148f07c7b75

Observation c9b93399-892d-4236-8494-030d768cd419 · outbound

This paper cites Quantized prompt for efficient generalization of vision-language models,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Quantized prompt for efficient generalization of vision-language models,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.199357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:12:14.925861Z digest=sha256:4a541b048f4336ac43460cb9bbbeb5425d4a8709ff259b9b13a31f248427005e

Observation 9dea9465-5870-4740-82a8-d4f9f4ac9a99 · outbound

This paper cites A closer look at the few-shot adaptation of large vision-language models,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding A closer look at the few-shot adaptation of large vision-language models,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.182253Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:12:14.931026Z digest=sha256:b4bd9b334822a1d042c7ebabf83ecc1ca59c19d12f6431dbad2cbbb7771616db

Observation 602b964a-3c95-48ed-a345-2fe974be4eb6 · outbound

This paper cites Analyzing fine-grained alignment and enhancing vision understanding in multimodal language models,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Analyzing fine-grained alignment and enhancing vision understanding in multimodal language models,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.165507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:12:14.935957Z digest=sha256:18be65e985eacfb5e2f9138dfc0579ed030e15fb744299e46fecb445beec3457

Observation 197454cb-838f-42f4-b2c4-f781031b7724 · outbound

This paper cites Hivg: Hierarchical multimodal fine-grained modulation for visual grounding,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Hivg: Hierarchical multimodal fine-grained modulation for visual grounding,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.148166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:12:14.940962Z digest=sha256:1a0877c7ea2e343ccd15ec3ad80bcb9af9de45896b54742eaaf203c7d05ef5bb

Observation 11e6eae3-c70b-4387-9a59-b2b84f91eb5c · outbound

This paper cites Towards grounded visual spatial reasoning in multi-modal vision language models,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Towards grounded visual spatial reasoning in multi-modal vision language models,

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T17:12:14.945803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T17:12:14.945803Z digest=sha256:0caac5c2fda2fe2c41d760fbc2de9bbec695901410b1b5f0c7124d2f3d9d8abb

Observation 8e1c078d-fdab-44de-a51a-c01d2ce734c0 · outbound

This paper cites Analyzing and boosting the power of fine-grained visual recognition for multi-modal large language models,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Analyzing and boosting the power of fine-grained visual recognition for multi-modal large language models,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.118424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:12:14.950408Z digest=sha256:23d0ce407592ac322f1809741ccf34d04004e9a04e21125a9eae13375b21e848

Observation 23b79258-c320-4949-80b8-c5131073abe0 · outbound

This paper cites Insectmamba: State space model with adaptive composite features for insect recognition,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Insectmamba: State space model with adaptive composite features for insect recognition,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.100576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:12:14.955278Z digest=sha256:9af6c6ee8488b9a8b2e11e00cb2965c34dbe03db29f3ae214ff5ed9cd5c2365f

Observation 0d18babe-25b9-47a8-bde5-e1a4d08a29ad · outbound

This paper cites Towards visual grounding: A survey,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Towards visual grounding: A survey,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.082407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:12:14.960186Z digest=sha256:4b3a7a7ac7c0738eb030772fc5fbe8368e864c473e14f3090d626a0d0b955162

Observation 4b99eb64-a073-44d3-835b-c14bb9e2902f · outbound

This paper cites Hallucination of multimodal large language models: A survey,.

Hierarchical Contextual Grounding LVLM: Enhancing Fine-Grained Visual-Language Understanding with Robust Grounding Hallucination of multimodal large language models: A survey,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T17:12:15.064307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-15T17:12:14.964857Z digest=sha256:2b02f80a1269f8d3ad44299f33c05f49caa0baf0468e9aa0eca6c75d2f660adf

Pith citing papers

No inbound Pith citation observations are available.