Pith. sign in

Paper Citation Record · LEDGER

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2504.18152.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.18152 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:36:14.442972Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T09:54:35.322765Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation da4ef8b4-87b9-44b4-bb5c-c0544257c32a · inbound

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context cites this paper.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:14.442972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:14.442972Z digest=sha256:5e79d44793a0699a16586c6366f3c1c7173d3cbb444cacd96810df9b318eb4f4

Observation 25fe8ebe-aa16-489f-aea8-7827e1cd60b3 · inbound

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization cites this paper.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:20.338147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:20.338147Z digest=sha256:b7dcfa316d7f665cbb0fa4b27415d88d134424f219fded25e12a0df2f628b94c

Observation f50d51cd-1905-4ed7-b73f-622651c3aa42 · inbound

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration cites this paper.

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:49.056011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T20:20:08.590407Z digest=sha256:418f900c8f11112969f3cef033d998e262b03c460d392d480b86eb5fa57632d8

Observation 690280ac-a2d5-4e5b-b0dc-d60ab562d5e4 · inbound

Mining Multi-Modality Spatio-Temporal Cues for Video Important Person Identification cites this paper.

Mining Multi-Modality Spatio-Temporal Cues for Video Important Person Identification ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:33:28.320458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T13:26:21.016303Z digest=sha256:1a359948ff2518c8386ca77ff15527169ab11b2999a027c6ff0ab4ca0db9d72d

Observation 251ad567-e0b9-480a-9e46-e318d7da0ca8 · inbound

HumanMoveVQA: Can Video MLLMs reason about human movement in videos? cites this paper.

HumanMoveVQA: Can Video MLLMs reason about human movement in videos? ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:13:53.056250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-29T04:53:11.830488Z digest=sha256:9a2e92e9cd8a134c6cfb19225a30de146d05e90d92791219f4fd504dc67b51f6

Observation b58fd2bb-2e51-4fd0-a17a-1a51846e9090 · inbound

HumanMoveVQA: Can Video MLLMs reason about human movement in videos? cites this paper.

HumanMoveVQA: Can Video MLLMs reason about human movement in videos? ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:54:35.324230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-06-30T09:46:56.063333Z digest=sha256:be57f86024d8183b5ee3ab1c190bf4e77384cf81189e9cd0bb4477503573255f