Pith. sign in

Paper Citation Record · LEDGER

ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2504.18152.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.18152 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:36:14.442972Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-30T09:54:35.322765Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation da4ef8b4-87b9-44b4-bb5c-c0544257c32a · inbound

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context cites this paper.

HumanOmniV2: From Understanding to Omni-Modal Reasoning with Context ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T22:36:14.442972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:36:14.442972Z digest=sha256:a796ef7f84a57716e467c83f23d45ca6c2eb8e7e167fc2963d3b9cd0072a62d4

Observation 25fe8ebe-aa16-489f-aea8-7827e1cd60b3 · inbound

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization cites this paper.

AVC-DPO: Aligned Video Captioning via Direct Preference Optimization ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T20:53:20.338147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T20:53:20.338147Z digest=sha256:db4fc0cc4c56e0f5089217d19b27a3405643d74b312b2a79d0e6f5e371eace91

Observation f50d51cd-1905-4ed7-b73f-622651c3aa42 · inbound

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration cites this paper.

SVAgent: Storyline-Guided Long Video Understanding via Cross-Modal Multi-Agent Collaboration ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:00:49.056011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T20:20:08.590407Z digest=sha256:dbfea17f936cd56f53db951aec760082b7c6fa9790a9012e7f92f5c051cd1cd9

Observation 690280ac-a2d5-4e5b-b0dc-d60ab562d5e4 · inbound

Mining Multi-Modality Spatio-Temporal Cues for Video Important Person Identification cites this paper.

Mining Multi-Modality Spatio-Temporal Cues for Video Important Person Identification ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-06-29T13:33:28.320458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T13:26:21.016303Z digest=sha256:14325ddce8ebdd627ee754a73babc7768c59cf667f186a955ea04c5610615952

Observation 251ad567-e0b9-480a-9e46-e318d7da0ca8 · inbound

HumanMoveVQA: Can Video MLLMs reason about human movement in videos? cites this paper.

HumanMoveVQA: Can Video MLLMs reason about human movement in videos? ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-29T19:13:53.056250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-29T04:53:11.830488Z digest=sha256:e5f1c5ab01991c59a5c8216ad2970ba9c3af2b33540d031cc5c2768718d235b8

Observation b58fd2bb-2e51-4fd0-a17a-1a51846e9090 · inbound

HumanMoveVQA: Can Video MLLMs reason about human movement in videos? cites this paper.

HumanMoveVQA: Can Video MLLMs reason about human movement in videos? ActionArt: Advancing Multimodal Large Models for Fine-Grained Human-Centric Video Understanding

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-06-30T09:54:35.324230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T09:46:56.063333Z digest=sha256:68f8025c172482086f1d71dda5b0c7c422c258a34c3f9185eb3b3f76a86439bc