Pith. sign in

Paper Citation Record · LEDGER

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering

As of 7 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 0 inbound Pith citation observations for arXiv:2508.03039.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.03039 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:49:42.909084Z

measured 53 of 53 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

53 of 53 outbound references displayed

  • verified exact4
  • verified fuzzy2
  • unresolved44
  • parse uncertain0
  • malformed identifier2
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d0fc53e2-d681-4262-9f49-feb7351126b9 · outbound

This paper cites The IKEA ASM Dataset: Understanding People Assembling Furniture through Actions, Objects and Pose.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering The IKEA ASM Dataset: Understanding People Assembling Furniture through Actions, Objects and Pose

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:49:44.836759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:37.193595Z digest=sha256:8e8367b9fbe1760253bfaec30ba4ef198ebd4a7f17fe8c1c2e5f2b094b33c1e2

Observation 8916b5d6-5696-4ab6-85bb-0c4570017a2e · outbound

This paper cites VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:37.262227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:37.262227Z digest=sha256:3d7f15129af5200ef5eb2778659d74aefbfddacae4493fe724fb117bec05dcd0

Observation dbbd0beb-e5af-453c-862f-f1b19e196efe · outbound

This paper cites A Short Note on the Kinetics-700 Human Action Dataset.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering A Short Note on the Kinetics-700 Human Action Dataset

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:37.375757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:37.375757Z digest=sha256:29f5f2956d4fb51837b0560fbf60ad8d6364db1a6d0f2ec83c6c7a4da2cab74a

Observation 457862a8-afb3-48f7-85d2-efa803b70330 · outbound

This paper cites ShareGPT4Video: Improving Video Understanding and Generation with Better Captions.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering ShareGPT4Video: Improving Video Understanding and Generation with Better Captions

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:37.464967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:37.464967Z digest=sha256:381e8ffc65f43d8d08700f61c3459153327213ceb3756396e3edbec930b6560e

Observation 4ed2ed85-0e9d-4dc1-a83a-a32078302cf0 · outbound

This paper cites Enhancing Long Video Understanding via Hierarchical Event-Based Memory.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Enhancing Long Video Understanding via Hierarchical Event-Based Memory

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:37.562864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:37.562864Z digest=sha256:30c16a4794c25b9aeeb65bf459636c8011c14bef196d46a4d1bee42b99584e0d

Observation e691ea34-fc65-4152-ac4c-de2035cca43e · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:37.643512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:37.643512Z digest=sha256:6435a52a762e233a34f26993dd15f64244039cc136249371982055f25494a187

Observation a74b6568-0c14-4b40-bcfd-65d75c08ab60 · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-06T04:49:45.982958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:37.738710Z digest=sha256:2febd8874c13d5ed1e966971cfd858c2d570133e4eef3fa91f7a5f5b5854dd83

Observation 59dd60d2-be20-446d-ac9c-a41ce768a327 · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 8

Resolution
metadata mismatch
raw_fallback, observed 2026-08-06T04:49:44.688355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:37.840755Z digest=sha256:27e6012f828554645233d59d0bad614010f6d5fb9329030190f314c9853e2619

Observation d8855dac-cc08-4c80-8e64-ec425fe59203 · outbound

This paper cites Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Video-CCAM: Enhancing Video-Language Understanding with Causal Cross-Attention Masks for Short and Long Videos

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:37.951565Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:37.951565Z digest=sha256:5de3e169218de39db78dd0c66f1b1636c61712717986efbded8439785bd66e8b

Observation 39f00fc1-9517-4c66-b420-af0ad379cf9d · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:38.032079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:38.032079Z digest=sha256:b217693c5e7e418a1295aff94af3d533d229f8b37fb08c42895ccde837f2f9e6

Observation 4c12f38d-35bd-4389-bdcc-01ab01d93e16 · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:38.149590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:38.149590Z digest=sha256:7cc344cce51739bdbdcc426575c5a33468ff97643c2c508b800ad2a3b14048b2

Observation c69bb0d7-baac-4ee6-80ee-9b06a3803f67 · outbound

This paper cites Video ReCap: Recursive Captioning of Hour-Long Videos.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Video ReCap: Recursive Captioning of Hour-Long Videos

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:38.388595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:38.388595Z digest=sha256:d4c4228152e5c7fa2ee9a40471a3300bfe650e1c82b54e44b15f31808ecf82c9

Observation ada452eb-d963-4838-8da9-cf70e5235c73 · outbound

This paper cites BIMBA: Selective-Scan Compression for Long-Range Video Question Answering.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering BIMBA: Selective-Scan Compression for Long-Range Video Question Answering

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:38.487988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:38.487988Z digest=sha256:50448b475a67ed74aa84b44f8951b8d4c72ab6e7397b39c6e0ad9fa6ae6d002f

Observation f4bcf83e-1057-4518-8e37-70fa5b0f5f89 · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-06T04:49:45.734096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:38.592793Z digest=sha256:74aa8c04f294ab327348c2a82f4773a538feec587be3a627f9fb73b537366dd0

Observation 3abd9c83-6526-4618-8dd2-b7e698badfed · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering LLaVA-OneVision: Easy Visual Task Transfer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:38.715526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:38.715526Z digest=sha256:a373e77c9656cc313712739e5405905b4cf1a5ada7e1c6dd9a05af493bf0fb07

Observation bb82cd92-7085-45ee-9194-c19c7c338fa8 · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:38.819208Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:38.819208Z digest=sha256:519d6c3243d7cd477dda25dc01d66fb44652e35dd34a8fd55306f60fd1465db8

Observation a6aafc0f-31fb-4274-a25b-06f213b47ac4 · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:38.925919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:38.925919Z digest=sha256:b2f3087c7e461775162585ede09b780b4fe5f7d3d94fda95c18b693ad1d2171f

Observation 215cecf1-67bb-41ca-83fc-6a30a3038aa5 · outbound

This paper cites VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering VideoChat-Flash: Hierarchical Compression for Long-Context Video Modeling

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:39.008817Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:39.008817Z digest=sha256:2c590706f66172d8a22951411485a5518f50d52d84d4d613ba3de0fafedba03c

Observation 880ab0e9-887c-47b5-bbbd-b542464beaca · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-06T04:49:45.656764Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:39.090836Z digest=sha256:d24d9c496c443b6709caa93dd137c11074a9dcb3d2617d56f544cc56bcde0c20

Observation 4be1116f-7d98-4a42-97ed-97c8dac97951 · outbound

This paper cites Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Commonsense Video Question Answering through Video-Grounded Entailment Tree Reasoning

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:49:44.268102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:39.256408Z digest=sha256:c2e6cc4da9b1638672c502f4644734ed74ad6fbf8784cc656dfaf8d385202dca

Observation 1fe38cbe-5629-477a-b57f-8a8f54bea77d · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:39.355423Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:39.355423Z digest=sha256:373ac6ae005db9869986bfd5e2ed1a023c4d60cb46530a0dd89e04f14b839417

Observation 43064c23-18b9-4948-b091-01583c16cbb1 · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 22

Resolution
malformed identifier
doi_truncated, observed 2026-08-06T04:49:43.187724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:39.449214Z digest=sha256:9816ebf1258cc30d738d1ad061075a04c923368d0174761b20919dd43bff9e72

Observation d55e937b-66a2-44c1-97c4-9820b5cc216c · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-06T04:49:45.517323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:39.548025Z digest=sha256:a1fe7b622e475bfe90e65a2efd595f5db06accfd9abe749aca4a4ee708487107

Observation f6d41882-8c87-4c39-8e1f-71d88ac17d0f · outbound

This paper cites FineAction: A Fine-Grained Video Dataset for Temporal Action Localization.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering FineAction: A Fine-Grained Video Dataset for Temporal Action Localization

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:49:43.956206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:39.607583Z digest=sha256:64e057aa38ba0fee3576094e638e96d4a3bf2bea1c20a836eb80edb4825cacf6

Observation c9ffc015-b9e8-4de1-b45f-f0574edc4c8a · outbound

This paper cites DrVideo: Document Retrieval Based Long Video Understanding.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering DrVideo: Document Retrieval Based Long Video Understanding

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:39.709238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:39.709238Z digest=sha256:c175a33138144fb46a4577a52608370f435fb2e5c0ea7cb1badc3a236a8b446b

Observation 4f84d07b-d150-4a22-a3e6-d5af65ef3d1d · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:39.834064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:39.834064Z digest=sha256:ca40804c71b93a656d004d313bc04c4503008d2cf1b0fcaa7465da567d187789

Observation 085a55b7-2b21-43a7-8a36-80c70466fbb1 · outbound

This paper cites Foundation Models for Video Understanding: A Survey.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Foundation Models for Video Understanding: A Survey

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:40.068720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:40.068720Z digest=sha256:b5b36412ade51da251a83f3cea67c919596f1b0bc126032c53d6b53a06956901

Observation a6dd47bf-e0be-4ea9-9140-bfcfb4aac0f3 · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:40.186187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:40.186187Z digest=sha256:7f089bc81c8582cd2e8f5f597697a1f3efee749d609ccd1e16b8b9928e6f60fe

Observation 58c488f5-236a-4915-8b1e-458f1a1cce65 · outbound

This paper cites InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024).

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering InProceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL 2024)

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:39.952238Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:39.952238Z digest=sha256:0d97de20bc5fefd28d37ceedce31ba265480e79af9f8752cdc9956e895890a80

Observation e049e15f-6f1d-43c2-b1ff-b5fc0c38de89 · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:40.413518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:40.413518Z digest=sha256:7960d0740dbf833b696d659bf5c96f0b707912fce64c32a69bedc37d9c885e9a

Observation 48bf7a61-c99d-485a-a7e5-114d49c266a4 · outbound

This paper cites Qasim, R.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Qasim, R

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:49:45.411683Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:40.562692Z digest=sha256:e23864c521fa1d12ab87b7a0f3338f4a78946d42b2ac507396022607ba7975ea

Observation bfe8e4e3-9465-4421-9f4b-1e4d1c9e580b · outbound

This paper cites Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Video-Bench: A Comprehensive Benchmark and Toolkit for Evaluating Video-based Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:40.297857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:40.297857Z digest=sha256:0e6abea9bbbcbd7e21c728f05a1249f5097a52885b147404d8ecc16d1186a595

Observation 0404bfd9-8da6-4e13-a968-9f43d86969a0 · outbound

This paper cites TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering TV-TREES: Multimodal Entailment Trees for Neuro-Symbolic Video Reasoning

Reference 33

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:49:43.602596Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:40.870463Z digest=sha256:1f6f16a49ed367e64a6b894fa63fb36868b8eca90f873e66e208c76e5bb86ec6

Observation 384eec76-6808-4fd5-8fea-1a8de3ccfbc5 · outbound

This paper cites Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Hollywood in Homes: Crowdsourcing Data Collection for Activity Understanding

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:41.021647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:41.021647Z digest=sha256:4de393987fed1517e00e39a14fd0afdb8beca8f2559122b6318c41cf4f162592

Observation 4d5c72c7-0faa-494f-b3f4-6453f13a14c2 · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:40.708535Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:40.708535Z digest=sha256:5597cd2ff955f2809089483b350e94d0fbf496e725f4028b849fce3d3c31dbfa

Observation eb9e33e5-fe94-4d8d-824e-90f52617c15d · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-06T04:49:45.322353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:41.240798Z digest=sha256:673f954a4242fc66df5271318acff7f17d642ab51e228c2cb61127646dff5f32

Observation cdb02b9b-67d9-4830-be24-be0b0b3b54f8 · outbound

This paper cites InternVideo: General Video Foundation Models via Generative and Discriminative Learning.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering InternVideo: General Video Foundation Models via Generative and Discriminative Learning

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:41.380659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:41.380659Z digest=sha256:94b498422c64c0ff56736b48491d8172a65e7082a6e088a07cb9395a6a95ab59

Observation 0dd44e4c-e53b-44f3-be17-484d8fd9ff9b · outbound

This paper cites ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering ChatVideo: A Tracklet-centric Multimodal and Versatile Video Understanding System

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:41.124471Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:41.124471Z digest=sha256:fbf360f1e4d7234e4545273e5efc3bd1b1fc78e4c16fcf4226c76f910a248090

Observation 72be860a-0026-4722-bb6d-14242836a9c7 · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 39

Resolution
unresolved
raw_fallback, observed 2026-08-06T04:49:45.166685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:41.636346Z digest=sha256:c985eb5395ab51ccfe1f035d0af96c8570e810b6fdfd6e3cbd9e117eed480e74

Observation 28fd947c-bfe8-4d0b-8954-e19b2f9e491f · outbound

This paper cites LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering LongVideoBench: A Benchmark for Long-context Interleaved Video-Language Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:41.776592Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:41.776592Z digest=sha256:f732a2c0718ddbb9d5e9ea8267e36185b35d388029b548b11aad49a5564dd247

Observation 2ac3d27f-7c44-4e0a-be78-6d8ba40ee92e · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:41.523576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:41.523576Z digest=sha256:53749a30057fedefa875ef50a5bc1b646530e1317eccaa5172c179d82726bccc

Observation 20751638-ba9e-4f60-84aa-533399259549 · outbound

This paper cites LLaVA-Critic: Learning to Evaluate Multimodal Models.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering LLaVA-Critic: Learning to Evaluate Multimodal Models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:42.003797Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:42.003797Z digest=sha256:c82a2c274fcd53dcf0d7542ee39ec00fc0df162dac181aeb72c877e61b27355a

Observation 7633f68b-5142-46fb-be42-0dcf7ba2cfcf · outbound

This paper cites mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:42.097293Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:42.097293Z digest=sha256:15bfe19a1b6a3caa0e580bb8dcd81335ca90463ad4a59ad5a2c45edea2ff3c90

Observation 0f9d426d-bf5f-4d1d-a70a-ea7ea25b8dd6 · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-06T04:49:45.086855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:41.911771Z digest=sha256:9a5e3f1e5c79fe6ac9e3948540e3da781321e60dfc8870b395cd0991dd9c110e

Observation 34e11361-29a3-4833-90e7-58ddcf13c67e · outbound

This paper cites A Simple LLM Framework for Long-Range Video Question-Answering.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering A Simple LLM Framework for Long-Range Video Question-Answering

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:42.312571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:42.312571Z digest=sha256:a729d60cd9e39ca1cd36e42328456448134b84f6630116f600d63194ae68e768

Observation e28f971e-4ef0-4741-b281-96adc07bbee8 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:42.423364Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:42.423364Z digest=sha256:575dc0f44918b1c7e6e1c81ec7b5f99df799e0ff3815b910f6ee8a6f2cfa34c2

Observation 0a16f30d-b100-4137-849c-bc3eb44ebdfa · outbound

This paper cites Self-Chained Image-Language Model for Video Localization and Question Answering.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Self-Chained Image-Language Model for Video Localization and Question Answering

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:42.185509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:42.185509Z digest=sha256:cb3b0418e7a5016c00b88b5aee8c23a3b46ac9b0609c95d87005886c075a8aee

Observation 9efbab19-1769-48cb-9f6a-b522e115e48d · outbound

This paper cites HACS: Human Action Clips and Segments Dataset for Recognition and Temporal Localization.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering HACS: Human Action Clips and Segments Dataset for Recognition and Temporal Localization

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:42.667270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:42.667270Z digest=sha256:e30d0f996a44e0befdd5dc30fe0b3fa1189c106bfe0f5f274e2afb08335168e5

Observation 22477e2d-58ab-4c4a-bd9a-878b1ce5b9da · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 49

Resolution
malformed identifier
no resolver link, observed 2026-08-06T04:49:42.770981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:42.770981Z digest=sha256:a28c334c61cffac0e871df165631cb0dc0a45da766b1cbdbc4a60f2386d79daa

Observation 274ddc68-04d3-4e29-b61a-c2eefc215d8d · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:42.523814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:42.523814Z digest=sha256:ec046a78cb4b3f2e8f093e88e65dc2657aaf697f2e8f4a7dd8033a41cdc0a022

Observation bb1fe118-e6d6-4eb4-a643-551b0be153ba · outbound

This paper cites an unresolved cited work.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering Unresolved cited work

Reference 53

Resolution
unresolved
raw_fallback, observed 2026-08-06T04:49:44.985780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:42.909084Z digest=sha256:c9d70f2304b8eb88da7bbb14f03e6cfb65ed09d49b1b03adbb7d16bb7418ee55

Observation ea88ecab-50cd-4258-909f-1cbeec8dc364 · outbound

This paper cites https://api.semanticscholar.org/CorpusID:1710722.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering https://api.semanticscholar.org/CorpusID:1710722

Reference 2015

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:49:45.879225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-06T04:49:38.295335Z digest=sha256:02e5043a44d5fde3e4f30ba847c3f62c9cf336b1662640cf761772a3c316a652

Observation 7d166b95-e89a-4822-9a73-8309863d4e2a · outbound

This paper cites VideoVista: A Versatile Benchmark for Video Understanding and Reasoning.

VideoForest: Person-Anchored Hierarchical Reasoning for Cross-Video Question Answering VideoVista: A Versatile Benchmark for Video Understanding and Reasoning

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T04:49:39.159644Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:49:39.159644Z digest=sha256:3a5dcff4c680904b003b261d4b8a4455848cda3aecdf2b2ae5874a74e76a9538

Pith citing papers

No inbound Pith citation observations are available.