Pith. sign in

Paper Citation Record · LEDGER

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes

As of 17 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2505.03581.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.03581 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:50:44.081050Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact3
  • verified fuzzy21
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2f0771a3-f1ca-4c86-8055-a9b79bb5f917 · outbound

This paper cites Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.709445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:43.879396Z digest=sha256:29f31c1d0be5640502291f50f11d89e4d4a23d4fe0c7bf0964efab841920ddee

Observation 4c7416d5-173b-4e37-be71-32794146e0e3 · outbound

This paper cites Beyond bare queries: Open-vocabulary object retrieval with 3d scene graph,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Beyond bare queries: Open-vocabulary object retrieval with 3d scene graph,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.700268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:43.883967Z digest=sha256:291a0849820e4310138724f3de09d78bfc540d40ab82ee412855a176efaf8b00

Observation 84232814-3510-4da1-a4f1-dd96325f2251 · outbound

This paper cites Search3d: Hierarchical open-vocabulary 3d segmenta- tion,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Search3d: Hierarchical open-vocabulary 3d segmenta- tion,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.689789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:43.888014Z digest=sha256:076d091dc4f77ca28eabfafca723bf50a159d9c5959e28e0bf64ed0e7559f6f5

Observation d7074e35-bd7b-4e54-99cb-ece5e1f239d7 · outbound

This paper cites Hierarchical open-vocabulary 3d scene graphs for language-grounded robot navigation,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Hierarchical open-vocabulary 3d scene graphs for language-grounded robot navigation,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.891225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.891225Z digest=sha256:996df2267cac11e85feccabd92783386ba7890f0c3617c078711b531b7e9fd8e

Observation b7d94e18-e615-42c0-b8b0-51c029add4f6 · outbound

This paper cites Clio: Real-time task-driven open-set 3d scene graphs,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Clio: Real-time task-driven open-set 3d scene graphs,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.895182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.895182Z digest=sha256:2510b6103019e3928a64fca6a8aaca7c950faf4e275b59f1e8798700070f41b9

Observation b9494be8-c440-42a5-b3bf-6204d5c977cd · outbound

This paper cites 4d panoptic scene graph generation,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes 4d panoptic scene graph generation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.668082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:43.898419Z digest=sha256:66860169a122e5cd4fc2d6f19a87c328946651a8197b8f67ef82ab8ba78e8712

Observation 46c15d20-7df7-4f50-a137-3da68dc4bdce · outbound

This paper cites G-retriever: Retrieval-augmented generation for textual graph understanding and question answering,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes G-retriever: Retrieval-augmented generation for textual graph understanding and question answering,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.658325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:43.902087Z digest=sha256:657ee9f5e6493913a4dd757d785faa360984e0f1bc97c901282e4a56d439b1b3

Observation 645bd11b-dc08-4e04-8761-17006d2928db · outbound

This paper cites Let Your Graph Do the Talking: Encoding Structured Data for LLMs.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Let Your Graph Do the Talking: Encoding Structured Data for LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.905097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.905097Z digest=sha256:da25ab76c9ba47eee55f17c946fcaee255cfa5922021e0fe3c90bbbfd338f921

Observation e8cc9849-0fb3-46d9-bcf3-73b8e8b9bccb · outbound

This paper cites Can llms enhance performance prediction for deep learning models?.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Can llms enhance performance prediction for deep learning models?

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.647941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:43.908611Z digest=sha256:f6a48436d0bc5092b21889b8f94fcb9ffa9e5ddd0380835771a113e16bcb35a9

Observation 8128cca7-1eee-44b0-a647-76d99cacdfdf · outbound

This paper cites STAR: A Benchmark for Situated Reasoning in Real-World Videos.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.912140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.912140Z digest=sha256:a3e2c0b6bba2c8f36ae1e5eb7f7701b7ed41ec42ce11513966f835f54d31ba7f

Observation 8c57e5aa-525e-46ac-9490-95e8720d9733 · outbound

This paper cites Agqa 2.0: An updated benchmark for compositional spatio-temporal reasoning,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Agqa 2.0: An updated benchmark for compositional spatio-temporal reasoning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.637619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:43.915839Z digest=sha256:04aef9c56f4da4edef95f95d1a591975e08eb9b2b76c0238a9536717963beebe

Observation 1885515b-95f2-4100-8240-6e09a118c67a · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Visual genome: Connecting language and vision using crowdsourced dense image annotations,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.922505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.922505Z digest=sha256:5a3ac7828e7fcb721f0c59b9cda550a50a5f90c96dfa6fa7b1bb79f2bcb5fd66

Observation f758b0d9-18c6-47ca-b6e9-be1bb9780696 · outbound

This paper cites Gqa: A new dataset for real- world visual reasoning and compositional question answering,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Gqa: A new dataset for real- world visual reasoning and compositional question answering,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.925544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.925544Z digest=sha256:853ca2846c8401545527156f82fb932b8eb69dc5e2dd2406456924ae59370279

Observation 9a767cad-a860-4189-86b5-bef289d6efa8 · outbound

This paper cites Panoptic scene graph generation,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Panoptic scene graph generation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.618448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:43.928530Z digest=sha256:4031ea696befd553bb59c7d86597fd7e3312c1c8556c36a998dd364f4a902ca6

Observation 65521ffd-71c5-42e0-b747-d37d6c798f85 · outbound

This paper cites Action genome: Actions as compositions of spatio-temporal scene graphs,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Action genome: Actions as compositions of spatio-temporal scene graphs,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.609510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:43.931740Z digest=sha256:79ddf40d1f0300f1f8849a3bcfc0c9adc1386ceb9472c3d7542136804b062f95

Observation 0fed5c8e-ac80-492e-a37b-dd2684cfbfa4 · outbound

This paper cites Panoptic video scene graph generation,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Panoptic video scene graph generation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.600160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:43.934810Z digest=sha256:f534b5ee40d44edd20689f75fd756010dbb2009eff6b20cce7f386eb547347b0

Observation a090021c-ddbf-4edd-b9a9-393dcf6de988 · outbound

This paper cites Egtr: Extracting graph from transformer for scene graph generation,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Egtr: Extracting graph from transformer for scene graph generation,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.938041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.938041Z digest=sha256:6fd8a9f3bf096b0f424593a8aa4d14e1e3c4dd96b6a1f4cbd5a8687cbea99e2a

Observation 30a67c2d-1411-459d-a581-f32df200fe56 · outbound

This paper cites Reltr: Relation transformer for scene graph generation,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Reltr: Relation transformer for scene graph generation,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.941400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.941400Z digest=sha256:364962638af58ac09aa8d3a6d0c9ae83b3ef40de483a8c839668e5e8deabf35b

Observation a6806e38-1f43-436e-a43a-bf1604f5a723 · outbound

This paper cites Oed: towards one-stage end-to- end dynamic scene graph generation,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Oed: towards one-stage end-to- end dynamic scene graph generation,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.580376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:43.945053Z digest=sha256:2b19e5a1705574a14ca379e12d8a15a55cb2ea4f9ec7005755b96461d0d172bc

Observation 1605903d-9619-4b17-b3b2-2cfbf2bf3ee1 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes SAM 2: Segment Anything in Images and Videos

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.948634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.948634Z digest=sha256:075acee18fcb7fdfeebec40760e50f8c666baf653bc4eb82678c2f6e3715c3e8

Observation 7e316375-8fa3-404a-83ff-1832345ffe1f · outbound

This paper cites Llavanext: Improved reasoning, ocr, and world knowledge,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Llavanext: Improved reasoning, ocr, and world knowledge,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.952811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.952811Z digest=sha256:cad95468208b9d8291e2018779bd1ee7eeba58086d283d9b7325d2ce7d4f611b

Observation 22188316-e050-47d6-9f6f-be776b314683 · outbound

This paper cites Yolo-world: Real-time open-vocabulary object detection,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Yolo-world: Real-time open-vocabulary object detection,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.956865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.956865Z digest=sha256:dffc031e044478c14008ba534d0499aef8b11270522dbcbc1898729e22a9fec9

Observation e3418715-8ace-498d-bcf7-ef9b3d605286 · outbound

This paper cites GPT-4 Technical Report.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes GPT-4 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.960371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.960371Z digest=sha256:980ca97dbb88fab7a07f70efa88e19acae249bb395d7cdcbfff1510989253cdf

Observation 5c725269-06a3-4ead-ae5e-0a3f44765277 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Yi: Open Foundation Models by 01.AI

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.963321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.963321Z digest=sha256:bc4d5bc7a2b71dfcc89d4d5a70c8da98d186300e431a902cfbbd36bda9d4425f

Observation 77f7d309-02d3-4c7c-b892-64d46f8a505c · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes NVILA: Efficient Frontier Visual Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.967414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.967414Z digest=sha256:c81cd633e4bfbffb2d20518b4d1f6dd748204355b709f7c44b662ef61cb6c6b7

Observation 740d89b9-7117-4b46-a85b-ff8648391f89 · outbound

This paper cites Qwen2.5-VL Technical Report.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Qwen2.5-VL Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.970524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.970524Z digest=sha256:76ba27d80cf25a84210dcc52a44887adc66273265a2b50e1a089434cb5b804b4

Observation 8a8e3370-8c8f-4ecf-ad26-162d9a30876b · outbound

This paper cites Graph-Based Captioning: Enhancing Visual Descriptions by Interconnecting Region Captions.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Graph-Based Captioning: Enhancing Visual Descriptions by Interconnecting Region Captions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.973960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.973960Z digest=sha256:ede6e2aacb5ae56acbc12e85b9585a4926be6d2df882dbe09205cd0877f5df00

Observation 04a36963-6c86-4e53-87ab-7c9073ebd39a · outbound

This paper cites ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.977587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.977587Z digest=sha256:b5d74dc746247547d7978a59267f5d0f464f18ff7c21df41253c761139f767dd

Observation f6eeaa08-33b7-48ca-b31f-e80b1c3944c1 · outbound

This paper cites Llm4sgg: large language models for weakly supervised scene graph generation,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Llm4sgg: large language models for weakly supervised scene graph generation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.560483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:43.980913Z digest=sha256:bc776721378dec92d09da989116822fd15d98407686c0a2833a7c3023770042f

Observation 93d64783-f5cc-43c8-b905-59f1660eccb1 · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.984048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.984048Z digest=sha256:0e181de8236467eebaad2428c0aaad943baa3b8417327c966720574c63c5ea5d

Observation db1551b9-ffdd-478d-b7fb-7678a1a290c6 · outbound

This paper cites From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.987304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.987304Z digest=sha256:3bfd3b9758cede17c77bc4a1497e7b0adb0a1555fd05c20276bfd63509b54b52

Observation 9818144f-18fa-476e-a85b-1e19b57e3fa5 · outbound

This paper cites Visual Large Language Models for Generalized and Specialized Applications.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Visual Large Language Models for Generalized and Specialized Applications

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.990891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.990891Z digest=sha256:b704e30fabcbfbebf097a36fcfa2f6f84d525149bae96e6e7ddc21c0f0a27ba8

Observation 3d0cf3e8-42d0-4d39-a3f8-79333589329e · outbound

This paper cites (2.5+ 1) d spatio- temporal scene graphs for video question answering,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes (2.5+ 1) d spatio- temporal scene graphs for video question answering,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.551272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:43.993936Z digest=sha256:a43bc6c39899d3b4c0c10827a206669582434bd07ffd62ad78ce450e7e7f5627

Observation 4cc7d048-e096-4191-b2f3-254a8c9b05f7 · outbound

This paper cites Action scene graphs for long-form understanding of egocentric videos,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Action scene graphs for long-form understanding of egocentric videos,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.541815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:43.997532Z digest=sha256:1ed44fa7af96041436c79db59612095f84b1e34229a6734bc88addcf7be662b0

Observation 1b858b1f-bfeb-47f1-b0bd-5e31f8284cb9 · outbound

This paper cites HyperGLM: HyperGraph for Video Scene Graph Generation and Anticipation.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes HyperGLM: HyperGraph for Video Scene Graph Generation and Anticipation

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-15T23:50:44.318336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:44.000451Z digest=sha256:2fa6c60f201e033374a21c024cec7ad2ca9df6172554d2bd3dc81eead39fe5a6

Observation d0f2412a-688e-4361-b79c-5c67e17a3d8c · outbound

This paper cites STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.003979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.003979Z digest=sha256:0ee2338aff272bd3ee5589d4bc43eae9374a9b6bd1bda43ca1b957208b93aade

Observation a1abcacc-c6b9-435a-a869-a612ea80fcd2 · outbound

This paper cites Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.007182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.007182Z digest=sha256:966cb8e11b786fe2329d5d877a40ac1f32336fdbbcd019a9c0b62d8cc64498f3

Observation d4250c15-2b0a-4a16-ae60-b28dee74ec5e · outbound

This paper cites Benchmarking graph neural networks,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Benchmarking graph neural networks,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.010411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.010411Z digest=sha256:037dfe9ec89c5fc2b94863c3e04b1aa68bfa8ef6588c1142f0787a092f089263

Observation d938cc35-c39c-4172-ac51-f9f679921791 · outbound

This paper cites Masked Label Prediction: Unified Message Passing Model for Semi-Supervised Classification.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Masked Label Prediction: Unified Message Passing Model for Semi-Supervised Classification

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.013299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.013299Z digest=sha256:4d703508f1814b864654d31b8c8ea3733a99952005904353f56d9b61e7be5e0f

Observation bd3c880a-c483-4c30-9183-6ac3d7a94fd0 · outbound

This paper cites Roformer: En- hanced transformer with rotary position embedding,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Roformer: En- hanced transformer with rotary position embedding,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.016593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.016593Z digest=sha256:df7ca719f3ef258e1d3ee7a88f3063f29a1181d2db762d3492c307733f70ba5c

Observation a12445ac-7fcb-4ea8-88e4-d207e0a54489 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.019556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.019556Z digest=sha256:570e6a7da0126d1737f13c396ed5b3342858ddcf6eaccdf6caaaa4e3831b6065

Observation 8cd2c11b-fd18-4e51-806c-f2e36ee7cb5b · outbound

This paper cites Agqa: A benchmark for compositional spatio-temporal reasoning,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Agqa: A benchmark for compositional spatio-temporal reasoning,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.516269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:44.022635Z digest=sha256:00c6f77174f87c3e46678ab5b7b98f5b97226f772db0ca815a03874a8c047829

Observation c168e277-342d-4817-9a93-37fe007677cd · outbound

This paper cites Lora: Low-rank adaptation of large language models.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Lora: Low-rank adaptation of large language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.025383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.025383Z digest=sha256:3b8a5b343e4a34c787e9d8801499a3491e9a218aa770930cb7b22a77e48c03f2

Observation c8f94f6c-05ae-411a-8abb-bfb10691e8ad · outbound

This paper cites The Llama 3 Herd of Models.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes The Llama 3 Herd of Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.029059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.029059Z digest=sha256:8a2e72b2ff1dc78333ecc9291fdbb56bc701b6c95abcf987707d76d53c748bdd

Observation 85b94c4e-2a7b-44a0-9418-7c43271d08d7 · outbound

This paper cites Do We Really Need Complicated Model Architectures For Temporal Networks?.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Do We Really Need Complicated Model Architectures For Temporal Networks?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.032856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.032856Z digest=sha256:adfe27a4a070bd28f4ec25527796a38c9a1f9beaa54d236d8954f63b7a54314b

Observation fc2ae432-b86d-49e9-82a7-11240a495aa5 · outbound

This paper cites Attention is all you need,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Attention is all you need,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.036476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.036476Z digest=sha256:6d068ee30f6c7e8dfc95837bb701c853f65b97f48dfa02fb9d62f22c878cffa6

Observation ea23b79a-d636-4a5f-9585-088fc0a12db6 · outbound

This paper cites Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.040102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.040102Z digest=sha256:84247c2aed55a2af935592270ac171249d4b32c1d3ce073669e9532be27279af

Observation f3d640f6-5953-462d-85d4-55aaab99866f · outbound

This paper cites Mist: Multi-modal iterative spatial-temporal transformer for long-form video question answering,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Mist: Multi-modal iterative spatial-temporal transformer for long-form video question answering,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.495603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:44.044120Z digest=sha256:768494592c58ec283512ec26c804e2928e817c821def403f6414d423a0d68682

Observation 41a135df-9a6f-469b-b794-3ed9cdc277a5 · outbound

This paper cites Self-chained image-language model for video localization and question answering,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Self-chained image-language model for video localization and question answering,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.485545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:44.047370Z digest=sha256:ccae7df0928450008bfd225ca893b8a3f4ef9398f518110a313071e28f42bc6b

Observation 219cc9f7-2e91-40e7-a4f4-2e2d071e4456 · outbound

This paper cites Vila: Efficient video-language alignment for video question answering,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Vila: Efficient video-language alignment for video question answering,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.476300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:44.051426Z digest=sha256:0a724f32a9f593d8edd69c11fe856908608aaacb4a55d701933f2331e0d58eb2

Observation 17ca962d-91c5-44f7-ae3e-ab8089b6aac4 · outbound

This paper cites End-to-End Video Question Answering with Frame Scoring Mechanisms and Adaptive Sampling.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes End-to-End Video Question Answering with Frame Scoring Mechanisms and Adaptive Sampling

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.054394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.054394Z digest=sha256:2ccf4597699af230cce58c3f15520431aeead9b75e19462ae7b9c5aa1ebf2eba

Observation 9c57fd8c-d894-4d96-baa6-3b39df3c6587 · outbound

This paper cites Look, Remember and Reason: Grounded reasoning in videos with language models.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Look, Remember and Reason: Grounded reasoning in videos with language models

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-15T23:50:44.244898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:44.057798Z digest=sha256:46992a7e4764d5a2fb78e07456d9a63c0dd60925899c7ea7a5197d21cc26e1e7

Observation 9fe86359-9402-4674-81ae-90c404ae6f0a · outbound

This paper cites Glance and focus: Memory prompt- ing for multi-event video question answering,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Glance and focus: Memory prompt- ing for multi-event video question answering,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.465775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:44.060839Z digest=sha256:deec232183a114158a2a1f5145e79289db6e6cf45dc6f0c861bef8d8aa94ba5f

Observation 204bab2a-2e42-4642-97bb-685fc96fff48 · outbound

This paper cites Learning to reason iteratively and parallelly for complex visual reasoning scenarios,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Learning to reason iteratively and parallelly for complex visual reasoning scenarios,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.456171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:44.064372Z digest=sha256:badcb83b829ec38f775dbcf0534afd0b212db9e39891d17201cb817e8bc2f987

Observation 7e94c167-1768-4650-bcc6-10a6cde4d138 · outbound

This paper cites Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding Bridge.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding Bridge

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.067588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.067588Z digest=sha256:24c5b77f61b3ddc32558f80f78fceebb35e5b8b9439fa5d688bd88bbc79274e2

Observation fec0ce2d-c820-44ef-b6b5-a7500e15d60f · outbound

This paper cites Learning Fine-Grained Visual Understanding for Video Question Answering via Decoupling Spatial-Temporal Modeling.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Learning Fine-Grained Visual Understanding for Video Question Answering via Decoupling Spatial-Temporal Modeling

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-15T23:50:44.134813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:44.071264Z digest=sha256:d1f702096fc5ce1567cdb2fda0fc8e3b8ac227cab4bc94ac89d790725265f121

Observation e527ce8f-8f68-4236-8d92-491033c3c6e2 · outbound

This paper cites G-Retriever: Retrieval-Augmented Generation for Textual Graph Understanding and Question Answering.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes G-Retriever: Retrieval-Augmented Generation for Textual Graph Understanding and Question Answering

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.074558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.074558Z digest=sha256:45e332fe306362885cd32fed02b58d1e5ebfe4d0aaf588e51d94b1b478745ec9

Observation 49832ac9-9274-4e89-9d63-778c489cd185 · outbound

This paper cites A note on the prize collecting traveling salesman problem,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes A note on the prize collecting traveling salesman problem,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.445522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:44.078054Z digest=sha256:f7b096844f49ea84717bcd4b5bfd311019a34399657eac2d15dc199c79ae796a

Observation 825f9578-8f73-459c-9c8b-16f2f5cc1acb · outbound

This paper cites FACTUAL: A Benchmark for Faithful and Consistent Textual Scene Graph Parsing.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes FACTUAL: A Benchmark for Faithful and Consistent Textual Scene Graph Parsing

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.081050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.081050Z digest=sha256:e89c278480d553750373b04427f35bde5dd44ca05572eb4b7adeca738d98f643

Observation c137e4b7-3635-4d63-91ba-aceac1a5e0ed · outbound

This paper cites AGQA 2.0: An Updated Benchmark for Compositional Spatio-Temporal Reasoning.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes AGQA 2.0: An Updated Benchmark for Compositional Spatio-Temporal Reasoning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.919104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.919104Z digest=sha256:308c53c6982348af8e32442de611546bf4ea5acd583dd95aa7f451db8f38639b

Pith citing papers

No inbound Pith citation observations are available.