Pith. sign in

Paper Citation Record · LEDGER

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes

As of 17 August 2026, this Paper Citation Record lists 60 of 60 outbound references and 0 inbound Pith citation observations for arXiv:2505.03581.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.03581 v1

Coverage vector

measured 60 of 60 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:50:44.081050Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

60 of 60 outbound references displayed

  • verified exact3
  • verified fuzzy21
  • unresolved36
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2f0771a3-f1ca-4c86-8055-a9b79bb5f917 · outbound

This paper cites Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Conceptgraphs: Open-vocabulary 3d scene graphs for perception and planning,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.709445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:43.879396Z digest=sha256:7cd149bd92ce454bea75a8b66f0ea0119a9adb1f389629e899d4b277c42a8188

Observation 4c7416d5-173b-4e37-be71-32794146e0e3 · outbound

This paper cites Beyond bare queries: Open-vocabulary object retrieval with 3d scene graph,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Beyond bare queries: Open-vocabulary object retrieval with 3d scene graph,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.700268Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:43.883967Z digest=sha256:982163a4d833b2db81a0024e86ce14c9a4e9ff53e2efbc38f2468bf7dd42cfce

Observation 84232814-3510-4da1-a4f1-dd96325f2251 · outbound

This paper cites Search3d: Hierarchical open-vocabulary 3d segmenta- tion,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Search3d: Hierarchical open-vocabulary 3d segmenta- tion,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.689789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:43.888014Z digest=sha256:705c78a10f5fa06c896418a852989ed328d5c8474ac4f2cef8f28981a407aef1

Observation d7074e35-bd7b-4e54-99cb-ece5e1f239d7 · outbound

This paper cites Hierarchical open-vocabulary 3d scene graphs for language-grounded robot navigation,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Hierarchical open-vocabulary 3d scene graphs for language-grounded robot navigation,

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.891225Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.891225Z digest=sha256:966f92cd604dd07814e1e492b0878946f63706832cf6f7e0713f46af0e5bf10d

Observation b7d94e18-e615-42c0-b8b0-51c029add4f6 · outbound

This paper cites Clio: Real-time task-driven open-set 3d scene graphs,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Clio: Real-time task-driven open-set 3d scene graphs,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.895182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.895182Z digest=sha256:8a962f6434aad18d75424f294a76b70e9b91c6fc7702f4476605608400d93b7e

Observation b9494be8-c440-42a5-b3bf-6204d5c977cd · outbound

This paper cites 4d panoptic scene graph generation,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes 4d panoptic scene graph generation,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.668082Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:43.898419Z digest=sha256:ba933df6ca89aedd528c78988cb45739fe95efea82d6ae92426789432737a680

Observation 46c15d20-7df7-4f50-a137-3da68dc4bdce · outbound

This paper cites G-retriever: Retrieval-augmented generation for textual graph understanding and question answering,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes G-retriever: Retrieval-augmented generation for textual graph understanding and question answering,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.658325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:43.902087Z digest=sha256:5f41278aa6ab85c1862c08aa17b45cb4f4ace8e45bcb4ebecda25e9188d48e68

Observation 645bd11b-dc08-4e04-8761-17006d2928db · outbound

This paper cites Let Your Graph Do the Talking: Encoding Structured Data for LLMs.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Let Your Graph Do the Talking: Encoding Structured Data for LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.905097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.905097Z digest=sha256:986642f96a646869e3f1c1ecbed5df711e8c97b28bad98c16eca711c6df8fcb2

Observation e8cc9849-0fb3-46d9-bcf3-73b8e8b9bccb · outbound

This paper cites Can llms enhance performance prediction for deep learning models?.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Can llms enhance performance prediction for deep learning models?

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.647941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:43.908611Z digest=sha256:0806e5ba2158dd32cb756ad14eb496810d69ae76266999573ff3286d4c2ebf5b

Observation 8128cca7-1eee-44b0-a647-76d99cacdfdf · outbound

This paper cites STAR: A Benchmark for Situated Reasoning in Real-World Videos.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.912140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.912140Z digest=sha256:ae0d57e10be940829bf529773a9e7dfef31d8dae20f83dcdb101ccc45b567c22

Observation 8c57e5aa-525e-46ac-9490-95e8720d9733 · outbound

This paper cites Agqa 2.0: An updated benchmark for compositional spatio-temporal reasoning,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Agqa 2.0: An updated benchmark for compositional spatio-temporal reasoning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.637619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:43.915839Z digest=sha256:81d84ccf8502c8fb5c9d440aab4104ba73b1ee8b9af5f1ad7b4238115088b0d9

Observation 1885515b-95f2-4100-8240-6e09a118c67a · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Visual genome: Connecting language and vision using crowdsourced dense image annotations,

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.922505Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.922505Z digest=sha256:6ee9b4885410bb98a511ac6907a56b7ab01665e34156a4b7954526fd3bfaefdd

Observation f758b0d9-18c6-47ca-b6e9-be1bb9780696 · outbound

This paper cites Gqa: A new dataset for real- world visual reasoning and compositional question answering,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Gqa: A new dataset for real- world visual reasoning and compositional question answering,

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.925544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.925544Z digest=sha256:c1e6db1a869d84bf6d545c75b1affdb1e1827f697214bba21a96b240c8ad29e6

Observation 9a767cad-a860-4189-86b5-bef289d6efa8 · outbound

This paper cites Panoptic scene graph generation,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Panoptic scene graph generation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.618448Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:43.928530Z digest=sha256:62c230524b138000a4af45635653ff8bce807c809cfba938b22faabc1b09542c

Observation 65521ffd-71c5-42e0-b747-d37d6c798f85 · outbound

This paper cites Action genome: Actions as compositions of spatio-temporal scene graphs,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Action genome: Actions as compositions of spatio-temporal scene graphs,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.609510Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:43.931740Z digest=sha256:b74ec7c5f611399c79a1dc3257dc49cca6308d82a489ed45aa9d62970812ec0a

Observation 0fed5c8e-ac80-492e-a37b-dd2684cfbfa4 · outbound

This paper cites Panoptic video scene graph generation,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Panoptic video scene graph generation,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.600160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:43.934810Z digest=sha256:329f5d2133279c73dfd34438f2204a441e48a5e1be71fb637996141436c37839

Observation a090021c-ddbf-4edd-b9a9-393dcf6de988 · outbound

This paper cites Egtr: Extracting graph from transformer for scene graph generation,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Egtr: Extracting graph from transformer for scene graph generation,

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.938041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.938041Z digest=sha256:7c6c9523ceff0b889194c8c520634b760546a5fbf2fab01ba1732ef31c2fcc12

Observation 30a67c2d-1411-459d-a581-f32df200fe56 · outbound

This paper cites Reltr: Relation transformer for scene graph generation,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Reltr: Relation transformer for scene graph generation,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.941400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.941400Z digest=sha256:6e6345ec6732b1092185bc42f827ab148ee27019aca7e66802d95eae5e327b31

Observation a6806e38-1f43-436e-a43a-bf1604f5a723 · outbound

This paper cites Oed: towards one-stage end-to- end dynamic scene graph generation,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Oed: towards one-stage end-to- end dynamic scene graph generation,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.580376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:43.945053Z digest=sha256:ebc7040d41e3beff1438d021527b0ac1b6b1aaf24b1864bcf8c59367ba8f1232

Observation 1605903d-9619-4b17-b3b2-2cfbf2bf3ee1 · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes SAM 2: Segment Anything in Images and Videos

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.948634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.948634Z digest=sha256:27db8eff5cc557bef0e2f0f0732fa06d243da110480b9057320520f51ab1230c

Observation 7e316375-8fa3-404a-83ff-1832345ffe1f · outbound

This paper cites Llavanext: Improved reasoning, ocr, and world knowledge,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Llavanext: Improved reasoning, ocr, and world knowledge,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.952811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.952811Z digest=sha256:64d55d48637f671248e4d92305007ea68cf7efe397088ea914a3453630f7c4c7

Observation 22188316-e050-47d6-9f6f-be776b314683 · outbound

This paper cites Yolo-world: Real-time open-vocabulary object detection,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Yolo-world: Real-time open-vocabulary object detection,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.956865Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.956865Z digest=sha256:76e36d84cc3c06ac034a936f3d1ae126ea0539f867fd2153b679df7ba349eed4

Observation e3418715-8ace-498d-bcf7-ef9b3d605286 · outbound

This paper cites GPT-4 Technical Report.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes GPT-4 Technical Report

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.960371Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.960371Z digest=sha256:7e7fda200d0c714b992da2f0e71006ddc9edaf2125998d5e04e9174a91e82c03

Observation 5c725269-06a3-4ead-ae5e-0a3f44765277 · outbound

This paper cites Yi: Open Foundation Models by 01.AI.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Yi: Open Foundation Models by 01.AI

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.963321Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.963321Z digest=sha256:4e33e2434e324d83c6ed9c1be96c28b13b2ae035a85066d7f8e52d6ea1cbc7a5

Observation 77f7d309-02d3-4c7c-b892-64d46f8a505c · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes NVILA: Efficient Frontier Visual Language Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.967414Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.967414Z digest=sha256:69c580a6f73de2299e5c7db9986ad2396f331fda181cc1a56cce961a099196dd

Observation 740d89b9-7117-4b46-a85b-ff8648391f89 · outbound

This paper cites Qwen2.5-VL Technical Report.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Qwen2.5-VL Technical Report

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.970524Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.970524Z digest=sha256:18f7347ce0a047c40e10bd6d6702302cae4122ff1bda16d739c262acd379cfa7

Observation 8a8e3370-8c8f-4ecf-ad26-162d9a30876b · outbound

This paper cites Graph-Based Captioning: Enhancing Visual Descriptions by Interconnecting Region Captions.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Graph-Based Captioning: Enhancing Visual Descriptions by Interconnecting Region Captions

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.973960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.973960Z digest=sha256:cd57a904143ac5983362f021393e0d1efcfae51efa7c9d14dfcccaa0494646f7

Observation 04a36963-6c86-4e53-87ab-7c9073ebd39a · outbound

This paper cites ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes ProVision: Programmatically Scaling Vision-centric Instruction Data for Multimodal Language Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.977587Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.977587Z digest=sha256:f8228f319e88f26f6efdc2363c798cd0e71bf69f4665a8492931a82501ea6514

Observation f6eeaa08-33b7-48ca-b31f-e80b1c3944c1 · outbound

This paper cites Llm4sgg: large language models for weakly supervised scene graph generation,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Llm4sgg: large language models for weakly supervised scene graph generation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.560483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:43.980913Z digest=sha256:74e9bee3b734e846b33c14b3de187d3a3d542c5533558ecaa712b158826d3d6a

Observation 93d64783-f5cc-43c8-b905-59f1660eccb1 · outbound

This paper cites Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Thinking in Space: How Multimodal Large Language Models See, Remember, and Recall Spaces

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.984048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.984048Z digest=sha256:367a70e9dcde3e63e92010361a050907283959424d6d96a8819868ac79ca92b5

Observation db1551b9-ffdd-478d-b7fb-7678a1a290c6 · outbound

This paper cites From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes From Seconds to Hours: Reviewing MultiModal Large Language Models on Comprehensive Long Video Understanding

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.987304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.987304Z digest=sha256:cd1c0c3848570da0ad9d57500d6237f647d0d616f8550f8c6e80545820f93465

Observation 9818144f-18fa-476e-a85b-1e19b57e3fa5 · outbound

This paper cites Visual Large Language Models for Generalized and Specialized Applications.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Visual Large Language Models for Generalized and Specialized Applications

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.990891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.990891Z digest=sha256:27c373f935b9c1c2b2aeee578a8acb78dc354cc6846506dab7466eb5b8e0eb9d

Observation 3d0cf3e8-42d0-4d39-a3f8-79333589329e · outbound

This paper cites (2.5+ 1) d spatio- temporal scene graphs for video question answering,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes (2.5+ 1) d spatio- temporal scene graphs for video question answering,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.551272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:43.993936Z digest=sha256:c55e033d40dd347debe1720620b3c5150be7f71db5ed44cfa1350928740acfca

Observation 4cc7d048-e096-4191-b2f3-254a8c9b05f7 · outbound

This paper cites Action scene graphs for long-form understanding of egocentric videos,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Action scene graphs for long-form understanding of egocentric videos,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.541815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:43.997532Z digest=sha256:c87f1626555d3644de6712abf2bd80f322aa12049e1512cb8c83c5040739c849

Observation 1b858b1f-bfeb-47f1-b0bd-5e31f8284cb9 · outbound

This paper cites HyperGLM: HyperGraph for Video Scene Graph Generation and Anticipation.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes HyperGLM: HyperGraph for Video Scene Graph Generation and Anticipation

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-15T23:50:44.318336Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:44.000451Z digest=sha256:fcf75ccb889639ad7a14bab42569dae9a1d492ef08d956c79f376a4faefe4982

Observation d0f2412a-688e-4361-b79c-5c67e17a3d8c · outbound

This paper cites STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes STEP: Enhancing Video-LLMs' Compositional Reasoning by Spatio-Temporal Graph-guided Self-Training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.003979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.003979Z digest=sha256:895261cd74b5c0ea03b3ebace49ed68e1fafcb56f42c28de3655915323319298

Observation a1abcacc-c6b9-435a-a869-a612ea80fcd2 · outbound

This paper cites Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Smarter, Better, Faster, Longer: A Modern Bidirectional Encoder for Fast, Memory Efficient, and Long Context Finetuning and Inference

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.007182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.007182Z digest=sha256:7bd14752f59a25151956a926d48539056221a6d237ced4ae93597448c7eb5c96

Observation d4250c15-2b0a-4a16-ae60-b28dee74ec5e · outbound

This paper cites Benchmarking graph neural networks,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Benchmarking graph neural networks,

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.010411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.010411Z digest=sha256:81e897ed93ddd3a6165d67804c46597ac5e1736bab06639d05a625558630e607

Observation d938cc35-c39c-4172-ac51-f9f679921791 · outbound

This paper cites Masked Label Prediction: Unified Message Passing Model for Semi-Supervised Classification.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Masked Label Prediction: Unified Message Passing Model for Semi-Supervised Classification

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.013299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.013299Z digest=sha256:284ca352650a821e9b5ee409a088b3de94e94161c90e0a95894fbd4148d6e86c

Observation bd3c880a-c483-4c30-9183-6ac3d7a94fd0 · outbound

This paper cites Roformer: En- hanced transformer with rotary position embedding,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Roformer: En- hanced transformer with rotary position embedding,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.016593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.016593Z digest=sha256:1faea35f0eecc40919d00b562d377857c1e289d806830d6a23482e580f385b82

Observation a12445ac-7fcb-4ea8-88e4-d207e0a54489 · outbound

This paper cites Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Blip-2: Bootstrapping language- image pre-training with frozen image encoders and large language models,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.019556Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.019556Z digest=sha256:5f40dc0ff67beb2a7528b45e86ca404bd1064932c8fddfa63e18670f23eac638

Observation 8cd2c11b-fd18-4e51-806c-f2e36ee7cb5b · outbound

This paper cites Agqa: A benchmark for compositional spatio-temporal reasoning,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Agqa: A benchmark for compositional spatio-temporal reasoning,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.516269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:44.022635Z digest=sha256:7a9bddb019fe7f5e72eab9fcd3ed35063ec6c0c5fbc79e0f501cbee8d84aa87c

Observation c168e277-342d-4817-9a93-37fe007677cd · outbound

This paper cites Lora: Low-rank adaptation of large language models.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Lora: Low-rank adaptation of large language models

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.025383Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.025383Z digest=sha256:6951fbfe4bd149b9a506d39c1f3a70f16e30c63d869655d0897b8558225bf9b3

Observation c8f94f6c-05ae-411a-8abb-bfb10691e8ad · outbound

This paper cites The Llama 3 Herd of Models.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes The Llama 3 Herd of Models

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.029059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.029059Z digest=sha256:465025f6ef080cdbaf3da2fe7e038f8e10a4948442f7828b0121d70755649fe7

Observation 85b94c4e-2a7b-44a0-9418-7c43271d08d7 · outbound

This paper cites Do We Really Need Complicated Model Architectures For Temporal Networks?.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Do We Really Need Complicated Model Architectures For Temporal Networks?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.032856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.032856Z digest=sha256:55e4b39221bc39469d147ca2f5a87583ed8ad295ef7509c0a7aed2c9c2ee3981

Observation fc2ae432-b86d-49e9-82a7-11240a495aa5 · outbound

This paper cites Attention is all you need,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Attention is all you need,

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.036476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.036476Z digest=sha256:7b17fd606fdeed22bf4d7eb061fdfdd1e26e6a741fc67b19177b232c875549c4

Observation ea23b79a-d636-4a5f-9585-088fc0a12db6 · outbound

This paper cites Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Question-Instructed Visual Descriptions for Zero-Shot Video Question Answering

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.040102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.040102Z digest=sha256:ac84d2daf4b7f307a2207179ca233724a7338f7557ebfd24d63629d58aa0bab8

Observation f3d640f6-5953-462d-85d4-55aaab99866f · outbound

This paper cites Mist: Multi-modal iterative spatial-temporal transformer for long-form video question answering,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Mist: Multi-modal iterative spatial-temporal transformer for long-form video question answering,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.495603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:44.044120Z digest=sha256:0d816c0f55be1d7d1fa64e8ab0958d3b9d43d4f4ee0e4392ef95edfed5b1cd34

Observation 41a135df-9a6f-469b-b794-3ed9cdc277a5 · outbound

This paper cites Self-chained image-language model for video localization and question answering,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Self-chained image-language model for video localization and question answering,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.485545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:44.047370Z digest=sha256:e7e7b8526e004cdc6206bc368d408ab52e556c238068fa539fd3e41536907720

Observation 219cc9f7-2e91-40e7-a4f4-2e2d071e4456 · outbound

This paper cites Vila: Efficient video-language alignment for video question answering,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Vila: Efficient video-language alignment for video question answering,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.476300Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:44.051426Z digest=sha256:0bb8ff38743ad3b03cfe269f3567c19bba7988541fb5ca0b546b32ed14af2596

Observation 17ca962d-91c5-44f7-ae3e-ab8089b6aac4 · outbound

This paper cites End-to-End Video Question Answering with Frame Scoring Mechanisms and Adaptive Sampling.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes End-to-End Video Question Answering with Frame Scoring Mechanisms and Adaptive Sampling

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.054394Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.054394Z digest=sha256:ad3b734575668c594c69239a7298e94107dc16bd82672589f2c0cdd679393823

Observation 9c57fd8c-d894-4d96-baa6-3b39df3c6587 · outbound

This paper cites Look, Remember and Reason: Grounded reasoning in videos with language models.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Look, Remember and Reason: Grounded reasoning in videos with language models

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-08-15T23:50:44.244898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:44.057798Z digest=sha256:28bd00515ce92117b68c593bc434deebf5bc6602b6fca73ab9cde5e5a188efff

Observation 9fe86359-9402-4674-81ae-90c404ae6f0a · outbound

This paper cites Glance and focus: Memory prompt- ing for multi-event video question answering,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Glance and focus: Memory prompt- ing for multi-event video question answering,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.465775Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:44.060839Z digest=sha256:f659ed7109f203b08be20b1e9f78f4b5e6675876257d16b3477121fdd59d4bef

Observation 204bab2a-2e42-4642-97bb-685fc96fff48 · outbound

This paper cites Learning to reason iteratively and parallelly for complex visual reasoning scenarios,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Learning to reason iteratively and parallelly for complex visual reasoning scenarios,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.456171Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:44.064372Z digest=sha256:d48442a83f92c1c776ae83f37993ad475f7da432f935de6daf9f3a8040a0b7d1

Observation 7e94c167-1768-4650-bcc6-10a6cde4d138 · outbound

This paper cites Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding Bridge.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Efficient Temporal Extrapolation of Multimodal Large Language Models with Temporal Grounding Bridge

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.067588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.067588Z digest=sha256:bf6780d4f2b14c600b1ade9fcdd9c2cc03719828fb1a077b8bb49fe5e7179bb6

Observation fec0ce2d-c820-44ef-b6b5-a7500e15d60f · outbound

This paper cites Learning Fine-Grained Visual Understanding for Video Question Answering via Decoupling Spatial-Temporal Modeling.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes Learning Fine-Grained Visual Understanding for Video Question Answering via Decoupling Spatial-Temporal Modeling

Reference 56

Resolution
verified exact
local_arxiv, observed 2026-08-15T23:50:44.134813Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:44.071264Z digest=sha256:5965ccdfb4a320dd2f5027d342c930cd8bf22dab04640784736898364923b379

Observation e527ce8f-8f68-4236-8d92-491033c3c6e2 · outbound

This paper cites G-Retriever: Retrieval-Augmented Generation for Textual Graph Understanding and Question Answering.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes G-Retriever: Retrieval-Augmented Generation for Textual Graph Understanding and Question Answering

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.074558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.074558Z digest=sha256:c6a314cbdfb36007f43fd0e84ec42d7ec894eac3dd89df49802867a1c419820a

Observation 49832ac9-9274-4e89-9d63-778c489cd185 · outbound

This paper cites A note on the prize collecting traveling salesman problem,.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes A note on the prize collecting traveling salesman problem,

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:50:44.445522Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-15T23:50:44.078054Z digest=sha256:9a067f2fc4ccce9cbc3152392c09c47624bf4dcf19c1373651a9f5f4710cb65e

Observation 825f9578-8f73-459c-9c8b-16f2f5cc1acb · outbound

This paper cites FACTUAL: A Benchmark for Faithful and Consistent Textual Scene Graph Parsing.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes FACTUAL: A Benchmark for Faithful and Consistent Textual Scene Graph Parsing

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:44.081050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:44.081050Z digest=sha256:fdea12372f4a114aecf61b8ce69b32117cfb068a8366ad9da7a98876c63bd012

Observation c137e4b7-3635-4d63-91ba-aceac1a5e0ed · outbound

This paper cites AGQA 2.0: An Updated Benchmark for Compositional Spatio-Temporal Reasoning.

DyGEnc: Encoding a Sequence of Textual Scene Graphs to Reason and Answer Questions in Dynamic Scenes AGQA 2.0: An Updated Benchmark for Compositional Spatio-Temporal Reasoning

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-15T23:50:43.919104Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:50:43.919104Z digest=sha256:13ca045074bde37812c4baeea89b706b5b296b78f1153413373b231956295129

Pith citing papers

No inbound Pith citation observations are available.