Pith. sign in

Paper Citation Record · LEDGER

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge

As of 20 August 2026, this Paper Citation Record lists 27 of 27 outbound references and 1 inbound Pith citation observation for arXiv:2506.08553.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08553 v1

Coverage vector

measured 27 of 27 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:11:19.983039Z

measured 28 of 28 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-01T06:21:09.996386Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

27 of 27 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation c2f6542a-0d1b-4646-bb97-d9376f273ce1 · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.316550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:11:19.878734Z digest=sha256:d093bc8f6d2bd6b97ddcce051016964cd850e1f07f7af17602d506cf6d9ca6b1

Observation be1adb73-2497-4037-9299-eb3f6afb6417 · outbound

This paper cites 3d scene graph: A structure for unified semantics, 3d space, and cam- era.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge 3d scene graph: A structure for unified semantics, 3d space, and cam- era

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.304373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:11:19.883010Z digest=sha256:10b06ff055207c9a6bcc422cdb5a9c1651f40cb900c0dbf9ad9cca5dd72120c0

Observation 4c63f293-f3c6-48e6-bcf6-0f1373285c4c · outbound

This paper cites Lamb Artur d’Avila Garcez.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Lamb Artur d’Avila Garcez

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.291760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:11:19.887059Z digest=sha256:3ffac3146b07b8e78cd840899af7e148bb86804145759c598624520ad159c71f

Observation 235cf7a3-ff0b-42d6-9078-43480f14a488 · outbound

This paper cites Comet: Com- monsense transformers for automatic knowledge graph con- struction.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Comet: Com- monsense transformers for automatic knowledge graph con- struction

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.280349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:11:19.891088Z digest=sha256:d0658c59dbd1c8aed9c78ee80e6d87c0b03e0ee04dcb59fc3ca5bf2471a941b7

Observation 21240ddd-d8de-4b24-b25c-01d998c275d9 · outbound

This paper cites Towards Neuro- Symbolic Video Understanding, page 220–236.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Towards Neuro- Symbolic Video Understanding, page 220–236

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.268509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:11:19.895044Z digest=sha256:e1659439fb2b8852533cf6938ccf25fb5e2265f55fa55794b6f91104c7a28d03

Observation 73ff61ec-5855-43d9-b6b6-04b1491bb20a · outbound

This paper cites The epic-kitchens dataset: Collection, challenges and base- lines.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge The epic-kitchens dataset: Collection, challenges and base- lines

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.256974Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:11:19.899376Z digest=sha256:c24f969bd263ee5067c89063b57300fbca21405a223bb5242d1fecb8b5f177ec

Observation 28a4f84c-5e7d-4872-9065-b822f31ad9fe · outbound

This paper cites HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge HuggingGPT: Solving AI Tasks with ChatGPT and its Friends in Hugging Face

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:11:19.903965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:11:19.903965Z digest=sha256:bf78802f5998f44551ef71a5a10b010c211163e3facb5f64f5afd32696b8643b

Observation fa42f414-e5fd-44bc-8d4f-b057a07f5006 · outbound

This paper cites Gemini 2.0 flash.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Gemini 2.0 flash

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.245182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:11:19.908376Z digest=sha256:1f368959a6185016a8779507b21e90430463940acc69325ee1cc028468d09e7d

Observation d8803587-dafd-4457-9d08-c24f064298b6 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Ego4d: Around the world in 3,000 hours of egocentric video

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:11:19.912469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:11:19.912469Z digest=sha256:9c16e5537729d50720a2e55b4953fadd78ca9ea4edf2628aaa078e1aca1707b7

Observation bdd20501-19e8-4cda-9108-a5bf667b1c25 · outbound

This paper cites Learning by asking questions.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Learning by asking questions

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.226951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:11:19.916164Z digest=sha256:35cc8610e6fe160781a2a573413576951f396a2c7c0dac9de9317570546c7897

Observation 81b095c9-5195-41f7-a11d-31a8aefc8eab · outbound

This paper cites Patching open- vocabulary vision models with commonsense.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Patching open- vocabulary vision models with commonsense

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.216102Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:11:19.920387Z digest=sha256:fd341c240945c9596735865c5d5020c8146b3b9adb682a3700400aebe52e0434

Observation 9465e5cb-9df3-4a3a-9914-43f908f42393 · outbound

This paper cites Shamma, Michael Bernstein, and Li Fei-Fei.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Shamma, Michael Bernstein, and Li Fei-Fei

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.205152Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:11:19.924190Z digest=sha256:6280f2824291e38f7bf1cde446d9c396c31f7da010a750af43b984d422aed794

Observation 53b74dbc-71cb-4134-8ecc-e7a59a79144a · outbound

This paper cites an unresolved cited work.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:11:20.194397Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:11:19.927911Z digest=sha256:6180aaf324e591c7aedf57171cbb26677c2898779f84ac9a36d505677c05ae83

Observation e3bc46e9-c2d9-4ca7-a005-b8171900903c · outbound

This paper cites BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge BLIP: Bootstrapping Language-Image Pre-training for Unified Vision-Language Understanding and Generation

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:11:19.932135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:11:19.932135Z digest=sha256:2879cd9ea43daf79793d584701273f6009b512449779be8a1a4e71c7f1915267

Observation f794b1ba-7f42-464f-a881-e32cc9fb03c8 · outbound

This paper cites Neuro-symbolic concept learner: Interpreting scenes, words, and sentences from nat- ural supervision.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Neuro-symbolic concept learner: Interpreting scenes, words, and sentences from nat- ural supervision

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.183469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:11:19.936145Z digest=sha256:c6ac22efd765e1fa7e5a8c413247bc5347ba8c468d4fd221c142652613ee95ab

Observation e30a909d-c479-445e-8313-a936e9425e0d · outbound

This paper cites Augmented common- sense knowledge for remote object grounding.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Augmented common- sense knowledge for remote object grounding

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.172186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:11:19.939700Z digest=sha256:d44fa143e22ee64d63ea1ce57a99ee208381f699c2bb306cb87b23a8c3f75242

Observation 3a6262cb-d8cc-4fd7-8038-7e5750166e15 · outbound

This paper cites Towards unbiased and ro- bust spatio-temporal scene graph generation and anticipa- tion, 2025.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Towards unbiased and ro- bust spatio-temporal scene graph generation and anticipa- tion, 2025

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.160515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:11:19.943784Z digest=sha256:e962d5868c60bce892e51e848790371514108bf5652d91530ec0fa728c94f25f

Observation 2576c7b1-7f67-475a-8946-afb9727e12cd · outbound

This paper cites Pedregosa, G.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Pedregosa, G

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.149184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:11:19.947463Z digest=sha256:b8bdc2972489bf7403733649fb366b5c80f5b96cfa0c7e60611d9ab119812718

Observation eaf39c6b-bb7a-49b8-a893-0c96c71070af · outbound

This paper cites Sentence-bert: Sentence embeddings using siamese bert-networks.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Sentence-bert: Sentence embeddings using siamese bert-networks

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.137882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:11:19.951109Z digest=sha256:61cf175c2b3fa94909023f01e7113e7f577570a3f175a91b9b426abb001b03af

Observation b7d9afb5-6413-4072-9713-88931f8cef68 · outbound

This paper cites Action scene graphs for long- form understanding of egocentric videos, 2023.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Action scene graphs for long- form understanding of egocentric videos, 2023

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.125175Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:11:19.954515Z digest=sha256:77a78fd4977c4eecb06a204ca9e385dd0a87a06b6636f88a91f9b1fb03a1a220

Observation 7063de9e-9a9b-43dd-be6c-d82e2d42bdb8 · outbound

This paper cites Atomic: An atlas of ma- chine commonsense for if-then reasoning.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Atomic: An atlas of ma- chine commonsense for if-then reasoning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.113042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:11:19.959085Z digest=sha256:2ea3abb5d1cca1106913372de5e78a8e870280ac63bb077f448a08bba3f450f1

Observation eb06ebe3-fb83-4fbe-a318-3a0aee9005a0 · outbound

This paper cites Concept- net 5.5: An open multilingual graph of general knowledge.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Concept- net 5.5: An open multilingual graph of general knowledge

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.100797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:11:19.963013Z digest=sha256:21d9138cc4b11e30b66d5ca0c0a0b2ccb6fe050071507deafdb32a1b3613b95b

Observation 61b922e8-d390-45f0-918f-9cc63241f79a · outbound

This paper cites Concept- net 5.5: An open multilingual graph of general knowledge,.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Concept- net 5.5: An open multilingual graph of general knowledge,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.087645Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:11:19.966591Z digest=sha256:5263e2e51dddd9e82f8ef3891ab42c5b80fa5e0df62088ff7870ab0daad2c468

Observation 26dc9ebc-4dac-454d-b468-42cee761571f · outbound

This paper cites Learning situation hyper-graphs for video question answer- ing.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Learning situation hyper-graphs for video question answer- ing

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.074898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:11:19.970671Z digest=sha256:226d1a8d504b65ee1d5901ac029e3dcffa60a0e27c72a24f242bbb1b026c7366

Observation 4f776e03-62e6-4fe9-b272-7deba1911577 · outbound

This paper cites Scene graph generation by iterative message passing.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Scene graph generation by iterative message passing

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.062960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:11:19.975080Z digest=sha256:332a7c7da06baa332a07e5fab8d41d38dcf44ffab68147fb3b550db25713bd35

Observation 957fbb91-98dc-4674-b2db-2fe9e6a755cd · outbound

This paper cites Neuro-symbolic visual reasoning: Disentangling ”vi- sual” from ”reasoning”.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Neuro-symbolic visual reasoning: Disentangling ”vi- sual” from ”reasoning”

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.050199Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:11:19.978859Z digest=sha256:9cff22a8e8326a6b6c81eaf0d2660f3543ac79b5bf30c805f7440e6cb408a2a2

Observation 692377ea-c82e-4a12-bfc1-c2e2479fe93c · outbound

This paper cites Jasper and stella: distillation of sota embedding models,.

From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge Jasper and stella: distillation of sota embedding models,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:11:20.037717Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:11:19.983039Z digest=sha256:ff9686fc78069418e3d8d96512ffac10f53655e5a4f8f9540d45b9f8620570b7

Pith citing papers

Observation 2f522f67-b64e-4061-b04d-5487d7ea1536 · inbound

Learning to Deny: Action Denial in Multimodal Large Language Models cites this paper.

Learning to Deny: Action Denial in Multimodal Large Language Models From Pixels to Graphs: using Scene and Knowledge Graphs for HD-EPIC VQA Challenge

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-07-01T06:25:26.952478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-07-01T06:21:09.996386Z digest=sha256:dafd43810dcbaca6242408bf9905aa7d8b8afc495601df7f168b1f5e5190e8a3