Pith. sign in

Paper Citation Record · LEDGER

Open World Scene Graph Generation using Vision Language Models

As of 21 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 3 inbound Pith citation observations for arXiv:2506.08189.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08189 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:22:12.194352Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T10:44:08.297552Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-14T20:42:58.205735Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact1
  • verified fuzzy29
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c0fc7967-51b6-447b-b055-61b378c7e6ae · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Open World Scene Graph Generation using Vision Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.087934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.087934Z digest=sha256:a2ab55d487d3f39cd48b617d8a7cc2c629d2ec5dd4fc8dbb8487f22b83a4ca9a

Observation f25714bf-0374-43a8-a38c-dad42bafb371 · outbound

This paper cites GPT4SGG: Synthesizing Scene Graphs from Holistic and Region-specific Narratives.

Open World Scene Graph Generation using Vision Language Models GPT4SGG: Synthesizing Scene Graphs from Holistic and Region-specific Narratives

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.091525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.091525Z digest=sha256:44009437575b10d5d4874d90bba6642d70dbd28a13f81b0b4ea36bff55d6f98e

Observation 58085406-ab58-40f7-89c5-380f19210270 · outbound

This paper cites Expanding scene graph boundaries: fully open-vocabulary scene graph generation via visual-concept alignment and retention.

Open World Scene Graph Generation using Vision Language Models Expanding scene graph boundaries: fully open-vocabulary scene graph generation via visual-concept alignment and retention

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.620025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.094288Z digest=sha256:6d1819fafdf0e96eeaae354a9254c30c584a495ab45c2491cffda79a7709d30d

Observation 12ae2883-1520-484f-a502-e06b47a57183 · outbound

This paper cites Reltr: Relation transformer for scene graph generation.IEEE Trans- actions on Pattern Analysis and Machine Intelligence, 45(9): 11169–11183, 2023.

Open World Scene Graph Generation using Vision Language Models Reltr: Relation transformer for scene graph generation.IEEE Trans- actions on Pattern Analysis and Machine Intelligence, 45(9): 11169–11183, 2023

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.612736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.096830Z digest=sha256:2a8227152770f53e2d1cbcd9dfd0ba501622e14e456caffeed83514c3ed5ec49

Observation 347bb022-99ea-4d76-b760-7e683b7f98cd · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Open World Scene Graph Generation using Vision Language Models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.099201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.099201Z digest=sha256:753bd62995f8bc50d7e4ce52901740c172cc565d6052289441c27a3d9b466072

Observation ab1b86b8-b2a4-407e-8684-09326248bc2c · outbound

This paper cites Prism-0: A predicate-rich scene graph genera- tion framework for zero-shot open-vocabulary tasks.arXiv preprint arXiv:2504.00844, 2025.

Open World Scene Graph Generation using Vision Language Models Prism-0: A predicate-rich scene graph genera- tion framework for zero-shot open-vocabulary tasks.arXiv preprint arXiv:2504.00844, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.101904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.101904Z digest=sha256:b4f07614489c264ddfe59b50d336adbce2e09a88af2f18aab93ac93c58334dc5

Observation 7e752d72-4886-4e04-9aa2-c3206c38b5b6 · outbound

This paper cites SimCSE: Simple Contrastive Learning of Sentence Embeddings.

Open World Scene Graph Generation using Vision Language Models SimCSE: Simple Contrastive Learning of Sentence Embeddings

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.104420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.104420Z digest=sha256:f2cd7bfa11bbadad2b799f3b86985ca3ae0a2495db143b9c940de1a5e216c06c

Observation e4064992-17d0-47ec-9142-2b387f3f5749 · outbound

This paper cites Open-vocabulary Object Detection via Vision and Language Knowledge Distillation.

Open World Scene Graph Generation using Vision Language Models Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.107012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.107012Z digest=sha256:5e20397a4c6cf398e1149a76aa5c7f94f7b552c9d36940b84a101e7747831616

Observation 5d2effcd-34b2-470b-8d48-da7c88326306 · outbound

This paper cites To- wards open-vocabulary scene graph generation with prompt- 7 based finetuning.

Open World Scene Graph Generation using Vision Language Models To- wards open-vocabulary scene graph generation with prompt- 7 based finetuning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.605387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.109479Z digest=sha256:f1a8335dc5a9c395940c54a53d8dcdde2824927e69a249dbcde347204194ff5c

Observation ebe5fd25-ee12-4f92-81b4-9fd2e32e8673 · outbound

This paper cites Scene Graph Reasoning for Visual Question Answering.

Open World Scene Graph Generation using Vision Language Models Scene Graph Reasoning for Visual Question Answering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.111648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.111648Z digest=sha256:d87a1c04ce80eb5331f5b3d85b64546610b8b5ba647591b4a873827b7b78a8d5

Observation 39fd54ed-26af-4a45-a6aa-a5146ce61db2 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Open World Scene Graph Generation using Vision Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.114195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.114195Z digest=sha256:ab7eec6af6f0f80e8ac030e1169f626d352e2fdd57660bf34b040fa90bd96487

Observation b4b5ce4d-ae68-41f6-aab5-5e5c72e934fe · outbound

This paper cites Enhancing scene graph generation with hierarchical relationships and commonsense knowledge.

Open World Scene Graph Generation using Vision Language Models Enhancing scene graph generation with hierarchical relationships and commonsense knowledge

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.594374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.116526Z digest=sha256:9b7bddf4db56de90e02fb9db62efc8cd0d14396f0e8976d826d8da8f7577a541

Observation 552fd6b8-9ebf-4403-a986-1561a10867d6 · outbound

This paper cites Image retrieval using scene graphs.

Open World Scene Graph Generation using Vision Language Models Image retrieval using scene graphs

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.587225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.118645Z digest=sha256:37c84761bcf54460a184b43ed21699719b7f7948f627f10f490d8b2622fa70bb

Observation 696ec004-e2bf-40ce-9e50-a1f600cdf44b · outbound

This paper cites Scene Graph Generation Strategy with Co-occurrence Knowledge and Learnable Term Frequency.

Open World Scene Graph Generation using Vision Language Models Scene Graph Generation Strategy with Co-occurrence Knowledge and Learnable Term Frequency

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:22:12.250576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.120736Z digest=sha256:739fb61ed5ecb635b642933b552a4f1f98b6d4ffe07f3a145180813bde227495

Observation e21e7d9e-d8eb-4955-8b94-c9698e1a00c8 · outbound

This paper cites Llm4sgg: large language models for weakly supervised scene graph generation.

Open World Scene Graph Generation using Vision Language Models Llm4sgg: large language models for weakly supervised scene graph generation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.579047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.122900Z digest=sha256:738bc23bde35733f86ed039ac66cb282136d377388e01ff0e7cb7383027619f1

Observation c181adb0-57cc-41bd-bad9-4d7bad47c8ea · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017.

Open World Scene Graph Generation using Vision Language Models Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.571701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.125074Z digest=sha256:9e3d3696b97d4fe4ff31a3fd0fb61117ca55b9dca22fbc956635394b9a5bda94

Observation 96942250-333e-4d51-94be-96c8e55e2d9c · outbound

This paper cites an unresolved cited work.

Open World Scene Graph Generation using Vision Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:22:12.563917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.126991Z digest=sha256:2f589d89ec7633108b6f1dee84b918d76aa0688c616677159e400f18bd3083ff

Observation 234f1add-b521-4887-b86a-e8b758effe78 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Open World Scene Graph Generation using Vision Language Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.556567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.129134Z digest=sha256:474d0e8c908a66372b4587c64340411a7bf1612cc2915fd83d8182f1993c4b3f

Observation 7060702d-f3f2-47d5-a9cf-5442e3405a86 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation.

Open World Scene Graph Generation using Vision Language Models Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.549301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.131213Z digest=sha256:eb5cd2faa32c8cf4fc1fbe777f10313351145b418fde1761c13cfca0ac9c0901

Observation cc27b857-6ec1-4e1f-82c9-3807eda49d7a · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Open World Scene Graph Generation using Vision Language Models Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.133250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.133250Z digest=sha256:d6268d040691ec907e91d058fa5d2c988617dec3490f88474d7fdd1668f7ff97

Observation 5b0f3a94-2ee0-4e36-8300-118e09f4fc59 · outbound

This paper cites Sgtr: End-to- end scene graph generation with transformer.

Open World Scene Graph Generation using Vision Language Models Sgtr: End-to- end scene graph generation with transformer

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.537311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.135454Z digest=sha256:ab69599414443af54a79fab859a3b8aafa7fb93e62cbd7fe836e9a7c05d31b15

Observation a643e086-5fad-41cd-8288-23c100320f15 · outbound

This paper cites From pixels to graphs: Open-vocabulary scene graph generation with vision-language models.

Open World Scene Graph Generation using Vision Language Models From pixels to graphs: Open-vocabulary scene graph generation with vision-language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.530043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.137587Z digest=sha256:d4f75300714f2f86dfc7f89548485fff393d6f12e7710167c790dd30e89c4896

Observation 2d53b965-e8b6-454c-b67f-ab098aab04df · outbound

This paper cites Gps-net: Graph property sensing network for scene graph generation.

Open World Scene Graph Generation using Vision Language Models Gps-net: Graph property sensing network for scene graph generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.523285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.139572Z digest=sha256:fd12dcbe57aca768637ba05947f79b8ac3b117097a98c714c9750a909c42c30c

Observation e5deb73b-4c2c-4de7-bd78-7351c852f8e3 · outbound

This paper cites Llava-next: Improved reason- ing, ocr, and world knowledge, 2024.

Open World Scene Graph Generation using Vision Language Models Llava-next: Improved reason- ing, ocr, and world knowledge, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.141548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.141548Z digest=sha256:13071c67fe35a90929005a161395a81c6682bc03244af0b5a8b75627089c5e66

Observation ba7951ae-a5d1-4c26-98c6-f0aad3cb06fa · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

Open World Scene Graph Generation using Vision Language Models Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.143903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.143903Z digest=sha256:64a3bae120e8f0112934e816b79b95962d5c49592c5a98fa0b88d4602f4ed5e3

Observation d68dbdc3-7d4b-4cf5-95da-3c6769641e7d · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

Open World Scene Graph Generation using Vision Language Models Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.146147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.146147Z digest=sha256:ad7b0e28a31f93d5dfc082f6c22934ea139ce0cb20eca4e8310f529412d67010

Observation 4f4b881e-e4f8-4a2b-9b7c-e647a1c58216 · outbound

This paper cites Relation-aware hierarchical prompt for open-vocabulary scene graph generation.

Open World Scene Graph Generation using Vision Language Models Relation-aware hierarchical prompt for open-vocabulary scene graph generation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.505101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.148127Z digest=sha256:52883956f1c1705fba942e28839e15fa05faf38ec5772c7974f672152617a6be

Observation 6637806e-20e0-4cd8-9a6c-e7ef2fca36c3 · outbound

This paper cites Visual relationship detection with language priors.

Open World Scene Graph Generation using Vision Language Models Visual relationship detection with language priors

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.498561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.150609Z digest=sha256:c16fe9a25dbd830ed5f2a03fd41208ad8eee7b354c096a6da2866aaedb95966a

Observation 988bbf7c-5eca-4d04-8db9-ae985e5cb788 · outbound

This paper cites hello gpt-4.

Open World Scene Graph Generation using Vision Language Models hello gpt-4

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.491902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.152786Z digest=sha256:2e32bd973fbf051a4a281eb23cb44f14c07b6115d67f8ad5f22a1bc70a8e3183

Observation 0eb9fa80-0d12-430a-a583-b5d2f23f67e6 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Open World Scene Graph Generation using Vision Language Models Learning transferable visual models from natural language supervi- sion

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.154832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.154832Z digest=sha256:1abee4ed0601b09362ba21dc8a725f53f8421e8c8770ff937577fdb1a52f4cf4

Observation 99885462-b79e-4011-8808-e3941a838617 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Open World Scene Graph Generation using Vision Language Models Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.156964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.156964Z digest=sha256:ce3292d4190c0a34490e7871e2515e50b07ae020e3259295787499a0cbedf103

Observation 812dc2d4-3b6a-4774-8498-ebbc3dfc3623 · outbound

This paper cites Learning to compose dynamic tree structures for visual contexts.

Open World Scene Graph Generation using Vision Language Models Learning to compose dynamic tree structures for visual contexts

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.481150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.159122Z digest=sha256:3e829a8b68de7a39e7abd5acb4a827120c786115b1160d823fe720efba0ef706

Observation 503d2f66-2ac8-4fb3-8522-b3c0e83d734c · outbound

This paper cites Unbiased scene graph generation from bi- ased training.

Open World Scene Graph Generation using Vision Language Models Unbiased scene graph generation from bi- ased training

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.473950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.161392Z digest=sha256:85d53981b1beb200ed96c23f288d279aa6f3cfd2aa9f9a3d12497c4ce853d653

Observation 65b3abb5-eb71-4ee9-b3a4-86a26d7bf75b · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Open World Scene Graph Generation using Vision Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.163539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.163539Z digest=sha256:d31362ef319d63bfb3867030cb0d8ed4167bf3a1d703cd67f93648ca49e9f9b4

Observation dd5f0b44-2676-4109-b62b-4b17ab9b1d5e · outbound

This paper cites Graph-structured representations for visual question answer- ing.

Open World Scene Graph Generation using Vision Language Models Graph-structured representations for visual question answer- ing

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.466765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.166021Z digest=sha256:e56d5dad2cd415d8c21189104277fe1571cb1df84258e70a357f3b172fbf774d

Observation f63bdcee-c9a4-4717-8cd2-4cd0d6011ecc · outbound

This paper cites Structured sparse r-cnn for direct scene graph generation.

Open World Scene Graph Generation using Vision Language Models Structured sparse r-cnn for direct scene graph generation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.460126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.168519Z digest=sha256:a614179067e625888fa87ec0f023dcb58df760a29ffa4af97c9ceb70d1b70ea0

Observation 95efe330-ac5c-42a4-88bf-5db75ab49aa4 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Open World Scene Graph Generation using Vision Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.170659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.170659Z digest=sha256:f9668b4cc2f328e3ba7878847df1147cf854836846dbc56f9c2a35a54fa04897

Observation de25351e-2698-4140-9294-4e3b9e37c319 · outbound

This paper cites Scene graph generation by iterative message passing.

Open World Scene Graph Generation using Vision Language Models Scene graph generation by iterative message passing

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.453061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.173090Z digest=sha256:89d069fe8fe28f1246c96522781bf4bf830181db7ee5a3e7cd5fe60ca494d504

Observation 61c721cb-7ad3-4e2d-a19c-af7d4c3059f5 · outbound

This paper cites Llava-spacesgg: Visual instruct tuning for open-vocabulary scene graph generation with enhanced spatial relations.

Open World Scene Graph Generation using Vision Language Models Llava-spacesgg: Visual instruct tuning for open-vocabulary scene graph generation with enhanced spatial relations

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.446937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.175109Z digest=sha256:e4f61063ec42a71673d504b0dc286f20e3da4d3c04aa7f2adcb1feac7d8d8864

Observation 4c8bca80-1123-4596-8c8e-04df8d28fe3a · outbound

This paper cites Panoptic scene graph gen- eration.

Open World Scene Graph Generation using Vision Language Models Panoptic scene graph gen- eration

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.440066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.177220Z digest=sha256:51323f062512d1cf6a255a7221f88c14292fa27abb5f6926e0ac0221c981e715

Observation 9d37b875-eaef-4109-87f9-bdf8b9876244 · outbound

This paper cites Depth anything v2.Advances in Neural Information Processing Systems, 37: 21875–21911, 2024.

Open World Scene Graph Generation using Vision Language Models Depth anything v2.Advances in Neural Information Processing Systems, 37: 21875–21911, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.433926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.179475Z digest=sha256:e7f50e8b5689dc67c5455284fff60ecd20690bb1fb22022ee26b81d3e7e87a0c

Observation 8fa9a057-99ad-4258-87da-e326a9954e9a · outbound

This paper cites Cross-modal rela- tionship inference for grounding referring expressions.

Open World Scene Graph Generation using Vision Language Models Cross-modal rela- tionship inference for grounding referring expressions

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.426876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.181569Z digest=sha256:750a7e32050d85840fefdb7ccb57b7ffab15fe18a9a0aa9bdfff78fc98867e4f

Observation ba671b8e-4c7b-42d5-91d1-bb93b9c4bb5c · outbound

This paper cites Visually-prompted language model for fine- grained scene graph generation in an open world.

Open World Scene Graph Generation using Vision Language Models Visually-prompted language model for fine- grained scene graph generation in an open world

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.418720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.183631Z digest=sha256:d82e3e52a4154d99e9fdef60f0be309dbc0753c243c967bbf6ca8884d84b5506

Observation 0d6453c6-b1c5-4294-8718-c055ff3cd14a · outbound

This paper cites Open-vocabulary object detection using captions.

Open World Scene Graph Generation using Vision Language Models Open-vocabulary object detection using captions

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.411258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.185895Z digest=sha256:6fd27693e47bb4cc93fb8337887dc3a0da9497189d38da915ce97b785a0bd3f5

Observation db98507f-16df-4147-ae74-ecbab09f577c · outbound

This paper cites Neural Motifs: Scene Graph Parsing with Global Context.

Open World Scene Graph Generation using Vision Language Models Neural Motifs: Scene Graph Parsing with Global Context

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.187935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.187935Z digest=sha256:c484f08b8bffacc2dcbb0e3e21be29464b80f3622b535f0fa4ed1cd806e9ec4b

Observation e0d9e044-a009-4de3-a517-797cc1b16b3a · outbound

This paper cites Graphical contrastive losses for scene graph parsing.

Open World Scene Graph Generation using Vision Language Models Graphical contrastive losses for scene graph parsing

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.403580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.190254Z digest=sha256:e1ba828eea524ee38d0a0f7e42280bc0c5c1ba5e1480d508e7d846775014b902

Observation eb752ca1-0f56-4403-8f17-e3ca98a7e91e · outbound

This paper cites Learning to generate language- supervised and open-vocabulary scene graph using pre-trained visual-semantic space.

Open World Scene Graph Generation using Vision Language Models Learning to generate language- supervised and open-vocabulary scene graph using pre-trained visual-semantic space

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.396290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.192405Z digest=sha256:2f4992eee6e60c08d73962d2bca957195df2d5df17a2dafbc0f9efbf8edd0381

Observation d22aa359-343e-48d7-ad0c-99ab81d9d6c2 · outbound

This paper cites There is aXin the image.

Open World Scene Graph Generation using Vision Language Models There is aXin the image

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.389239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-07T05:22:12.194352Z digest=sha256:32469a09b6641b816d97ecb032258661e5a83678cbe98983bb01a5cc3f9b0151

Pith citing papers

Observation 6e2d48bd-84b8-45e0-be97-51eef825d1f8 · inbound

KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering cites this paper.

KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering Open World Scene Graph Generation using Vision Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T10:44:08.297552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:44:08.297552Z digest=sha256:ffdee1d44d492df49e76676a7a8ac733a7c0c4252b77520acbc4e6c31072075a

Observation 773895a4-abe5-4930-ac1f-05077e3c49f5 · inbound

SceneGraphVLM: Dynamic Scene Graph Generation from Video with Vision-Language Models cites this paper.

SceneGraphVLM: Dynamic Scene Graph Generation from Video with Vision-Language Models Open World Scene Graph Generation using Vision Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:42:58.209549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-14T20:39:56.448121Z digest=sha256:294d2c45bfbe01d2fa94a5e941f584bd25a3c45a49242aa64713dcbc743f9934

Observation 2f9aae3f-2b1b-492e-b041-e6c39371b4bf · inbound

GraphVid: Interactive Graph-Controllable Video Generation cites this paper.

GraphVid: Interactive Graph-Controllable Video Generation Open World Scene Graph Generation using Vision Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T07:04:59.338909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:04:59.338909Z digest=sha256:1ee9594e467049f5105fb7098e630d7a6e4191b3ac193a519d8d2f8913db9d63