Pith. sign in

Paper Citation Record · LEDGER

Open World Scene Graph Generation using Vision Language Models

As of 21 August 2026, this Paper Citation Record lists 48 of 48 outbound references and 3 inbound Pith citation observations for arXiv:2506.08189.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08189 v1

Coverage vector

measured 48 of 48 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:22:12.194352Z

measured 51 of 51 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T10:44:08.297552Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-14T20:42:58.205735Z

Reference resolution

48 of 48 outbound references displayed

  • verified exact1
  • verified fuzzy29
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c0fc7967-51b6-447b-b055-61b378c7e6ae · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Open World Scene Graph Generation using Vision Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.087934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.087934Z digest=sha256:a2ab55d487d3f39cd48b617d8a7cc2c629d2ec5dd4fc8dbb8487f22b83a4ca9a

Observation f25714bf-0374-43a8-a38c-dad42bafb371 · outbound

This paper cites GPT4SGG: Synthesizing Scene Graphs from Holistic and Region-specific Narratives.

Open World Scene Graph Generation using Vision Language Models GPT4SGG: Synthesizing Scene Graphs from Holistic and Region-specific Narratives

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.091525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.091525Z digest=sha256:44009437575b10d5d4874d90bba6642d70dbd28a13f81b0b4ea36bff55d6f98e

Observation 58085406-ab58-40f7-89c5-380f19210270 · outbound

This paper cites Expanding scene graph boundaries: fully open-vocabulary scene graph generation via visual-concept alignment and retention.

Open World Scene Graph Generation using Vision Language Models Expanding scene graph boundaries: fully open-vocabulary scene graph generation via visual-concept alignment and retention

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.620025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.094288Z digest=sha256:f9224039136192cb1f283da50a6451bf06ddf72848002f9211520707403d40f9

Observation 12ae2883-1520-484f-a502-e06b47a57183 · outbound

This paper cites Reltr: Relation transformer for scene graph generation.IEEE Trans- actions on Pattern Analysis and Machine Intelligence, 45(9): 11169–11183, 2023.

Open World Scene Graph Generation using Vision Language Models Reltr: Relation transformer for scene graph generation.IEEE Trans- actions on Pattern Analysis and Machine Intelligence, 45(9): 11169–11183, 2023

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.612736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.096830Z digest=sha256:50d885f31f13f4f8357877d86b178038b36f726ba1b239706fc54793851c278c

Observation 347bb022-99ea-4d76-b760-7e683b7f98cd · outbound

This paper cites Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models.

Open World Scene Graph Generation using Vision Language Models Molmo and PixMo: Open Weights and Open Data for State-of-the-Art Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.099201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.099201Z digest=sha256:753bd62995f8bc50d7e4ce52901740c172cc565d6052289441c27a3d9b466072

Observation ab1b86b8-b2a4-407e-8684-09326248bc2c · outbound

This paper cites Prism-0: A predicate-rich scene graph genera- tion framework for zero-shot open-vocabulary tasks.arXiv preprint arXiv:2504.00844, 2025.

Open World Scene Graph Generation using Vision Language Models Prism-0: A predicate-rich scene graph genera- tion framework for zero-shot open-vocabulary tasks.arXiv preprint arXiv:2504.00844, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.101904Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.101904Z digest=sha256:b4f07614489c264ddfe59b50d336adbce2e09a88af2f18aab93ac93c58334dc5

Observation 7e752d72-4886-4e04-9aa2-c3206c38b5b6 · outbound

This paper cites SimCSE: Simple Contrastive Learning of Sentence Embeddings.

Open World Scene Graph Generation using Vision Language Models SimCSE: Simple Contrastive Learning of Sentence Embeddings

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.104420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.104420Z digest=sha256:f2cd7bfa11bbadad2b799f3b86985ca3ae0a2495db143b9c940de1a5e216c06c

Observation e4064992-17d0-47ec-9142-2b387f3f5749 · outbound

This paper cites Open-vocabulary Object Detection via Vision and Language Knowledge Distillation.

Open World Scene Graph Generation using Vision Language Models Open-vocabulary Object Detection via Vision and Language Knowledge Distillation

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.107012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.107012Z digest=sha256:5e20397a4c6cf398e1149a76aa5c7f94f7b552c9d36940b84a101e7747831616

Observation 5d2effcd-34b2-470b-8d48-da7c88326306 · outbound

This paper cites To- wards open-vocabulary scene graph generation with prompt- 7 based finetuning.

Open World Scene Graph Generation using Vision Language Models To- wards open-vocabulary scene graph generation with prompt- 7 based finetuning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.605387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.109479Z digest=sha256:5822d1a2491328875c7c1095aad6ad94a593579aabd7abe9d28b9ba3ca6cb4ee

Observation ebe5fd25-ee12-4f92-81b4-9fd2e32e8673 · outbound

This paper cites Scene Graph Reasoning for Visual Question Answering.

Open World Scene Graph Generation using Vision Language Models Scene Graph Reasoning for Visual Question Answering

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.111648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.111648Z digest=sha256:d87a1c04ce80eb5331f5b3d85b64546610b8b5ba647591b4a873827b7b78a8d5

Observation 39fd54ed-26af-4a45-a6aa-a5146ce61db2 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

Open World Scene Graph Generation using Vision Language Models Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.114195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.114195Z digest=sha256:ab7eec6af6f0f80e8ac030e1169f626d352e2fdd57660bf34b040fa90bd96487

Observation b4b5ce4d-ae68-41f6-aab5-5e5c72e934fe · outbound

This paper cites Enhancing scene graph generation with hierarchical relationships and commonsense knowledge.

Open World Scene Graph Generation using Vision Language Models Enhancing scene graph generation with hierarchical relationships and commonsense knowledge

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.594374Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.116526Z digest=sha256:2f5a38fa228319f036056bd2294e0c1fd2ef300f28e2b95933ddee0f13f4a5f9

Observation 552fd6b8-9ebf-4403-a986-1561a10867d6 · outbound

This paper cites Image retrieval using scene graphs.

Open World Scene Graph Generation using Vision Language Models Image retrieval using scene graphs

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.587225Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.118645Z digest=sha256:24def03dd6f97a77e1f5ca265e7f2785f15a811d1b4f2f3e68efb47dbfa7e0b3

Observation 696ec004-e2bf-40ce-9e50-a1f600cdf44b · outbound

This paper cites Scene Graph Generation Strategy with Co-occurrence Knowledge and Learnable Term Frequency.

Open World Scene Graph Generation using Vision Language Models Scene Graph Generation Strategy with Co-occurrence Knowledge and Learnable Term Frequency

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:22:12.250576Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.120736Z digest=sha256:af6fa1aca059b88b1fe9e151b86e64c2dbc0c3c4dd3afcd465c9fedfb0fa6a27

Observation e21e7d9e-d8eb-4955-8b94-c9698e1a00c8 · outbound

This paper cites Llm4sgg: large language models for weakly supervised scene graph generation.

Open World Scene Graph Generation using Vision Language Models Llm4sgg: large language models for weakly supervised scene graph generation

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.579047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.122900Z digest=sha256:db8b4f31ef52f53dc4955472d441fd8118f0edc6df76b77bc1a56d36b9f199ad

Observation c181adb0-57cc-41bd-bad9-4d7bad47c8ea · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017.

Open World Scene Graph Generation using Vision Language Models Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123:32–73, 2017

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.571701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.125074Z digest=sha256:75e27169f7c07b837fd1c2d5439ecb4b141598223dc698b3242137e597704812

Observation 96942250-333e-4d51-94be-96c8e55e2d9c · outbound

This paper cites an unresolved cited work.

Open World Scene Graph Generation using Vision Language Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:22:12.563917Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.126991Z digest=sha256:67c9972af549454cec1bd3ab64e95b166553dfe8a45b9a4723a4ccdd8b3bee43

Observation 234f1add-b521-4887-b86a-e8b758effe78 · outbound

This paper cites Gonzalez, Hao Zhang, and Ion Stoica.

Open World Scene Graph Generation using Vision Language Models Gonzalez, Hao Zhang, and Ion Stoica

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.556567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.129134Z digest=sha256:edf78d659fd4fae43273cdb08df9a555bdf696591287f0776832e36ac8971c57

Observation 7060702d-f3f2-47d5-a9cf-5442e3405a86 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation.

Open World Scene Graph Generation using Vision Language Models Blip: Bootstrapping language-image pre-training for unified vision- language understanding and generation

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.549301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.131213Z digest=sha256:c84783435aa33ff23f9044b4f0d258714194c5aa5419431c262f6b512e63a59c

Observation cc27b857-6ec1-4e1f-82c9-3807eda49d7a · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Open World Scene Graph Generation using Vision Language Models Blip- 2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.133250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.133250Z digest=sha256:d6268d040691ec907e91d058fa5d2c988617dec3490f88474d7fdd1668f7ff97

Observation 5b0f3a94-2ee0-4e36-8300-118e09f4fc59 · outbound

This paper cites Sgtr: End-to- end scene graph generation with transformer.

Open World Scene Graph Generation using Vision Language Models Sgtr: End-to- end scene graph generation with transformer

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.537311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.135454Z digest=sha256:627b4c1d8990e71abd171954955e63f733c0fa05eaa0594ccaa82339a37699d0

Observation a643e086-5fad-41cd-8288-23c100320f15 · outbound

This paper cites From pixels to graphs: Open-vocabulary scene graph generation with vision-language models.

Open World Scene Graph Generation using Vision Language Models From pixels to graphs: Open-vocabulary scene graph generation with vision-language models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.530043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.137587Z digest=sha256:54c8e933bfaf421591352ff027b61a495f80ea453c130080ffcfdd436375beeb

Observation 2d53b965-e8b6-454c-b67f-ab098aab04df · outbound

This paper cites Gps-net: Graph property sensing network for scene graph generation.

Open World Scene Graph Generation using Vision Language Models Gps-net: Graph property sensing network for scene graph generation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.523285Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.139572Z digest=sha256:1fea5f09ce5f70934db45a67470a28c68eb21a87c01cf9cb377ca586f49778a5

Observation e5deb73b-4c2c-4de7-bd78-7351c852f8e3 · outbound

This paper cites Llava-next: Improved reason- ing, ocr, and world knowledge, 2024.

Open World Scene Graph Generation using Vision Language Models Llava-next: Improved reason- ing, ocr, and world knowledge, 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.141548Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.141548Z digest=sha256:13071c67fe35a90929005a161395a81c6682bc03244af0b5a8b75627089c5e66

Observation ba7951ae-a5d1-4c26-98c6-f0aad3cb06fa · outbound

This paper cites Visual instruction tuning.Advances in neural information processing systems, 36, 2024.

Open World Scene Graph Generation using Vision Language Models Visual instruction tuning.Advances in neural information processing systems, 36, 2024

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.143903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.143903Z digest=sha256:64a3bae120e8f0112934e816b79b95962d5c49592c5a98fa0b88d4602f4ed5e3

Observation d68dbdc3-7d4b-4cf5-95da-3c6769641e7d · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

Open World Scene Graph Generation using Vision Language Models Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.146147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.146147Z digest=sha256:ad7b0e28a31f93d5dfc082f6c22934ea139ce0cb20eca4e8310f529412d67010

Observation 4f4b881e-e4f8-4a2b-9b7c-e647a1c58216 · outbound

This paper cites Relation-aware hierarchical prompt for open-vocabulary scene graph generation.

Open World Scene Graph Generation using Vision Language Models Relation-aware hierarchical prompt for open-vocabulary scene graph generation

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.505101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.148127Z digest=sha256:7a0126df59f006ca7ebe582f5d5351ef94c342c17e28e01f2cfb53602dfc659d

Observation 6637806e-20e0-4cd8-9a6c-e7ef2fca36c3 · outbound

This paper cites Visual relationship detection with language priors.

Open World Scene Graph Generation using Vision Language Models Visual relationship detection with language priors

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.498561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.150609Z digest=sha256:b84b2fcefd3b56d6f00aa4041205875efcfb79e7b8869ec96bd1b56378db5da9

Observation 988bbf7c-5eca-4d04-8db9-ae985e5cb788 · outbound

This paper cites hello gpt-4.

Open World Scene Graph Generation using Vision Language Models hello gpt-4

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.491902Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.152786Z digest=sha256:6efdc9441642783acc7aad0b3903ca3b329d83114eb56e1d6d429447e8dbec7b

Observation 0eb9fa80-0d12-430a-a583-b5d2f23f67e6 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Open World Scene Graph Generation using Vision Language Models Learning transferable visual models from natural language supervi- sion

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.154832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.154832Z digest=sha256:1abee4ed0601b09362ba21dc8a725f53f8421e8c8770ff937577fdb1a52f4cf4

Observation 99885462-b79e-4011-8808-e3941a838617 · outbound

This paper cites Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks.

Open World Scene Graph Generation using Vision Language Models Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.156964Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.156964Z digest=sha256:ce3292d4190c0a34490e7871e2515e50b07ae020e3259295787499a0cbedf103

Observation 812dc2d4-3b6a-4774-8498-ebbc3dfc3623 · outbound

This paper cites Learning to compose dynamic tree structures for visual contexts.

Open World Scene Graph Generation using Vision Language Models Learning to compose dynamic tree structures for visual contexts

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.481150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.159122Z digest=sha256:efc85d009fe90b3d340147703d984e04bc047394984b2aa115389676466525bd

Observation 503d2f66-2ac8-4fb3-8522-b3c0e83d734c · outbound

This paper cites Unbiased scene graph generation from bi- ased training.

Open World Scene Graph Generation using Vision Language Models Unbiased scene graph generation from bi- ased training

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.473950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.161392Z digest=sha256:9a151a5258c66ece7bb80e5e5af616fec664f17aa637fe8d62e62009da4ff18d

Observation 65b3abb5-eb71-4ee9-b3a4-86a26d7bf75b · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Open World Scene Graph Generation using Vision Language Models Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.163539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.163539Z digest=sha256:d31362ef319d63bfb3867030cb0d8ed4167bf3a1d703cd67f93648ca49e9f9b4

Observation dd5f0b44-2676-4109-b62b-4b17ab9b1d5e · outbound

This paper cites Graph-structured representations for visual question answer- ing.

Open World Scene Graph Generation using Vision Language Models Graph-structured representations for visual question answer- ing

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.466765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.166021Z digest=sha256:74cf6d37e314f94f188fb70c62b98d7ffcb619dbbae7a2943bb877593d9e7796

Observation f63bdcee-c9a4-4717-8cd2-4cd0d6011ecc · outbound

This paper cites Structured sparse r-cnn for direct scene graph generation.

Open World Scene Graph Generation using Vision Language Models Structured sparse r-cnn for direct scene graph generation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.460126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.168519Z digest=sha256:50488a9c58ac6527b936aac6958ce83878dec433f4fe3392ae81d1ea928b12b4

Observation 95efe330-ac5c-42a4-88bf-5db75ab49aa4 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Open World Scene Graph Generation using Vision Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.170659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.170659Z digest=sha256:f9668b4cc2f328e3ba7878847df1147cf854836846dbc56f9c2a35a54fa04897

Observation de25351e-2698-4140-9294-4e3b9e37c319 · outbound

This paper cites Scene graph generation by iterative message passing.

Open World Scene Graph Generation using Vision Language Models Scene graph generation by iterative message passing

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.453061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.173090Z digest=sha256:eb2d9dc47caf111c712b49b8dea481de24854aad35906ae0193ffe152b97cec4

Observation 61c721cb-7ad3-4e2d-a19c-af7d4c3059f5 · outbound

This paper cites Llava-spacesgg: Visual instruct tuning for open-vocabulary scene graph generation with enhanced spatial relations.

Open World Scene Graph Generation using Vision Language Models Llava-spacesgg: Visual instruct tuning for open-vocabulary scene graph generation with enhanced spatial relations

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.446937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.175109Z digest=sha256:77916a11c6ff53de2dd63253df0a85e38d04b8ce37b75295d23eb64ac355283e

Observation 4c8bca80-1123-4596-8c8e-04df8d28fe3a · outbound

This paper cites Panoptic scene graph gen- eration.

Open World Scene Graph Generation using Vision Language Models Panoptic scene graph gen- eration

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.440066Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.177220Z digest=sha256:0240ac552123518d5ea31ea99ba2746c8a31869d72c822cb5a0039255de06c1d

Observation 9d37b875-eaef-4109-87f9-bdf8b9876244 · outbound

This paper cites Depth anything v2.Advances in Neural Information Processing Systems, 37: 21875–21911, 2024.

Open World Scene Graph Generation using Vision Language Models Depth anything v2.Advances in Neural Information Processing Systems, 37: 21875–21911, 2024

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.433926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.179475Z digest=sha256:4717aa7e69924ba1529d9a85f8fb42847abc27087c34030acf034ae50bc804f3

Observation 8fa9a057-99ad-4258-87da-e326a9954e9a · outbound

This paper cites Cross-modal rela- tionship inference for grounding referring expressions.

Open World Scene Graph Generation using Vision Language Models Cross-modal rela- tionship inference for grounding referring expressions

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.426876Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.181569Z digest=sha256:09d4d362709c5ae313daebff70315ae9815014e3b7ca1d165c8565f0aed50375

Observation ba671b8e-4c7b-42d5-91d1-bb93b9c4bb5c · outbound

This paper cites Visually-prompted language model for fine- grained scene graph generation in an open world.

Open World Scene Graph Generation using Vision Language Models Visually-prompted language model for fine- grained scene graph generation in an open world

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.418720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.183631Z digest=sha256:bef94072de9b3d255ba2b61e42f3c24f686d53adbaf260ab1ac7ff35d3dd12bb

Observation 0d6453c6-b1c5-4294-8718-c055ff3cd14a · outbound

This paper cites Open-vocabulary object detection using captions.

Open World Scene Graph Generation using Vision Language Models Open-vocabulary object detection using captions

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.411258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.185895Z digest=sha256:c77a00f32169cdfc4e74df7a9082eaebc25444874b4e3570c653f7262ffa77bb

Observation db98507f-16df-4147-ae74-ecbab09f577c · outbound

This paper cites Neural Motifs: Scene Graph Parsing with Global Context.

Open World Scene Graph Generation using Vision Language Models Neural Motifs: Scene Graph Parsing with Global Context

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:22:12.187935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:22:12.187935Z digest=sha256:c484f08b8bffacc2dcbb0e3e21be29464b80f3622b535f0fa4ed1cd806e9ec4b

Observation e0d9e044-a009-4de3-a517-797cc1b16b3a · outbound

This paper cites Graphical contrastive losses for scene graph parsing.

Open World Scene Graph Generation using Vision Language Models Graphical contrastive losses for scene graph parsing

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.403580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.190254Z digest=sha256:68e8c24747ecc3c54a92ce1cb37ca0aaeb56d201807409fcd1bf12952378ab82

Observation eb752ca1-0f56-4403-8f17-e3ca98a7e91e · outbound

This paper cites Learning to generate language- supervised and open-vocabulary scene graph using pre-trained visual-semantic space.

Open World Scene Graph Generation using Vision Language Models Learning to generate language- supervised and open-vocabulary scene graph using pre-trained visual-semantic space

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.396290Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.192405Z digest=sha256:f24a64f820d97328a233e7792155bb8f13570eaf1e3024c48807fda0b0cbc472

Observation d22aa359-343e-48d7-ad0c-99ab81d9d6c2 · outbound

This paper cites There is aXin the image.

Open World Scene Graph Generation using Vision Language Models There is aXin the image

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:22:12.389239Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-07T05:22:12.194352Z digest=sha256:a5fe905e886fc449fbebfd9be5ac4ce61a8df5526b207519011d89017392ec11

Pith citing papers

Observation 6e2d48bd-84b8-45e0-be97-51eef825d1f8 · inbound

KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering cites this paper.

KG-ViP: Bridging Knowledge Grounding and Visual Perception in Multi-modal LLMs for Visual Question Answering Open World Scene Graph Generation using Vision Language Models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T10:44:08.297552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T10:44:08.297552Z digest=sha256:ffdee1d44d492df49e76676a7a8ac733a7c0c4252b77520acbc4e6c31072075a

Observation 773895a4-abe5-4930-ac1f-05077e3c49f5 · inbound

SceneGraphVLM: Dynamic Scene Graph Generation from Video with Vision-Language Models cites this paper.

SceneGraphVLM: Dynamic Scene Graph Generation from Video with Vision-Language Models Open World Scene Graph Generation using Vision Language Models

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-14T20:42:58.209549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-05-14T20:39:56.448121Z digest=sha256:6d0169c51ce629b4b633bbfaca2f47ed5d98595e18ef07b172e47083bf271e3e

Observation 2f9aae3f-2b1b-492e-b041-e6c39371b4bf · inbound

GraphVid: Interactive Graph-Controllable Video Generation cites this paper.

GraphVid: Interactive Graph-Controllable Video Generation Open World Scene Graph Generation using Vision Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T07:04:59.338909Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T07:04:59.338909Z digest=sha256:1ee9594e467049f5105fb7098e630d7a6e4191b3ac193a519d8d2f8913db9d63