Pith. sign in

Paper Citation Record · LEDGER

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings

As of 9 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 1 inbound Pith citation observation for arXiv:2607.15418.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2607.15418 v1

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-01T23:29:02.328455Z

measured 59 of 59 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T14:28:20.930744Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved58
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 76b7eee8-e047-4e70-b446-990f9ee50122 · outbound

This paper cites Vqa-med: Overview of the medical visual question answering task at imageclef 2019.CLEF (working notes), 2(6):1–11, 2019.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Vqa-med: Overview of the medical visual question answering task at imageclef 2019.CLEF (working notes), 2(6):1–11, 2019

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:55.541852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:55.541852Z digest=sha256:4bae21ebdd26325b12ce64db05b83f3dc4073a75648ec85c9255304b45c9611a

Observation 20382dab-aa51-499d-9783-e570ebdb1495 · outbound

This paper cites GPT-4 Technical Report.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings GPT-4 Technical Report

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:55.639864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:55.639864Z digest=sha256:70b0821b2efb972a24a726f2d4551b0946ede509d3257f6c2b3cebcb298fa7e7

Observation 66eafda0-53fa-4f20-ba74-c92f267e8f35 · outbound

This paper cites OCR, Knowledge, Reasoning (Vi- sual, Alignment).

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings OCR, Knowledge, Reasoning (Vi- sual, Alignment)

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:59.622925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:59.622925Z digest=sha256:ceb9ba2e097d761118400fae1d9b8faaa1d537123f7cadc14c92a356ab5df1a5

Observation 0ff2ebf1-9fc2-4df4-a004-10ea0e7a75d2 · outbound

This paper cites Vqa: Visual question answering.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Vqa: Visual question answering

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:55.807715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:55.807715Z digest=sha256:efe5711969de9e91289e2bcfff10b8e89b268ef0b29a50226139e2cb25d0acbc

Observation ae858427-a851-449d-baad-cb9207fb2196 · outbound

This paper cites Improved baselines with visual instruction tuning, 2023.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Improved baselines with visual instruction tuning, 2023

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:57.125012Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:57.125012Z digest=sha256:a7f5db016f78f7be02f1e38ad792068bb5b3051748dccb2b5438232f268fe310

Observation 7a7c0566-6bc5-4121-ae02-d37742b6c459 · outbound

This paper cites Are large pre-trained vision language models effective construction safety inspec- tors?, 2025.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Are large pre-trained vision language models effective construction safety inspec- tors?, 2025

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:56.035092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:56.035092Z digest=sha256:f88438e4ff135d7da5447790d180b88c94a9ed8a298918756173962344a0e542

Observation 88f1199e-4a9c-44a6-ab65-bf3dcee3e951 · outbound

This paper cites Doris, Daniele Grandi, Ryan Tomich, Md Ferdous Alam, Mohammadmehdi Ataei, Hyunmin Cheong, and Faez Ahmed.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Doris, Daniele Grandi, Ryan Tomich, Md Ferdous Alam, Mohammadmehdi Ataei, Hyunmin Cheong, and Faez Ahmed

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:56.131198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:56.131198Z digest=sha256:03ce6b9fd0c460d3300cdb5916fc40f4f14c7b205f01910ea468f7779b237780

Observation aa62faaf-d6be-412b-b000-b6ab8e44c544 · outbound

This paper cites Galaz- Montoya, Yuhui Zhang, Yuchang Su, Disha Bhowmik, Zachary Coman, Sarina M Hasan, Alexandra Johannesson, William D.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Galaz- Montoya, Yuhui Zhang, Yuchang Su, Disha Bhowmik, Zachary Coman, Sarina M Hasan, Alexandra Johannesson, William D

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:55.871341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:55.871341Z digest=sha256:7e7d0464b0ce76ba747530d71a20b8b64fdcaf2c5cfe22d915f7360c9f6b1bcc

Observation 9f717891-e5f7-4cfe-a76a-ace092f3eda4 · outbound

This paper cites Floorplancad: A large-scale cad draw- ing dataset for panoptic symbol spotting.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Floorplancad: A large-scale cad draw- ing dataset for panoptic symbol spotting

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:56.224966Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:56.224966Z digest=sha256:391b6a4a57fd82831f8ae5964968a9df66a38a3fca129f871d4b237fef180019

Observation 40e0678f-95b9-423a-ae4c-9a7a23386ced · outbound

This paper cites Llava- onevision-1.5: Fully open framework for democratized mul- timodal training, 2025.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Llava- onevision-1.5: Fully open framework for democratized mul- timodal training, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:55.745606Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:55.745606Z digest=sha256:82cbe4821b7b5f5ebdebc294914e008e3398516f88f5713af9e56fa0bc3b5859

Observation fb6ca684-c494-4e48-a087-ab570573eced · outbound

This paper cites Mme-finance: A multimodal finance benchmark for expert-level understanding and rea- soning.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Mme-finance: A multimodal finance benchmark for expert-level understanding and rea- soning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:56.361635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:56.361635Z digest=sha256:f7d7c102e03fac3a9f6eb3523e1126104716e8935c442bcd989bc314a0533666

Observation 711d3b18-7220-444f-b81a-59a6652bf826 · outbound

This paper cites RBench: Graduate-level multi-disciplinary benchmarks for LLM &; MLLM complex reasoning evalu- ation.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings RBench: Graduate-level multi-disciplinary benchmarks for LLM &; MLLM complex reasoning evalu- ation

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:56.469991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:56.469991Z digest=sha256:acc052ea056e16301af8d7599b43128f4d1abfc8f9eae2fdacc00d431a5d48dd

Observation 3494111d-505f-4b23-9129-0ff386ca773f · outbound

This paper cites Omnimedvqa: A new large- scale comprehensive evaluation benchmark for medical lvlm.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Omnimedvqa: A new large- scale comprehensive evaluation benchmark for medical lvlm

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:56.541810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:56.541810Z digest=sha256:cb37c446b79b37ff9cadeffc443e37db51d5964d64a35e930c5c154dfa57d735

Observation ad722b03-8a32-47ca-8a59-13428b3896f9 · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:56.642228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:56.642228Z digest=sha256:181970adde725c957f3153c965f102a2d3f469111ad89faa30874c429da3e316

Observation b53d956d-2528-4347-bc65-1928ae420c61 · outbound

This paper cites an unresolved cited work.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Unresolved cited work

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:56.727583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:56.727583Z digest=sha256:9149bb1eb213731705a7989d5adf7a653a7eabad83a61488f875b5798a8abfb2

Observation 3069b9a2-20c5-444a-b713-17aa6d5a0527 · outbound

This paper cites Seed-bench: Benchmark- ing multimodal large language models.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Seed-bench: Benchmark- ing multimodal large language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:56.780619Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:56.780619Z digest=sha256:45b1aefcb1dad1012a836c61c3bdb5feb78fefcf28daf1a50f1a22cbb82a48b8

Observation 095f1efe-096c-4d8c-848c-b6601d525c74 · outbound

This paper cites Drafterbench: Bench- marking large language models for tasks automation in civil engineering, 2025.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Drafterbench: Bench- marking large language models for tasks automation in civil engineering, 2025

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:56.846980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:56.846980Z digest=sha256:632a2d932e72dd17b27f61275d796d0d8c85121d5ee0be11c69def4d826bd196

Observation 65ae23af-9a24-46c3-970c-a3adf9f4d60b · outbound

This paper cites SceMQA: A scientific college entrance level multimodal question answering benchmark.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings SceMQA: A scientific college entrance level multimodal question answering benchmark

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:56.923864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:56.923864Z digest=sha256:2d6bce93b9810780f871502038e0bd46411235a0fef564cac815ee974b3c561e

Observation 7adec746-f8e4-43aa-9419-43887b6f14a7 · outbound

This paper cites Gemex: A large-scale, groundable, and explainable medical vqa benchmark for chest x-ray diagnosis.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Gemex: A large-scale, groundable, and explainable medical vqa benchmark for chest x-ray diagnosis

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:57.037988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:57.037988Z digest=sha256:cfe5f93098784b528411b564d92eed204efcaf34ef42055519552786a43f1ea3

Observation 57477684-b1ff-4ea1-9b22-503575081b9d · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InComputer Vision – ECCV 2024, pages 216–233, Cham, 2025.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Mmbench: Is your multi-modal model an all-around player? InComputer Vision – ECCV 2024, pages 216–233, Cham, 2025

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:57.207144Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:57.207144Z digest=sha256:a960de5de284f0c73ca4c338ecadb63eff931ee54fcaed645231cbd55e42a4a2

Observation 6fca4903-4c21-499f-8647-c663136fe150 · outbound

This paper cites Micro-bench: A microscopy benchmark for vision- language understanding.Advances in Neural Information Processing Systems, 37:30670–30685, 2024.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Micro-bench: A microscopy benchmark for vision- language understanding.Advances in Neural Information Processing Systems, 37:30670–30685, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:57.269886Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:57.269886Z digest=sha256:434f67fc3fccba6d6bc938049d2923e025ab3ebff8eb7d7e699403bfefd22f11

Observation aa0d21de-d5ac-47d7-81a3-1210962f63d0 · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:57.365011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:57.365011Z digest=sha256:3fa6405bd40b40d5744e9925a02f5df05c4af2c28e5b16fde4c5e9d7bdc3d188

Observation 2d8a451e-d030-4b3f-b729-2fc62808a74f · outbound

This paper cites FinMME: Benchmark dataset for financial multi-modal reasoning evaluation.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings FinMME: Benchmark dataset for financial multi-modal reasoning evaluation

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:57.438314Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:57.438314Z digest=sha256:c9bca108e74605d9cd792015506cdbe4999d43dd20dfad8f451ac7d9dd1115c2

Observation 5972c82d-a80f-4ba5-b43e-24de803a10c1 · outbound

This paper cites Archcad-400k: A large-scale cad drawings dataset and new baseline for panoptic symbol spotting.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Archcad-400k: A large-scale cad drawings dataset and new baseline for panoptic symbol spotting

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:57.536496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:57.536496Z digest=sha256:6839e2271854272ab66dec36986ceef4349d94d282f5e8db03c449baea4fe642

Observation 83a48c90-5a79-4a48-90d1-83dab1183777 · outbound

This paper cites Residential floor plan recognition and reconstruction.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Residential floor plan recognition and reconstruction

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:57.623965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:57.623965Z digest=sha256:035956fb94a6214023e62d93d05f2b1a73751483b2dbb4a8213ae9e9d44949c3

Observation 7919f1b7-f60d-4a64-82ef-351f87ca6c45 · outbound

This paper cites Phi-4-mini technical re- port: Compact yet powerful multimodal language models via mixture-of-loras, 2025.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Phi-4-mini technical re- port: Compact yet powerful multimodal language models via mixture-of-loras, 2025

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:57.755065Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:57.755065Z digest=sha256:ceb470dbd4c857f6cd80864586e378cd98cb0da335d25169806054aca24f7160

Observation 79fbece5-5677-4035-a50e-b88430875197 · outbound

This paper cites A-okvqa: A bench- mark for visual question answering using world knowledge.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings A-okvqa: A bench- mark for visual question answering using world knowledge

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:57.858309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:57.858309Z digest=sha256:f9b4201b34a7511fbf5c1cbee35c8afaca4d598fbb315607cdc6122ca2bc43e9

Observation 4bbcd799-cbc6-4213-8dc1-80a7213455c4 · outbound

This paper cites Qwen3 technical report, 2025.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Qwen3 technical report, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:57.977980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:57.977980Z digest=sha256:22e201171289291fd28025f42642e5cc1a4338abf342da08ab5ba2b17d6d1830

Observation 1fccdb11-93ab-4db5-bf61-051684911134 · outbound

This paper cites CEQuest: Benchmarking Large Language Models for Construction Estimation.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings CEQuest: Benchmarking Large Language Models for Construction Estimation

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:58.158507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:58.158507Z digest=sha256:a3244e9b9dd06379fdc8696c313a30d85a5a326271de86bb130449eb6147a9c7

Observation 8c0e8b3a-8f54-4d74-b175-1853727d286b · outbound

This paper cites Can ai master construction management (cm)? benchmarking state-of-the-art large language models on cm certification exams, 2025.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Can ai master construction management (cm)? benchmarking state-of-the-art large language models on cm certification exams, 2025

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:58.218481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:58.218481Z digest=sha256:f7da83935abd7d1d6a68e8fa9657e3dfdb5d5ff50d60c136d3bed5400afb4ffc

Observation b2476120-54fe-45b1-98cc-5147ccc25607 · outbound

This paper cites Mm-vet: evaluating large multimodal models for integrated capabilities.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Mm-vet: evaluating large multimodal models for integrated capabilities

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:58.301832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:58.301832Z digest=sha256:b9b5e6842a09065a839c17b5a0aed02e2a05ffcad69bbc051843327c0193f1d8

Observation 292867be-aa4a-48d4-805b-44d922936c36 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:58.449248Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:58.449248Z digest=sha256:1a850c13a5cb0a0f09ff0596b52d4bc0f2cb7a1bf102cc48cc23b1335488ff56

Observation a921f603-1234-44eb-a8fc-951c96c268bf · outbound

This paper cites Deep floor plan recognition using a multi-task network with room-boundary-guided attention.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Deep floor plan recognition using a multi-task network with room-boundary-guided attention

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:58.572597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:58.572597Z digest=sha256:a6a9651d33b4aa0a957689c4b14300937ef42583b55178fe345752e34ed20639

Observation 750768d7-62d8-42e5-b931-48a1f6376cad · outbound

This paper cites M3exam: A multilingual, multi- modal, multilevel benchmark for examining large language models.Advances in Neural Information Processing Systems, 36:5484–5505, 2023.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings M3exam: A multilingual, multi- modal, multilevel benchmark for examining large language models.Advances in Neural Information Processing Systems, 36:5484–5505, 2023

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:58.677698Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:58.677698Z digest=sha256:8781754b3f6333c03a63daf20e96fa37050f38bcf7c2e83ce8b3168690e3e07e

Observation 8e1c463e-dfef-459f-8e8a-edbbdb02e95c · outbound

This paper cites Pmc-vqa: Visual in- struction tuning for medical visual question answering, 2024.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Pmc-vqa: Visual in- struction tuning for medical visual question answering, 2024

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:58.750022Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:58.750022Z digest=sha256:4b8c2058db152be4eea864ce52897e87a4e75d7ea0d40d056085f959e927b84d

Observation 0c2032f8-8431-4294-bc79-06bf0e96825f · outbound

This paper cites Overview ofDrawingVQA.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Overview ofDrawingVQA

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:58.792929Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:58.792929Z digest=sha256:48aed25c76aff856fa7e7d30a8cfb3f6a777941ac037f57a21748d20719c51c7

Observation aa8cc483-10f4-4e1a-8c0a-2b0cb98711db · outbound

This paper cites Baselines.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Baselines

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:59.028585Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:59.028585Z digest=sha256:8b189b0c66258d941fb068613df981cf925f21c35dd898fcf2a0cb2badc078b0

Observation ddf8d110-8f27-44c4-9b57-c527d7537ccf · outbound

This paper cites expert bottleneck.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings expert bottleneck

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:59.108292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:59.108292Z digest=sha256:78ae20da68435fed3c55e5603931e4533d592e31705ca7faabcdd2bcbb81c370

Observation 7048dea6-ea1c-4021-a1dd-827730c0e9be · outbound

This paper cites OCR, Visual Perception.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings OCR, Visual Perception

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:59.203113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:59.203113Z digest=sha256:3225f6aaf05f2e42ca8dca4266eecddade18975adf327b808aa7dad984ec3b27

Observation 198b13df-189f-419d-ba60-7019fc510cf0 · outbound

This paper cites Visual Perception, Reasoning (Alignment, Visual), OCR.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Visual Perception, Reasoning (Alignment, Visual), OCR

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:59.448156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:59.448156Z digest=sha256:091622b2c76509c14f10b5474803cca10f678f9e09edad74064dcfe1ea244a3a

Observation 342a8184-2faa-44f2-b99a-6d04e987ba9b · outbound

This paper cites Reasoning (Alignment, Spatial), OCR.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Reasoning (Alignment, Spatial), OCR

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:59.800016Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:59.800016Z digest=sha256:bd939c0817872a8929a6bc9d0f8180903d281796035c2bf87f31b90e6154aabf

Observation 1a1b7d2e-5c1a-4d89-9b44-c096889a32bb · outbound

This paper cites Knowledge, OCR, Reasoning (Vi- sual, Spatial).

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Knowledge, OCR, Reasoning (Vi- sual, Spatial)

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-01T23:28:59.940466Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:28:59.940466Z digest=sha256:62d23c986d1ed0f21724e7503c893e2cedb8e3fff4d621e976602747a50e8c65

Observation 0c82cade-8d65-4b35-a682-d9c6aa47ce9a · outbound

This paper cites Reasoning (Alignment), Visual Per- ception, OCR.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Reasoning (Alignment), Visual Per- ception, OCR

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-01T23:29:00.140935Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:29:00.140935Z digest=sha256:06a0ac2af47d543cf54988beed2690b15b052f695ca59c3b31f00a74dc12d918

Observation 031f92f9-01a6-4c72-985f-e074b851d2d6 · outbound

This paper cites The answer is (X)\.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings The answer is (X)\

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-01T23:29:00.254896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:29:00.254896Z digest=sha256:50e7bfd638aa64affe9ea01c4fe5b7314205681fadc268ebb488d27ecee349bb

Observation e78abfa0-db37-455c-b52c-57f6e77740c0 · outbound

This paper cites Can AI Master Construction Management (CM)? Benchmarking State-of-the-Art Large Language Models on CM Certification Exams.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Can AI Master Construction Management (CM)? Benchmarking State-of-the-Art Large Language Models on CM Certification Exams

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-01T23:29:00.488752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:29:00.488752Z digest=sha256:9d371a2f99bfc7d654f55cf93b88b584d12cf60223224131456f9043d07fbc28

Observation 2bee8ed8-15a9-4efa-b8bf-44ce3ac0f99f · outbound

This paper cites DesignQA: A Multimodal Benchmark for Evaluating Large Language Models' Understanding of Engineering Documentation.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings DesignQA: A Multimodal Benchmark for Evaluating Large Language Models' Understanding of Engineering Documentation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-01T23:29:00.582559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:29:00.582559Z digest=sha256:2f326ca6dd2b1a6c8afc0c2f97ccb96a25beb2f8f1b4effecc01f44bcb29b01e

Observation dfaf8944-d82a-4aba-bcd1-997dde3b7c7d · outbound

This paper cites See What You Are Told: Visual Attention Sink in Large Multimodal Models.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-01T23:29:00.730157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:29:00.730157Z digest=sha256:82c0c6b38bbf836dba50e473453498b49a998c43f852e9602018f4b24ac0506a

Observation cc5658fa-3d52-452f-a5c0-b2f5dd5cb64e · outbound

This paper cites Llama-3.2-11b-vision-instruct.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Llama-3.2-11b-vision-instruct

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-01T23:29:00.898117Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:29:00.898117Z digest=sha256:f7c63a6a0e79d776c356b6fcc245e1ce97c25536675fb1615100a22848e35a53

Observation 1d5ddab6-ecfd-4492-8a70-49f81cfce455 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Improved Baselines with Visual Instruction Tuning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-01T23:29:01.030233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:29:01.030233Z digest=sha256:88a77e3e160c56cdbfa386334d924b89b3fb92642f3470b2e393dddd13b80992

Observation 025aeada-05ea-46f4-9e7f-59639cfb5c49 · outbound

This paper cites The llama 4 herd: The beginning of a new era of natively multimodal ai innovation, April 2025.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings The llama 4 herd: The beginning of a new era of natively multimodal ai innovation, April 2025

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-01T23:29:01.146198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:29:01.146198Z digest=sha256:b2d3e113e22dee92878ad3345ed32d7a0a91df4c44dd2b5bbfdf382a6fa84cc5

Observation 4d9e78c9-09e4-4148-b100-fc2aaec7711f · outbound

This paper cites Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-01T23:29:01.307092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:29:01.307092Z digest=sha256:84a11ca7b248d9375560e2de9be84182fdbd3e637b347c6c1ca40e8a693c51b3

Observation e695c2ef-bb05-4d72-99ac-84b65e4ad564 · outbound

This paper cites Burgess et al.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Burgess et al

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-01T23:29:01.445864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:29:01.445864Z digest=sha256:31ed14b413233954b9ce5dd275cc6e75496bf56a819f494edc6ba306dd027571

Observation 66942b01-7847-4f0c-8993-73a45deee2b6 · outbound

This paper cites Yue et al.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Yue et al

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-01T23:29:01.594857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:29:01.594857Z digest=sha256:e097d7c17d07d6501df42f64eaef130620842717614f4cbf9e3dfd14c24d0f83

Observation dbb8dbdd-2c32-43a8-a908-ee7e87b1748c · outbound

This paper cites LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-01T23:29:01.719021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:29:01.719021Z digest=sha256:0cb902466fd2b1cc53c4115ede55db37253c3e13baa72b175969b782cbfd8198

Observation 94998587-451f-4e6e-8757-fe17647b6b45 · outbound

This paper cites Qwen3 Technical Report.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Qwen3 Technical Report

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-01T23:29:01.869242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:29:01.869242Z digest=sha256:2fdd948af29821047659c58a45efeae6674c185b8c1c693b06fee07d7a3ab28c

Observation 39568c5b-0251-4f88-8aa2-18951ac7297e · outbound

This paper cites InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-01T23:29:01.995396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:29:01.995396Z digest=sha256:2f879c989557a017d206cf7f509c9d858ab2040d28cba6a720a72af569c1a055

Observation 452419ac-2c48-4ef9-99bf-a11a2c7a10f4 · outbound

This paper cites an unresolved cited work.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-01T23:29:02.149807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:29:02.149807Z digest=sha256:fde273a8c48f5af287636f95f8a83013f9848fe730a39b17830a5d6024bddf4c

Observation 4c31efcf-b031-4334-93ee-712b9eb599cb · outbound

This paper cites Large Language Models Are Not Robust Multiple Choice Selectors.

DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings Large Language Models Are Not Robust Multiple Choice Selectors

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-01T23:29:02.328455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:29:02.328455Z digest=sha256:38f269a69c34bac7d9b32061f2b65e889ccd25036958f14e65ac9fc94cf879ad

Pith citing papers

Observation 3ef6c10b-1baa-424e-825a-831dbc143560 · inbound

Evidence-Grounded Constraint Checking in Construction Documents cites this paper.

Evidence-Grounded Constraint Checking in Construction Documents DrawingVQA: A Real-World Benchmark for Multi-Depth Visual-Textual Reasoning on Construction Drawings

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T14:28:20.930744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T14:28:20.930744Z digest=sha256:1275464ed03cafcfad3e4f899257b1859cea92abd7e47bb6741ffc70cac3815e