Pith. sign in

Paper Citation Record · LEDGER

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs

As of 18 August 2026, this Paper Citation Record lists 100 of 106 outbound references and 2 inbound Pith citation observations for arXiv:2507.04664.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04664 v1

Coverage vector

measured 100 of 106 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:46:12.494249Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:19:20.192163Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T15:19:20.262812Z

Reference resolution

100 of 106 outbound references displayed

  • verified exact1
  • verified fuzzy47
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d8210e97-721d-4377-98c5-7fbcf411d4a1 · outbound

This paper cites write newline.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:10.920396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:10.920396Z digest=sha256:6a6d9e0ed3f5d240cf6c668f7b4476c555e1f0642e74ec80f32663d8c84cc507

Observation 284f28ca-349f-44b6-a49a-aee047b86458 · outbound

This paper cites Efficient interactive annotation of segmentation datasets with polygon-rnn++.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Efficient interactive annotation of segmentation datasets with polygon-rnn++

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:11.024550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:11.024550Z digest=sha256:8110b8d04f271f4d6724426637f373825f84685627db8a2e88a4cd61dd80432d

Observation 4ed2ba49-0e13-4285-bc0b-562f5b2176d5 · outbound

This paper cites Qwen Technical Report.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:11.125140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:11.125140Z digest=sha256:8c782b6c56c77686cfdeed571ee7b65777e765dcd509ea9454c45f2ab1bb8011

Observation 63f16326-a8d3-47a3-9f67-0cafaf669179 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:11.275037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:11.275037Z digest=sha256:d527f4f622d283a4026576e7696ecb98b87720b5446c1e9fd67370c4f9f83a3b

Observation 55a2e43a-de79-4c6a-94f2-51ed9ffe9dff · outbound

This paper cites Multi-task learning for segmentation of building footprints with deep neural networks.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Multi-task learning for segmentation of building footprints with deep neural networks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:11.444690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:11.444690Z digest=sha256:e211a62ec6c427575f540a223111438d56a9092e3e6fc61c6154329df73c865b

Observation efa7c1c8-6e2f-4fcd-982f-bd06561d09ba · outbound

This paper cites J., 2019.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs J., 2019

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:11.605083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:11.605083Z digest=sha256:60af7ac3d2536c7726406b04f1a1a095ce02d766b80494f1438b94e674d4d472

Observation 79d0203d-3bb3-4a7b-beee-c4685cd45c41 · outbound

This paper cites InternLM2 Technical Report.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs InternLM2 Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:11.735318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:11.735318Z digest=sha256:316b008c78d4eb1c52cac777909c39021770a8d0642c43a204cf955b71e0f14f

Observation d361c8c1-8a8f-44a8-a1ab-8d1ea92df04e · outbound

This paper cites Annotating object instances with a polygon-rnn.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Annotating object instances with a polygon-rnn

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:11.861213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:11.861213Z digest=sha256:469152eb7118d885b4462ae7af1467bb1bc81d0834e7ee4574a0a2da46f95cae

Observation d5c7e7e8-8b92-4f60-8cc8-4beaa3a6d456 · outbound

This paper cites ASF-Net: Adaptive Screening Feature Network for Building Footprint Extraction From Remote-Sensing Images.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs ASF-Net: Adaptive Screening Feature Network for Building Footprint Extraction From Remote-Sensing Images

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.035880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.035880Z digest=sha256:e979387d9386626a7538cdcda445b85255a39889cff7476206c4f5652391c66e

Observation e9ed3f77-ac4b-48de-844d-899ff0a7eb35 · outbound

This paper cites RSPrompter: Learning to prompt for remote sensing instance segmentation based on visual foundation model.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs RSPrompter: Learning to prompt for remote sensing instance segmentation based on visual foundation model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.044430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.044430Z digest=sha256:ffc078570d3414bddef8176388386b58be8c20f7de40a73f2328166fea6b7731

Observation 676c25f8-9990-409e-9867-bed03307dcbc · outbound

This paper cites L., Liu, X., 2020.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs L., Liu, X., 2020

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.108746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.108746Z digest=sha256:48088aff8ac6bde9af4dcde09a9ec21743df86a91506b3799f4b0ecc153cea29

Observation 0f688f2a-513d-4336-8c6b-2217dd18266d · outbound

This paper cites CGSANet: A Contour-Guided and Local Structure-Aware Encoder--Decoder Network for Accurate Building Extraction From Very High-Resolution Remote Sensing Imagery.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs CGSANet: A Contour-Guided and Local Structure-Aware Encoder--Decoder Network for Accurate Building Extraction From Very High-Resolution Remote Sensing Imagery

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.113348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.113348Z digest=sha256:4c0907f4b2ccf981f42d4db4ab0babec302ddc91cc558f9410f61d4574611a10

Observation 2c0fb206-cf12-4d86-85e9-95a803d45295 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.118150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.118150Z digest=sha256:f7e246c2dc3cd4de68189a6734aa359bdf27ba93b5e68f62c65ed059555e2d5a

Observation 098777dd-1745-4fc0-8d9d-fa06d2e57f0b · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.123184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.123184Z digest=sha256:72680fba2a7bd5ee662d6d0e36fae798fe52edc7bad77ac181d51f465b024227

Observation 37d15ac5-d595-4ae6-85ce-e59ac22db8af · outbound

This paper cites et al., 2024d.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs et al., 2024d

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.128509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.128509Z digest=sha256:bec6a07588acf80181bdf27d4634afcdac7c4731b467c0262f9bd17ec775503e

Observation 4a946523-bd2f-4276-853f-43490e74626e · outbound

This paper cites G., Kirillov, A., Girdhar, R., 2022.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs G., Kirillov, A., Girdhar, R., 2022

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.132674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.132674Z digest=sha256:52f12fa48c53b0539ee4128194e554ef1bed440223461cfb41e10a4e50144dd2

Observation ee994916-5335-4e80-af27-bb9bd9bc5ecf · outbound

This paper cites E., Stoica, I., Xing, E.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs E., Stoica, I., Xing, E

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.140442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.140442Z digest=sha256:7b817609cf5538afd7469a26fca5786ea4e4ea92ea13b55ecda2007be130bf1f

Observation 03a7438f-6450-4b4f-b99e-11d5ef313a01 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.145107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.145107Z digest=sha256:25e37c724221430288368b5b63ec433397ccd0268399d87b4834a4828eeacc4c

Observation 72388a81-c250-4eec-a565-219e61eab721 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.149426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.149426Z digest=sha256:9ccfd002ff1a25d45040ab82d8b0042df90f7accca83609c375b4bdee965052c

Observation f9a3c9a8-4cf6-4041-a84b-2654ce84d911 · outbound

This paper cites The Llama 3 Herd of Models.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs The Llama 3 Herd of Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.153773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.153773Z digest=sha256:fe4ac66f882b518173dab42ed97d8cee2176b69d75455a65b0013263d13cc29e

Observation 402e5d75-ed0f-4680-a137-99357aa52267 · outbound

This paper cites GeoLLaVA: Efficient Fine-Tuned Vision-Language Models for Temporal Change Detection in Remote Sensing.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs GeoLLaVA: Efficient Fine-Tuned Vision-Language Models for Temporal Change Detection in Remote Sensing

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.158072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.158072Z digest=sha256:c72966ea49fb55509eee6d91b5d23ef0ad385cf2a80d84fa8c63014d66dd187f

Observation 2644427a-76f4-458a-ae86-b42183c9ab8c · outbound

This paper cites Instances as queries.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Instances as queries

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.162712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.162712Z digest=sha256:51770f16f8c8b1b8079e40cf089d4625ee7442fae543d27a880fedeae48a0933

Observation 3417c8d8-f6f3-49f5-80c0-f6220469bb90 · outbound

This paper cites GPT-3: Its nature, scope, limits, and consequences.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs GPT-3: Its nature, scope, limits, and consequences

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.166733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.166733Z digest=sha256:e65ace37fa0004e9f69f3eaa2b4c4d7f2980e7483981cf89634477ffb422069b

Observation 22fb0106-5f5a-4949-b4c3-7d169df1fd6e · outbound

This paper cites Polygonal building extraction by frame field learning.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Polygonal building extraction by frame field learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.170835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.170835Z digest=sha256:0fad17ee48cf46ca49d176c55e6d4531b10f2104052855e505b970072740819f

Observation fec73877-e797-40a4-a79b-5ad3e5d28eec · outbound

This paper cites Remote Sensing ChatGPT: Solving Remote Sensing Tasks with ChatGPT and Visual Models.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Remote Sensing ChatGPT: Solving Remote Sensing Tasks with ChatGPT and Visual Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.175344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.175344Z digest=sha256:94fb2457f2984c854758ec02b8a1852fcf8dca647e251cd1dc5fd083ecd1e084

Observation f238ad8e-3b6b-43ad-8503-5e493a6da9c9 · outbound

This paper cites HigherNet-DST: Higher resolution network with dynamic scale training for rooftop delineation.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs HigherNet-DST: Higher resolution network with dynamic scale training for rooftop delineation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.179579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.179579Z digest=sha256:ec16c7fb84650994bea6352f5e01786fc896318fddbf57c6458d132bc5d18615

Observation c8a427c6-96d6-4c74-a134-891fa92077fc · outbound

This paper cites Mask r-cnn.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Mask r-cnn

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.183884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.183884Z digest=sha256:17635fac6860696aa23a32311c7dc60ed96a2922ccf0d38862941e63442ccb9b

Observation d390bb96-5275-4feb-ae74-c349d1d523d6 · outbound

This paper cites Rsgpt: A remote sensing vision language model and benchmark.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Rsgpt: A remote sensing vision language model and benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.187715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.187715Z digest=sha256:738e94f796037ee25f040db95e1f97f3680f6e82f8b0e17ca68a3c8221bbb9b0

Observation 7a2c0461-438f-4b96-87cb-4627b4b0c473 · outbound

This paper cites Sequentially delineation of rooftops with holes from VHR aerial images using a convolutional recurrent neural network.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Sequentially delineation of rooftops with holes from VHR aerial images using a convolutional recurrent neural network

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.191988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.191988Z digest=sha256:54f9254752e2ea1ed670c5c152fb51d6e2f6492ba696e5bc098200bc1a6a8315

Observation 4336df52-0d5d-43e9-9c6c-c7e0fb34713f · outbound

This paper cites OEC-RNN: Object-oriented delineation of rooftops with edges and corners using the recurrent neural network from the aerial images.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs OEC-RNN: Object-oriented delineation of rooftops with edges and corners using the recurrent neural network from the aerial images

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.196441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.196441Z digest=sha256:87371439e61b5b6d0bf18ee4b852a35850d049b05192b0a75d30edcedf54492c

Observation 09dde30f-882b-419e-83a3-f1e85ca77997 · outbound

This paper cites TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.200812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.200812Z digest=sha256:d11bbfa08c6d6831012d6f435386b5e5962fe007264411b109c1acf3c869957c

Observation 7d30da99-0abc-453d-b115-07c3af35c33e · outbound

This paper cites Fully convolutional networks for multisource building extraction from an open aerial and satellite imagery data set.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Fully convolutional networks for multisource building extraction from an open aerial and satellite imagery data set

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.205089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.205089Z digest=sha256:3a3c8ff4237718cba1d2b819fdeeaae382f4d633fa85c4aa8134d10923fe7246

Observation dcd086c9-29ac-424a-8f5d-3342a2617651 · outbound

This paper cites A scale robust convolutional neural network for automatic building extraction from aerial and satellite imagery.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs A scale robust convolutional neural network for automatic building extraction from aerial and satellite imagery

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.209483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.209483Z digest=sha256:edef1d7d2bdbeaeb58bfd38a6baace965c81b7d92086b7fa5a3fe7ef22db021f

Observation 889edfd5-384c-43d1-83c6-c19aae744b8c · outbound

This paper cites Segment Anything.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Segment Anything

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.214127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.214127Z digest=sha256:537d8d4ef053e991395c4ab294e7835859fa0178141aff616259fb8fb2a90781

Observation 54e8859c-9ea6-43ac-868d-0d2187b98f52 · outbound

This paper cites C., Lo, W.-Y.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs C., Lo, W.-Y

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.218573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.218573Z digest=sha256:1b1b673efcc3936145da3a5c06bbfe2f579e07ef1bb6a513070d5036c325ead7

Observation 62b9df7f-f2c1-4fa1-99c1-dd732c0569ce · outbound

This paper cites E., 2017.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs E., 2017

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.222970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.222970Z digest=sha256:c129e44594d008ed0910ee66b793c793139275082511f987280b6a3776e58f76

Observation a82c35c5-a1ed-4198-92a3-a2b2040f4336 · outbound

This paper cites S., Naseer, M., Das, A., Khan, S., Khan, F.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs S., Naseer, M., Das, A., Khan, S., Khan, F

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.227214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.227214Z digest=sha256:6bedb23ac909dc4f16da0a408d2b5b711a85c90c1e9d5992029e1c27bc38d66b

Observation bf74aaef-e9c3-4173-8e2e-12842ec9cdaa · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Lisa: Reasoning segmentation via large language model

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.768609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.231466Z digest=sha256:345ae75cc520c752811d4d0ee1b8de25a389ee18f5ae0bb9db7321da68e712b4

Observation 83f5ad00-4f16-4c13-ae3c-d5093f990ca2 · outbound

This paper cites VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.235897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.235897Z digest=sha256:405397531ad5b5ff4eca16c6858db54c5aa2c230f8754669580e4f9325198679

Observation c7f65599-2566-4ca8-946f-9be25206b819 · outbound

This paper cites Topological Map Extraction from Overhead Images.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Topological Map Extraction from Overhead Images

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T19:46:12.810473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.240475Z digest=sha256:60b410170dbeaca805c8cecdf36192ff922b8218bc2870c7f931c13c568f35ad

Observation 3ce7b139-c1b1-4957-8785-0d8f624a62bc · outbound

This paper cites D., Lucchi, A., 2019.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs D., Lucchi, A., 2019

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.750718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.247118Z digest=sha256:5598d6b5c1e16f142063ebc36184db39ba3931e4f08516ae8548e78242e6e676

Observation f2c10974-566d-4cc7-91ad-596ab7bb7a93 · outbound

This paper cites Polytransform: Deep polygon transformer for instance segmentation.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Polytransform: Deep polygon transformer for instance segmentation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.735684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.251308Z digest=sha256:bfd5aa207442bae513d783aea130a9127e2df499f064669adfa2e3dba93f32df

Observation d37fc12e-ba8d-4aed-8492-ab0056e2f0a9 · outbound

This paper cites L., 2014.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs L., 2014

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.255657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.255657Z digest=sha256:627745596e100997fe208f399b0169a7f700c6b0b22fbd3dec476d1368f7c081

Observation 827c00df-b728-4e03-9d0f-d55d15100aa4 · outbound

This paper cites Draw-and-understand: Leveraging visual prompts to enable mllms to comprehend what you want.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Draw-and-understand: Leveraging visual prompts to enable mllms to comprehend what you want

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.709894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.259708Z digest=sha256:d662ce2612d5b79ef3221832f1a1540482243e63a5436ccdfab8d0c595fe9ff8

Observation d0bbce28-e468-439e-99ca-15afbfcb788e · outbound

This paper cites Change-agent: Towards interactive comprehensive remote sensing change interpretation and analysis.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Change-agent: Towards interactive comprehensive remote sensing change interpretation and analysis

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.696936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.263754Z digest=sha256:96594bec0c867317c804b18991e6f4d051d79a857308481056e1058768378123

Observation d3ad047d-f059-4b44-aadd-03e46058c842 · outbound

This paper cites J., 2023a.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs J., 2023a

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.683383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.268556Z digest=sha256:6f2842c4fb3fa6be1cc883abb944906992aa9b73c2db304cac3a31f0426d425a

Observation c9f2d888-2725-4ade-8a56-de7b935ab042 · outbound

This paper cites J., 2023b.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs J., 2023b

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.669499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.272742Z digest=sha256:2edc23249f1180681dfce19a3ba93274ad36aefcf7f490e3131326573ebfb1ee

Observation 78abfdac-54bd-486c-a69d-ede748ad8ac3 · outbound

This paper cites Path aggregation network for instance segmentation.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Path aggregation network for instance segmentation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.655042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.277029Z digest=sha256:b67423366035d125419c8fc2634af73d245c1d21cf6f0e50cbaa25ca86d36d7e

Observation e24ef6bb-7b5e-4eb5-b2b1-cf06837ae3e4 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Swin transformer: Hierarchical vision transformer using shifted windows

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.641484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.281467Z digest=sha256:0e109eeb8242340106d30e67b9983428f0f6d99b64e3c6fa121e30790859113d

Observation ce1e00f3-4929-46f3-b4f9-6d7406b15fc6 · outbound

This paper cites Building Outline Delineation From VHR Remote Sensing Images Using the Convolutional Recurrent Neural Network Embedded With Line Segment Information.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Building Outline Delineation From VHR Remote Sensing Images Using the Convolutional Recurrent Neural Network Embedded With Line Segment Information

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.627700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.286109Z digest=sha256:9786967f0109f8fba8728253bee9dac68d12ead382bb866cfaabce44d93c27bc

Observation b13b024e-9a19-4236-ac5f-2454e73df9d4 · outbound

This paper cites SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.290250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.290250Z digest=sha256:f6fc68c3d0494a78c550d2381b1365992522d1941780b848a22816e800b586d5

Observation 1b14cb8b-87f3-46a5-958f-6c3c4d137129 · outbound

This paper cites Cross-spatiotemporal land-cover classification from VHR remote sensing images with deep learning based domain adaptation.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Cross-spatiotemporal land-cover classification from VHR remote sensing images with deep learning based domain adaptation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.613358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.295054Z digest=sha256:0e32e634cecd30425105334069e7b3dcfad71c8872227a1b902744a646662047

Observation 686281e2-c5b5-4920-8bb8-32a113dcae3c · outbound

This paper cites SAM-RSIS: Progressively adapting SAM with box prompting to remote sensing image instance segmentation.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs SAM-RSIS: Progressively adapting SAM with box prompting to remote sensing image instance segmentation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.597708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.299089Z digest=sha256:9f3ba81d73b7b7b99dcca591fd061d83eaca971d047964e8f99956c7361a68a2

Observation 29af8f0c-1af2-430f-bd2f-6a6bca28fb56 · outbound

This paper cites P., 2019 (accessed November 10, 2019).

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs P., 2019 (accessed November 10, 2019)

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.583455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.303073Z digest=sha256:dfcc4d3e1171f40355f3f6a901e922614cd736c447c29bda3f4b0c2b5bc648ae

Observation 8df539d9-a371-4125-b177-b49eacccc809 · outbound

This paper cites LHRS-Bot: Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language Model.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs LHRS-Bot: Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language Model

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.307217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.307217Z digest=sha256:d44f5264032c15128406ec688bd5d089908f90e60b3bbb11b90ab4017f4ef279

Observation 128196e5-76c8-4d90-b068-f37701d3c630 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs DINOv2: Learning Robust Visual Features without Supervision

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.311807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.311807Z digest=sha256:33885f09b09326e93e66837c00a8ff9e572d7edecc394680efc177ceb498a537

Observation 3f4aace0-37a0-4b07-b11b-ba9fbb559ba7 · outbound

This paper cites VHM: Versatile and Honest Vision Language Model for Remote Sensing Image Analysis.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs VHM: Versatile and Honest Vision Language Model for Remote Sensing Image Analysis

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.316595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.316595Z digest=sha256:5a0c23a06b0fb49f818ac8ee752db76a2cbcde63b94c36fcce2678832b6bd045

Observation f887536e-9793-4d10-97a4-26108a357f69 · outbound

This paper cites Deep snake for real-time instance segmentation.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Deep snake for real-time instance segmentation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.567120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.320655Z digest=sha256:08a18aabccfedcfae6506a494316bae11bae90e394f37f20c7dcfb70c9b5cc2e

Observation 4c97ddd7-5460-4792-b3ba-aa1775f656b2 · outbound

This paper cites W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.552334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.324853Z digest=sha256:80619069b28fdd73f27b63fff42efb241878075083c7ecaecdd6ca2edca7d612

Observation c54176be-9c7f-44e2-8812-74181bd3f3f5 · outbound

This paper cites D., Ermon, S., Finn, C., 2024.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs D., Ermon, S., Finn, C., 2024

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.536754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.328747Z digest=sha256:071bc7d35d74d3942a29c0eb88c655205f5667668c54f7fb3e23a38b59db7820

Observation 0cd113dd-ee83-4552-b263-db1eb4fa46d9 · outbound

This paper cites Am-radio: Agglomerative vision foundation model reduce all domains into one.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Am-radio: Agglomerative vision foundation model reduce all domains into one

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.521226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.332948Z digest=sha256:0a01245324d544ebba950237643587ff8becd89155bd55d9aba473301f32509a

Observation aba2b761-0c8c-4030-8fed-7672c527f1fd · outbound

This paper cites M., Xing, E., Yang, M.-H., Khan, F.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs M., Xing, E., Yang, M.-H., Khan, F

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.504453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.336760Z digest=sha256:1a768a5a6bbdbcc3ea08218088ea0ecc21538ec50a21ece0765eb58e1a4eb014

Observation 9aad2ccb-9034-46e9-a8aa-bd39225ce36f · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Pixellm: Pixel reasoning with large multimodal model

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.490327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.340577Z digest=sha256:392253b82d88fe80697a60895d8d828610736440b4ee734e6f2460ebe4984e45

Observation 540c0d6d-2bb8-433d-807a-fc007055224f · outbound

This paper cites Geollm-engine: A realistic environment for building geospatial copilots.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Geollm-engine: A realistic environment for building geospatial copilots

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.473850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.345906Z digest=sha256:c12880a4dd3f9ff97c3bd7f6d1404791d988d6c603eaa666a88aca132b4a67d1

Observation cc1b2af8-ef4a-471e-9c88-355c64381b81 · outbound

This paper cites Internlm: A multilingual language model with progressively enhanced capabilities.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Internlm: A multilingual language model with progressively enhanced capabilities

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.459916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.350574Z digest=sha256:db9d573cb140ce779236298ff1bb5851377d6327606417d452137096ea7c1691

Observation df3fd4f0-9393-4133-a26b-81021c816039 · outbound

This paper cites Fcos: Fully convolutional one-stage object detection.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Fcos: Fully convolutional one-stage object detection

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.446080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.354993Z digest=sha256:cd152eb5dd8cd57674549d88e66f5e145ea85d9c8cfba7aafaeff4a029e9dbfd

Observation c75b4a2d-c816-4c9d-afe0-9087289aa521 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.359024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.359024Z digest=sha256:fce4a028149875271601092d3f2a0cf95046059217639bb104bcab294917ecc2

Observation 5fb2122f-b754-48a1-937a-73536643f55e · outbound

This paper cites From image transfer to object transfer: Cross-domain instance segmentation based on center point feature alignment.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs From image transfer to object transfer: Cross-domain instance segmentation based on center point feature alignment

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.432782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.363048Z digest=sha256:b7f6128748ade342b442382c7ddfff0f50b1a37eb7b18a692c619c9c00433218

Observation 33524a30-b01d-41a9-a075-6ef26917ce84 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.367046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.367046Z digest=sha256:7fa427ba44be27b3e79042e6cb2431adf7d5299243c1d890e156b1c7440596cf

Observation e4eb0793-f138-4d14-bd7d-67ba16e7bf46 · outbound

This paper cites et al., 2024b.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs et al., 2024b

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.418753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.371187Z digest=sha256:ab4d89b787cebe335320b5233b4ea43f6c92f5dc55d612bc8355f806ddb0d20a

Observation 54bb2a9e-a2c6-4242-9dfc-8fcac75dd936 · outbound

This paper cites Solo: Segmenting objects by locations.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Solo: Segmenting objects by locations

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.405054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.375385Z digest=sha256:35977197daeb69150648616c7720812d9f36868f6e09f6d316fd824119c06b70

Observation 120e7511-4cf5-4df1-9071-b019c8b8f415 · outbound

This paper cites Skyscript: A large and semantically diverse vision-language dataset for remote sensing.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Skyscript: A large and semantically diverse vision-language dataset for remote sensing

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.391338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.379117Z digest=sha256:2f9fc2a8e3c73184ea4c4f8466accce2b7b88c30a0e709bd2259f05b54dd0083

Observation 20acfc7f-f109-4a06-ac92-5b3d82a53cde · outbound

This paper cites Graph convolutional networks for the automated production of building vector maps from aerial images.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Graph convolutional networks for the automated production of building vector maps from aerial images

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.376449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.383357Z digest=sha256:581d23fe84baf950b56bdba377ff6269f0a89e2a6622759cc8b8cc7e3f844a8b

Observation bc4e1410-b3d9-43c9-9dc1-bae1c1a1f4d7 · outbound

This paper cites Toward automatic building footprint delineation from aerial images using CNN and regularization.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Toward automatic building footprint delineation from aerial images using CNN and regularization

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.361186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.387107Z digest=sha256:d26d289d2051642f08fd37804bdb1260c401c99aafa64d18b69858188bbbd276

Observation b74a44fc-cae6-4e8c-811c-15f90523039e · outbound

This paper cites A Concentric Loop Convolutional Neural Network for Manual Delineation-Level Building Boundary Segmentation From Remote-Sensing Images.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs A Concentric Loop Convolutional Neural Network for Manual Delineation-Level Building Boundary Segmentation From Remote-Sensing Images

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.346545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.391234Z digest=sha256:d9564a57f7f8959f31162747d09b8036d8c216cb7e3a4e4772e3b98d42fed2a5

Observation 681bfab7-2255-4e1e-8b9c-52960724f664 · outbound

This paper cites BuildMapper: A fully learnable framework for vectorized building contour extraction.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs BuildMapper: A fully learnable framework for vectorized building contour extraction

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.332317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.395043Z digest=sha256:edfe06b66987d9f41700ce29e1035319676896922348f18e7bc935c0fd245c23

Observation 33d5c8b8-0ad6-43f7-b51f-561cfea87f99 · outbound

This paper cites From lines to Polygons: Polygonal building contour extraction from High-Resolution remote sensing imagery.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs From lines to Polygons: Polygonal building contour extraction from High-Resolution remote sensing imagery

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.319037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.399002Z digest=sha256:45626eb05df7133e7db7f0f7261a5d02c7e3a3f36a4757ed66caf6c6207c519a

Observation d6e94cea-aa35-4323-b476-edf1d9b58ca1 · outbound

This paper cites A., 2022.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs A., 2022

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.305215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.403050Z digest=sha256:c5132e379ef25e508da15fcf8e3554e4145ad58f1603aa5049a702116aabbe03

Observation 9e79cef5-9028-487b-a278-f3bfd73a6cdd · outbound

This paper cites Vectorizing historical maps with topological consistency: A hybrid approach using transformers and contour-based instance segmentation.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Vectorizing historical maps with topological consistency: A hybrid approach using transformers and contour-based instance segmentation

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.291113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.406984Z digest=sha256:f479743b1775b1a297296a0edf3071b22c335fa06b0073a9a6c2df395bcf09e4

Observation 05abb26e-ca00-4925-a7c1-12782aa2dbdc · outbound

This paper cites Video instance segmentation is all you need for linking geographic entities from historical maps.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Video instance segmentation is all you need for linking geographic entities from historical maps

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.275145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.410892Z digest=sha256:dffa9d55a29eb457fd689bdafd0c52f07f64c0256eaf329a35a291accd8d8db5

Observation d787037c-3fbf-4ccc-b604-62d610df6fde · outbound

This paper cites Polarmask: Single shot instance segmentation with polar representation.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Polarmask: Single shot instance segmentation with polar representation

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.261366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.415058Z digest=sha256:c2bb4e13dd83611b4c5eaf05a4cf7bbaf6142aa0e2d221b8b9bd1291ae2fbbd0

Observation 939ef4fa-5a08-4f74-aba5-74d9dfcd08bf · outbound

This paper cites HiSup: Accurate polygonal mapping of buildings in satellite imagery with hierarchical supervision.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs HiSup: Accurate polygonal mapping of buildings in satellite imagery with hierarchical supervision

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.246162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.419407Z digest=sha256:b40382e603e5916a0bda1301369d29276b17700130d86c8ed95686943aa00c7d

Observation 61e0fd6c-7086-41b8-a439-e08f2b35bd76 · outbound

This paper cites Qwen3 Technical Report.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Qwen3 Technical Report

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.423489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.423489Z digest=sha256:0a9065394a8e4baceb89fe5b24ad648d98b80e45b04050aac3376cb879f6c1e8

Observation 45a3c82f-e62d-44f6-be72-1e0b36778ccb · outbound

This paper cites Qwen2 Technical Report.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Qwen2 Technical Report

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.427552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.427552Z digest=sha256:04769d3e47c99559685c3669787d5024748be46be893f46acdffac9b52c5de59

Observation 02d2e529-04d6-463c-b965-a3a9370fd2c3 · outbound

This paper cites A review of recurrent neural networks: LSTM cells and network architectures.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs A review of recurrent neural networks: LSTM cells and network architectures

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.231469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.431291Z digest=sha256:400248bc796f417cf407b1aa6b0159a8afc1c83960396d42df9cfeb9814c7864

Observation b59a960f-b2db-46c5-b692-a14e5a8fb91e · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.435161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.435161Z digest=sha256:82ac62f1fd8e7a18b2210baf3cf64f60b1cb1e2f71005c2e2a6eefca8d5affd1

Observation 90967185-4b8f-4448-80df-44aa4b22d059 · outbound

This paper cites Learning building extraction in aerial scenes with convolutional networks.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Learning building extraction in aerial scenes with convolutional networks

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.216943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.439134Z digest=sha256:8a4529efee4546213038c849abf28216b53193fa3307b187cb2410ba8e2fe815

Observation 184b47f0-cbea-4043-99d8-a666a269c6b0 · outbound

This paper cites Osprey: Pixel understanding with visual instruction tuning.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Osprey: Pixel understanding with visual instruction tuning

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.202555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.443193Z digest=sha256:03e78fca0b9d2d560191339f0f290b7673d6e0b292e5d0063863b440591d3267

Observation 22a3060a-2a1b-45da-a5e6-2f7ab87ee0dc · outbound

This paper cites SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.447693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.447693Z digest=sha256:70ebfdeea63251fae17b2115433aca7f88131cd85cd0eb1f8b9e52cd24dcc0bc

Observation 73189925-8303-4776-8873-e0f744cf96dd · outbound

This paper cites HiT: Building Mapping with Hierarchical Transformers.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs HiT: Building Mapping with Hierarchical Transformers

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.188633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.452471Z digest=sha256:1af6f2aeb08cf5ac1aab5a49ed9365ed9701b6b5206464fb6b19b72f010b1813

Observation f294c6ba-585a-46ea-8119-03ae8b83043c · outbound

This paper cites Gpt4roi: Instruction tuning large language model on region-of-interest.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Gpt4roi: Instruction tuning large language model on region-of-interest

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.174907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.456326Z digest=sha256:aa258570711ef95b1566297a64ea26f9ea578b07a2d2b8d73e02824d0bb433a2

Observation d7f53afa-67ab-4129-8fb1-ac41a3f9eff3 · outbound

This paper cites OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.460373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.460373Z digest=sha256:4eeb345a1fc6bdacb627c23a36cb274b6c5032bdeca10962aca07211164d14c3

Observation 5d347ad1-9a51-40ec-851f-1a87105d9ce2 · outbound

This paper cites Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.464562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.464562Z digest=sha256:937b1b8e4064a700858ab832b065443a693ec9a3718e3ef3436974902aee1779

Observation f1eeaea0-1738-4a20-860f-427110480de5 · outbound

This paper cites E2ec: An end-to-end contour-based method for high-quality high-speed instance segmentation.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs E2ec: An end-to-end contour-based method for high-quality high-speed instance segmentation

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.160009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.468585Z digest=sha256:42a63f4c2072fd9e272673735953e2c1151411d0b39b36181852506646a3f638

Observation 75124c75-c625-4f0c-9a5d-d9db4f6e4197 · outbound

This paper cites P2PFormer: A Primitive-to-polygon Method for Regular Building Contour Extraction from Remote Sensing Images.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs P2PFormer: A Primitive-to-polygon Method for Regular Building Contour Extraction from Remote Sensing Images

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.145458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.472846Z digest=sha256:ae2d3a8155f607fa12437579ce48530a1483c1f4f17da1fa1e273b8a1b577085

Observation e38dff98-5106-4db9-b4a2-82426754e431 · outbound

This paper cites Popeye: A Unified Visual-Language Model for Multi-Source Ship Detection from Remote Sensing Imagery.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Popeye: A Unified Visual-Language Model for Multi-Source Ship Detection from Remote Sensing Imagery

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:46:12.597719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.477081Z digest=sha256:d199baf4202af2f08b173a181cd83c3432adfebf7c08c2c0cbe4a2337fe5e8a4

Observation f2d2a953-4bb8-4b02-a514-c83c161a84e7 · outbound

This paper cites Earthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domain.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Earthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domain

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.130723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.481956Z digest=sha256:e2cce3ab6adb8a24c530fe0e05fc4ad446d3cd154f58c4f6164010b19dc9cf74

Observation a09fa0d4-3ad7-4110-88ca-09c776f96563 · outbound

This paper cites RS5M and GeoRSCLIP: A large scale vision-language dataset and a large vision-language model for remote sensing.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs RS5M and GeoRSCLIP: A large scale vision-language dataset and a large vision-language model for remote sensing

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.117110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.486032Z digest=sha256:189beda6883ea96969243b8ca41ddd7238fb8e18179d93246301fd19a4aa393b

Observation 3a167d90-4e72-487e-a4cc-b679bf24355a · outbound

This paper cites Building extraction from satellite images using mask r-cnn with building boundary regularization.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Building extraction from satellite images using mask r-cnn with building boundary regularization

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.102277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.490111Z digest=sha256:5077dec2a2b106dbbca293c582232d98771a639022857af2569f0964754f8fa4

Observation c48c7d0f-3e53-4584-8f84-413a13286d3f · outbound

This paper cites Building instance segmentation and boundary regularization from high-resolution remote sensing images.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Building instance segmentation and boundary regularization from high-resolution remote sensing images

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.085675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.494249Z digest=sha256:30db8744cec6bbda54c5e70567419c3c787822d102799b36414a97decefcf532

Pith citing papers

Observation 03578315-1c94-4a01-a3a4-eb074b2e5e40 · inbound

HoliTracer: Holistic Vectorization of Geographic Objects from Large-Size Remote Sensing Imagery cites this paper.

HoliTracer: Holistic Vectorization of Geographic Objects from Large-Size Remote Sensing Imagery VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:19:20.268713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T15:19:20.192163Z digest=sha256:2a8cb492e32d027ca1104578dda1cfe0fd0a67c38f7bfd9e06f4c7107fda2e88

Observation 87847722-2f3c-4d59-87fc-5c8b6bb0e045 · inbound

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO cites this paper.

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-14T04:38:05.237334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:38:05.237334Z digest=sha256:7719c0044f6b97f7f3b3b0b3e9dfef5395dda7178911a910fa471c6d648a619a