Pith. sign in

Paper Citation Record · LEDGER

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs

As of 16 August 2026, this Paper Citation Record lists 100 of 106 outbound references and 2 inbound Pith citation observations for arXiv:2507.04664.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.04664 v1

Coverage vector

measured 100 of 106 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T19:46:12.494249Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:19:20.192163Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T15:19:20.262812Z

Reference resolution

100 of 106 outbound references displayed

  • verified exact1
  • verified fuzzy47
  • unresolved51
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d8210e97-721d-4377-98c5-7fbcf411d4a1 · outbound

This paper cites write newline.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:10.920396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:10.920396Z digest=sha256:6a6d9e0ed3f5d240cf6c668f7b4476c555e1f0642e74ec80f32663d8c84cc507

Observation 284f28ca-349f-44b6-a49a-aee047b86458 · outbound

This paper cites Efficient interactive annotation of segmentation datasets with polygon-rnn++.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Efficient interactive annotation of segmentation datasets with polygon-rnn++

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:11.024550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:11.024550Z digest=sha256:8110b8d04f271f4d6724426637f373825f84685627db8a2e88a4cd61dd80432d

Observation 4ed2ba49-0e13-4285-bc0b-562f5b2176d5 · outbound

This paper cites Qwen Technical Report.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Qwen Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:11.125140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:11.125140Z digest=sha256:8c782b6c56c77686cfdeed571ee7b65777e765dcd509ea9454c45f2ab1bb8011

Observation 63f16326-a8d3-47a3-9f67-0cafaf669179 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:11.275037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:11.275037Z digest=sha256:d527f4f622d283a4026576e7696ecb98b87720b5446c1e9fd67370c4f9f83a3b

Observation 55a2e43a-de79-4c6a-94f2-51ed9ffe9dff · outbound

This paper cites Multi-task learning for segmentation of building footprints with deep neural networks.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Multi-task learning for segmentation of building footprints with deep neural networks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:11.444690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:11.444690Z digest=sha256:e211a62ec6c427575f540a223111438d56a9092e3e6fc61c6154329df73c865b

Observation efa7c1c8-6e2f-4fcd-982f-bd06561d09ba · outbound

This paper cites J., 2019.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs J., 2019

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:11.605083Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:11.605083Z digest=sha256:60af7ac3d2536c7726406b04f1a1a095ce02d766b80494f1438b94e674d4d472

Observation 79d0203d-3bb3-4a7b-beee-c4685cd45c41 · outbound

This paper cites InternLM2 Technical Report.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs InternLM2 Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:11.735318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:11.735318Z digest=sha256:316b008c78d4eb1c52cac777909c39021770a8d0642c43a204cf955b71e0f14f

Observation d361c8c1-8a8f-44a8-a1ab-8d1ea92df04e · outbound

This paper cites Annotating object instances with a polygon-rnn.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Annotating object instances with a polygon-rnn

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:11.861213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:11.861213Z digest=sha256:469152eb7118d885b4462ae7af1467bb1bc81d0834e7ee4574a0a2da46f95cae

Observation d5c7e7e8-8b92-4f60-8cc8-4beaa3a6d456 · outbound

This paper cites ASF-Net: Adaptive Screening Feature Network for Building Footprint Extraction From Remote-Sensing Images.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs ASF-Net: Adaptive Screening Feature Network for Building Footprint Extraction From Remote-Sensing Images

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.035880Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.035880Z digest=sha256:e979387d9386626a7538cdcda445b85255a39889cff7476206c4f5652391c66e

Observation e9ed3f77-ac4b-48de-844d-899ff0a7eb35 · outbound

This paper cites RSPrompter: Learning to prompt for remote sensing instance segmentation based on visual foundation model.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs RSPrompter: Learning to prompt for remote sensing instance segmentation based on visual foundation model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.044430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.044430Z digest=sha256:ffc078570d3414bddef8176388386b58be8c20f7de40a73f2328166fea6b7731

Observation 676c25f8-9990-409e-9867-bed03307dcbc · outbound

This paper cites L., Liu, X., 2020.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs L., Liu, X., 2020

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.108746Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.108746Z digest=sha256:48088aff8ac6bde9af4dcde09a9ec21743df86a91506b3799f4b0ecc153cea29

Observation 0f688f2a-513d-4336-8c6b-2217dd18266d · outbound

This paper cites CGSANet: A Contour-Guided and Local Structure-Aware Encoder--Decoder Network for Accurate Building Extraction From Very High-Resolution Remote Sensing Imagery.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs CGSANet: A Contour-Guided and Local Structure-Aware Encoder--Decoder Network for Accurate Building Extraction From Very High-Resolution Remote Sensing Imagery

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.113348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.113348Z digest=sha256:4c0907f4b2ccf981f42d4db4ab0babec302ddc91cc558f9410f61d4574611a10

Observation 2c0fb206-cf12-4d86-85e9-95a803d45295 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.118150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.118150Z digest=sha256:f7e246c2dc3cd4de68189a6734aa359bdf27ba93b5e68f62c65ed059555e2d5a

Observation 098777dd-1745-4fc0-8d9d-fa06d2e57f0b · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.123184Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.123184Z digest=sha256:a60aa83a33351eafb76b4bfb96c6ba30d95d2e6af9f717e5426c47ba1d96c71e

Observation 37d15ac5-d595-4ae6-85ce-e59ac22db8af · outbound

This paper cites et al., 2024d.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs et al., 2024d

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.128509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.128509Z digest=sha256:bec6a07588acf80181bdf27d4634afcdac7c4731b467c0262f9bd17ec775503e

Observation 4a946523-bd2f-4276-853f-43490e74626e · outbound

This paper cites G., Kirillov, A., Girdhar, R., 2022.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs G., Kirillov, A., Girdhar, R., 2022

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.132674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.132674Z digest=sha256:52f12fa48c53b0539ee4128194e554ef1bed440223461cfb41e10a4e50144dd2

Observation ee994916-5335-4e80-af27-bb9bd9bc5ecf · outbound

This paper cites E., Stoica, I., Xing, E.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs E., Stoica, I., Xing, E

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.140442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.140442Z digest=sha256:7b817609cf5538afd7469a26fca5786ea4e4ea92ea13b55ecda2007be130bf1f

Observation 03a7438f-6450-4b4f-b99e-11d5ef313a01 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.145107Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.145107Z digest=sha256:25e37c724221430288368b5b63ec433397ccd0268399d87b4834a4828eeacc4c

Observation 72388a81-c250-4eec-a565-219e61eab721 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.149426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.149426Z digest=sha256:9ccfd002ff1a25d45040ab82d8b0042df90f7accca83609c375b4bdee965052c

Observation f9a3c9a8-4cf6-4041-a84b-2654ce84d911 · outbound

This paper cites The Llama 3 Herd of Models.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs The Llama 3 Herd of Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.153773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.153773Z digest=sha256:fe4ac66f882b518173dab42ed97d8cee2176b69d75455a65b0013263d13cc29e

Observation 402e5d75-ed0f-4680-a137-99357aa52267 · outbound

This paper cites GeoLLaVA: Efficient Fine-Tuned Vision-Language Models for Temporal Change Detection in Remote Sensing.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs GeoLLaVA: Efficient Fine-Tuned Vision-Language Models for Temporal Change Detection in Remote Sensing

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.158072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.158072Z digest=sha256:c72966ea49fb55509eee6d91b5d23ef0ad385cf2a80d84fa8c63014d66dd187f

Observation 2644427a-76f4-458a-ae86-b42183c9ab8c · outbound

This paper cites Instances as queries.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Instances as queries

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.162712Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.162712Z digest=sha256:51770f16f8c8b1b8079e40cf089d4625ee7442fae543d27a880fedeae48a0933

Observation 3417c8d8-f6f3-49f5-80c0-f6220469bb90 · outbound

This paper cites GPT-3: Its nature, scope, limits, and consequences.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs GPT-3: Its nature, scope, limits, and consequences

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.166733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.166733Z digest=sha256:e65ace37fa0004e9f69f3eaa2b4c4d7f2980e7483981cf89634477ffb422069b

Observation 22fb0106-5f5a-4949-b4c3-7d169df1fd6e · outbound

This paper cites Polygonal building extraction by frame field learning.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Polygonal building extraction by frame field learning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.170835Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.170835Z digest=sha256:0fad17ee48cf46ca49d176c55e6d4531b10f2104052855e505b970072740819f

Observation fec73877-e797-40a4-a79b-5ad3e5d28eec · outbound

This paper cites Remote Sensing ChatGPT: Solving Remote Sensing Tasks with ChatGPT and Visual Models.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Remote Sensing ChatGPT: Solving Remote Sensing Tasks with ChatGPT and Visual Models

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.175344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.175344Z digest=sha256:94fb2457f2984c854758ec02b8a1852fcf8dca647e251cd1dc5fd083ecd1e084

Observation f238ad8e-3b6b-43ad-8503-5e493a6da9c9 · outbound

This paper cites HigherNet-DST: Higher resolution network with dynamic scale training for rooftop delineation.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs HigherNet-DST: Higher resolution network with dynamic scale training for rooftop delineation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.179579Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.179579Z digest=sha256:ec16c7fb84650994bea6352f5e01786fc896318fddbf57c6458d132bc5d18615

Observation c8a427c6-96d6-4c74-a134-891fa92077fc · outbound

This paper cites Mask r-cnn.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Mask r-cnn

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.183884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.183884Z digest=sha256:17635fac6860696aa23a32311c7dc60ed96a2922ccf0d38862941e63442ccb9b

Observation d390bb96-5275-4feb-ae74-c349d1d523d6 · outbound

This paper cites Rsgpt: A remote sensing vision language model and benchmark.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Rsgpt: A remote sensing vision language model and benchmark

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.187715Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.187715Z digest=sha256:738e94f796037ee25f040db95e1f97f3680f6e82f8b0e17ca68a3c8221bbb9b0

Observation 7a2c0461-438f-4b96-87cb-4627b4b0c473 · outbound

This paper cites Sequentially delineation of rooftops with holes from VHR aerial images using a convolutional recurrent neural network.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Sequentially delineation of rooftops with holes from VHR aerial images using a convolutional recurrent neural network

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.191988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.191988Z digest=sha256:54f9254752e2ea1ed670c5c152fb51d6e2f6492ba696e5bc098200bc1a6a8315

Observation 4336df52-0d5d-43e9-9c6c-c7e0fb34713f · outbound

This paper cites OEC-RNN: Object-oriented delineation of rooftops with edges and corners using the recurrent neural network from the aerial images.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs OEC-RNN: Object-oriented delineation of rooftops with edges and corners using the recurrent neural network from the aerial images

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.196441Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.196441Z digest=sha256:87371439e61b5b6d0bf18ee4b852a35850d049b05192b0a75d30edcedf54492c

Observation 09dde30f-882b-419e-83a3-f1e85ca77997 · outbound

This paper cites TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.200812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.200812Z digest=sha256:d11bbfa08c6d6831012d6f435386b5e5962fe007264411b109c1acf3c869957c

Observation 7d30da99-0abc-453d-b115-07c3af35c33e · outbound

This paper cites Fully convolutional networks for multisource building extraction from an open aerial and satellite imagery data set.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Fully convolutional networks for multisource building extraction from an open aerial and satellite imagery data set

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.205089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.205089Z digest=sha256:3a3c8ff4237718cba1d2b819fdeeaae382f4d633fa85c4aa8134d10923fe7246

Observation dcd086c9-29ac-424a-8f5d-3342a2617651 · outbound

This paper cites A scale robust convolutional neural network for automatic building extraction from aerial and satellite imagery.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs A scale robust convolutional neural network for automatic building extraction from aerial and satellite imagery

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.209483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.209483Z digest=sha256:edef1d7d2bdbeaeb58bfd38a6baace965c81b7d92086b7fa5a3fe7ef22db021f

Observation 889edfd5-384c-43d1-83c6-c19aae744b8c · outbound

This paper cites Segment Anything.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Segment Anything

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.214127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.214127Z digest=sha256:537d8d4ef053e991395c4ab294e7835859fa0178141aff616259fb8fb2a90781

Observation 54e8859c-9ea6-43ac-868d-0d2187b98f52 · outbound

This paper cites C., Lo, W.-Y.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs C., Lo, W.-Y

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.218573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.218573Z digest=sha256:1b1b673efcc3936145da3a5c06bbfe2f579e07ef1bb6a513070d5036c325ead7

Observation 62b9df7f-f2c1-4fa1-99c1-dd732c0569ce · outbound

This paper cites E., 2017.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs E., 2017

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.222970Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.222970Z digest=sha256:c129e44594d008ed0910ee66b793c793139275082511f987280b6a3776e58f76

Observation a82c35c5-a1ed-4198-92a3-a2b2040f4336 · outbound

This paper cites S., Naseer, M., Das, A., Khan, S., Khan, F.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs S., Naseer, M., Das, A., Khan, S., Khan, F

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.227214Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.227214Z digest=sha256:6bedb23ac909dc4f16da0a408d2b5b711a85c90c1e9d5992029e1c27bc38d66b

Observation bf74aaef-e9c3-4173-8e2e-12842ec9cdaa · outbound

This paper cites Lisa: Reasoning segmentation via large language model.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Lisa: Reasoning segmentation via large language model

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.768609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.231466Z digest=sha256:9ed49c61e0a9b9e7c5359f7528b417c0465a6e425ed5352b9d1436ceb45c72e5

Observation 83f5ad00-4f16-4c13-ae3c-d5093f990ca2 · outbound

This paper cites VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.235897Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.235897Z digest=sha256:405397531ad5b5ff4eca16c6858db54c5aa2c230f8754669580e4f9325198679

Observation c7f65599-2566-4ca8-946f-9be25206b819 · outbound

This paper cites Topological Map Extraction from Overhead Images.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Topological Map Extraction from Overhead Images

Reference 40

Resolution
metadata mismatch
local_arxiv, observed 2026-08-06T19:46:12.810473Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.240475Z digest=sha256:abcbda01611eca138bba8001aad39d19be4e125e4331be2936c0ee426680b0cd

Observation 3ce7b139-c1b1-4957-8785-0d8f624a62bc · outbound

This paper cites D., Lucchi, A., 2019.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs D., Lucchi, A., 2019

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.750718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.247118Z digest=sha256:6bc99430df80714a9233c8046ae7b539ac773419d526cc3ad81c00f9329f3e92

Observation f2c10974-566d-4cc7-91ad-596ab7bb7a93 · outbound

This paper cites Polytransform: Deep polygon transformer for instance segmentation.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Polytransform: Deep polygon transformer for instance segmentation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.735684Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.251308Z digest=sha256:8311fe4e1157189606b04d3ca0efdfc1ada2e4ac87cf5c1c4695b57fa6aa5d07

Observation d37fc12e-ba8d-4aed-8492-ab0056e2f0a9 · outbound

This paper cites L., 2014.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs L., 2014

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.255657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.255657Z digest=sha256:627745596e100997fe208f399b0169a7f700c6b0b22fbd3dec476d1368f7c081

Observation 827c00df-b728-4e03-9d0f-d55d15100aa4 · outbound

This paper cites Draw-and-understand: Leveraging visual prompts to enable mllms to comprehend what you want.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Draw-and-understand: Leveraging visual prompts to enable mllms to comprehend what you want

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.709894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.259708Z digest=sha256:4c081d9c2b49412250210308b8554b1ee8beb2afbd55da8e67c18524c9e9f59e

Observation d0bbce28-e468-439e-99ca-15afbfcb788e · outbound

This paper cites Change-agent: Towards interactive comprehensive remote sensing change interpretation and analysis.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Change-agent: Towards interactive comprehensive remote sensing change interpretation and analysis

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.696936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.263754Z digest=sha256:4da2f191f15fc55a5264738f848d8aa8e1cd22a66787c2ed7a59aee16988aa75

Observation d3ad047d-f059-4b44-aadd-03e46058c842 · outbound

This paper cites J., 2023a.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs J., 2023a

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.683383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.268556Z digest=sha256:ff1fb5340d58b211d63953e1005f83475bdfc3b7293b6d3b8cb6be534eec6af3

Observation c9f2d888-2725-4ade-8a56-de7b935ab042 · outbound

This paper cites J., 2023b.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs J., 2023b

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.669499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.272742Z digest=sha256:85a8b1a91f76e3cad2abc97bcb1eb20038f1483bcb0113fd6ced3af5abfa6f78

Observation 78abfdac-54bd-486c-a69d-ede748ad8ac3 · outbound

This paper cites Path aggregation network for instance segmentation.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Path aggregation network for instance segmentation

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.655042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.277029Z digest=sha256:10106858ab33375650ab83bc94ea7017b22f19bc3509479e637bda3facdcdc45

Observation e24ef6bb-7b5e-4eb5-b2b1-cf06837ae3e4 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Swin transformer: Hierarchical vision transformer using shifted windows

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.641484Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.281467Z digest=sha256:7342c9d67e1911a6d080af683b164140b02828d33861c22ac7faacfd7928b744

Observation ce1e00f3-4929-46f3-b4f9-6d7406b15fc6 · outbound

This paper cites Building Outline Delineation From VHR Remote Sensing Images Using the Convolutional Recurrent Neural Network Embedded With Line Segment Information.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Building Outline Delineation From VHR Remote Sensing Images Using the Convolutional Recurrent Neural Network Embedded With Line Segment Information

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.627700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.286109Z digest=sha256:6fbbe811285e777a49ecd288bf18bee872dd571802acd930e5e8bf4d67a40b1a

Observation b13b024e-9a19-4236-ac5f-2454e73df9d4 · outbound

This paper cites SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.290250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.290250Z digest=sha256:f6fc68c3d0494a78c550d2381b1365992522d1941780b848a22816e800b586d5

Observation 1b14cb8b-87f3-46a5-958f-6c3c4d137129 · outbound

This paper cites Cross-spatiotemporal land-cover classification from VHR remote sensing images with deep learning based domain adaptation.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Cross-spatiotemporal land-cover classification from VHR remote sensing images with deep learning based domain adaptation

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.613358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.295054Z digest=sha256:db329109e098a16911a35cd7c48c4b0a742517a1545c2889d5f7075055ebfb53

Observation 686281e2-c5b5-4920-8bb8-32a113dcae3c · outbound

This paper cites SAM-RSIS: Progressively adapting SAM with box prompting to remote sensing image instance segmentation.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs SAM-RSIS: Progressively adapting SAM with box prompting to remote sensing image instance segmentation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.597708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.299089Z digest=sha256:30e1c233297d155b34651fe2716979129bd700d9582eb6f412e5810c22cdcfe5

Observation 29af8f0c-1af2-430f-bd2f-6a6bca28fb56 · outbound

This paper cites P., 2019 (accessed November 10, 2019).

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs P., 2019 (accessed November 10, 2019)

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.583455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.303073Z digest=sha256:2dad956c8890c0f02fa7f8c7a00b928f90cb91996930256208a73df7329f5a3e

Observation 8df539d9-a371-4125-b177-b49eacccc809 · outbound

This paper cites LHRS-Bot: Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language Model.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs LHRS-Bot: Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language Model

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.307217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.307217Z digest=sha256:d44f5264032c15128406ec688bd5d089908f90e60b3bbb11b90ab4017f4ef279

Observation 128196e5-76c8-4d90-b068-f37701d3c630 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs DINOv2: Learning Robust Visual Features without Supervision

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.311807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.311807Z digest=sha256:27a31b04a320b34adb6c67b836725009d075e43fe7c49cd3a3aa1dcf3797faad

Observation 3f4aace0-37a0-4b07-b11b-ba9fbb559ba7 · outbound

This paper cites VHM: Versatile and Honest Vision Language Model for Remote Sensing Image Analysis.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs VHM: Versatile and Honest Vision Language Model for Remote Sensing Image Analysis

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.316595Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.316595Z digest=sha256:5a0c23a06b0fb49f818ac8ee752db76a2cbcde63b94c36fcce2678832b6bd045

Observation f887536e-9793-4d10-97a4-26108a357f69 · outbound

This paper cites Deep snake for real-time instance segmentation.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Deep snake for real-time instance segmentation

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.567120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.320655Z digest=sha256:6d131abf5a5701bfcb602741bd363bdfe9931626bc508b784d915e088f12a7dd

Observation 4c97ddd7-5460-4792-b3ba-aa1775f656b2 · outbound

This paper cites W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs W., Hallacy, C., Ramesh, A., Goh, G., Agarwal, S., Sastry, G., Askell, A., Mishkin, P., Clark, J

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.552334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.324853Z digest=sha256:7dc68e43a5f1b2e723eccf8b2be28d5fa37ac8c334a1568fed5f40156a1bdbc3

Observation c54176be-9c7f-44e2-8812-74181bd3f3f5 · outbound

This paper cites D., Ermon, S., Finn, C., 2024.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs D., Ermon, S., Finn, C., 2024

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.536754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.328747Z digest=sha256:9010f11818f010eb25d975fc069a9038abef6576ac38de195629b0bc8aea9b60

Observation 0cd113dd-ee83-4552-b263-db1eb4fa46d9 · outbound

This paper cites Am-radio: Agglomerative vision foundation model reduce all domains into one.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Am-radio: Agglomerative vision foundation model reduce all domains into one

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.521226Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.332948Z digest=sha256:e2ef0de76a1ae4160a4b9762820145b81d312209458a1482bb65d3f5e2cc3301

Observation aba2b761-0c8c-4030-8fed-7672c527f1fd · outbound

This paper cites M., Xing, E., Yang, M.-H., Khan, F.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs M., Xing, E., Yang, M.-H., Khan, F

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.504453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.336760Z digest=sha256:61ac7f5bad5507bd06ec7195434bec98ec7f4ff1dcac26e250cb51ee90f1fe29

Observation 9aad2ccb-9034-46e9-a8aa-bd39225ce36f · outbound

This paper cites Pixellm: Pixel reasoning with large multimodal model.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Pixellm: Pixel reasoning with large multimodal model

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.490327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.340577Z digest=sha256:68add70bbe1d65806bae4b62bc1f1afaf7b4756750217711355618d869de816d

Observation 540c0d6d-2bb8-433d-807a-fc007055224f · outbound

This paper cites Geollm-engine: A realistic environment for building geospatial copilots.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Geollm-engine: A realistic environment for building geospatial copilots

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.473850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.345906Z digest=sha256:b5169a07b22635cbcf34314eaad6d628f7ccc24d03b14eb82798617cb997bb42

Observation cc1b2af8-ef4a-471e-9c88-355c64381b81 · outbound

This paper cites Internlm: A multilingual language model with progressively enhanced capabilities.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Internlm: A multilingual language model with progressively enhanced capabilities

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.459916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.350574Z digest=sha256:8914867151384cf8123e26e411944d2043e7629291feb441adb7d976f80b5bcb

Observation df3fd4f0-9393-4133-a26b-81021c816039 · outbound

This paper cites Fcos: Fully convolutional one-stage object detection.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Fcos: Fully convolutional one-stage object detection

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.446080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.354993Z digest=sha256:a4a1fbda0dd53722b6fc45c89f45d5b633e0656b64ff95e54af3a582e98ea98f

Observation c75b4a2d-c816-4c9d-afe0-9087289aa521 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.359024Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.359024Z digest=sha256:fce4a028149875271601092d3f2a0cf95046059217639bb104bcab294917ecc2

Observation 5fb2122f-b754-48a1-937a-73536643f55e · outbound

This paper cites From image transfer to object transfer: Cross-domain instance segmentation based on center point feature alignment.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs From image transfer to object transfer: Cross-domain instance segmentation based on center point feature alignment

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.432782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.363048Z digest=sha256:5fe7a501d19051b8bd63138213f7a01e480ed2adeee618c1943ab4ad81f752fa

Observation 33524a30-b01d-41a9-a075-6ef26917ce84 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.367046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.367046Z digest=sha256:7fa427ba44be27b3e79042e6cb2431adf7d5299243c1d890e156b1c7440596cf

Observation e4eb0793-f138-4d14-bd7d-67ba16e7bf46 · outbound

This paper cites et al., 2024b.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs et al., 2024b

Reference 70

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.418753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.371187Z digest=sha256:feb61263c36e912782ceeee4b76f65fec536bcf30b7fea913e6eb935bb516096

Observation 54bb2a9e-a2c6-4242-9dfc-8fcac75dd936 · outbound

This paper cites Solo: Segmenting objects by locations.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Solo: Segmenting objects by locations

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.405054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.375385Z digest=sha256:bf48f34fd5fe85adfb5d04ea256c12758f729cb87bb5452f4cdcf38a2d0783f9

Observation 120e7511-4cf5-4df1-9071-b019c8b8f415 · outbound

This paper cites Skyscript: A large and semantically diverse vision-language dataset for remote sensing.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Skyscript: A large and semantically diverse vision-language dataset for remote sensing

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.391338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.379117Z digest=sha256:8f294506c983b086b1a513f56e967c8b9d7cd84869b19db5c279139958edec3a

Observation 20acfc7f-f109-4a06-ac92-5b3d82a53cde · outbound

This paper cites Graph convolutional networks for the automated production of building vector maps from aerial images.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Graph convolutional networks for the automated production of building vector maps from aerial images

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.376449Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.383357Z digest=sha256:bcc6a388650556415ae39cb13caf2fa24ff4f622d6c560e05b4cb473fa0a2ec6

Observation bc4e1410-b3d9-43c9-9dc1-bae1c1a1f4d7 · outbound

This paper cites Toward automatic building footprint delineation from aerial images using CNN and regularization.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Toward automatic building footprint delineation from aerial images using CNN and regularization

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.361186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.387107Z digest=sha256:dfb4991488f204b21f621e070e7c5b4227dc9e52d0e2bbe3bf7d75377a84ac02

Observation b74a44fc-cae6-4e8c-811c-15f90523039e · outbound

This paper cites A Concentric Loop Convolutional Neural Network for Manual Delineation-Level Building Boundary Segmentation From Remote-Sensing Images.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs A Concentric Loop Convolutional Neural Network for Manual Delineation-Level Building Boundary Segmentation From Remote-Sensing Images

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.346545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.391234Z digest=sha256:eb94c3d879c752bded7c6cae0e1734ab8b8ffcf84d0d729c847ca5ee8d202ebd

Observation 681bfab7-2255-4e1e-8b9c-52960724f664 · outbound

This paper cites BuildMapper: A fully learnable framework for vectorized building contour extraction.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs BuildMapper: A fully learnable framework for vectorized building contour extraction

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.332317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.395043Z digest=sha256:c248c6b3ac7ee17597842d31aff9ee9048c30b499cd5a2f091949554fe6a66d2

Observation 33d5c8b8-0ad6-43f7-b51f-561cfea87f99 · outbound

This paper cites From lines to Polygons: Polygonal building contour extraction from High-Resolution remote sensing imagery.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs From lines to Polygons: Polygonal building contour extraction from High-Resolution remote sensing imagery

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.319037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.399002Z digest=sha256:bab76e781694074cf606f4608c6a219f76112a5bd048c70d9f0e026833806a22

Observation d6e94cea-aa35-4323-b476-edf1d9b58ca1 · outbound

This paper cites A., 2022.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs A., 2022

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.305215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.403050Z digest=sha256:3167732caeca849ca6326489c367c6aeed9477e257316e9e57c612dc9368dc77

Observation 9e79cef5-9028-487b-a278-f3bfd73a6cdd · outbound

This paper cites Vectorizing historical maps with topological consistency: A hybrid approach using transformers and contour-based instance segmentation.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Vectorizing historical maps with topological consistency: A hybrid approach using transformers and contour-based instance segmentation

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.291113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.406984Z digest=sha256:801cee95492b9b63a3266fa31c7ded131d67b52af917b8ebf14efc6622ff9df1

Observation 05abb26e-ca00-4925-a7c1-12782aa2dbdc · outbound

This paper cites Video instance segmentation is all you need for linking geographic entities from historical maps.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Video instance segmentation is all you need for linking geographic entities from historical maps

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.275145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.410892Z digest=sha256:d28365aa409c189bc3e52af34238477299b4e352608bdb4cfc2f3c2b21de27f6

Observation d787037c-3fbf-4ccc-b604-62d610df6fde · outbound

This paper cites Polarmask: Single shot instance segmentation with polar representation.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Polarmask: Single shot instance segmentation with polar representation

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.261366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.415058Z digest=sha256:02adee47092952d1aaf97035d8915ce9bece5f99dc20e15ada1ab2d4b46815f5

Observation 939ef4fa-5a08-4f74-aba5-74d9dfcd08bf · outbound

This paper cites HiSup: Accurate polygonal mapping of buildings in satellite imagery with hierarchical supervision.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs HiSup: Accurate polygonal mapping of buildings in satellite imagery with hierarchical supervision

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.246162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.419407Z digest=sha256:1193a6cac62432765a74d402e03ec21d8b55013f9a548bdb7c15d52e10ee73c9

Observation 61e0fd6c-7086-41b8-a439-e08f2b35bd76 · outbound

This paper cites Qwen3 Technical Report.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Qwen3 Technical Report

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.423489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.423489Z digest=sha256:0a9065394a8e4baceb89fe5b24ad648d98b80e45b04050aac3376cb879f6c1e8

Observation 45a3c82f-e62d-44f6-be72-1e0b36778ccb · outbound

This paper cites Qwen2 Technical Report.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Qwen2 Technical Report

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.427552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.427552Z digest=sha256:01e76280b2b08e400e455c156fbab4cca009ccb12c6734db5754c9641d0bbc74

Observation 02d2e529-04d6-463c-b965-a3a9370fd2c3 · outbound

This paper cites A review of recurrent neural networks: LSTM cells and network architectures.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs A review of recurrent neural networks: LSTM cells and network architectures

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.231469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.431291Z digest=sha256:fd3fa7167abe60842aa84c1d5b7b1c4bfaa512861e6fd3f0371702642e7685e1

Observation b59a960f-b2db-46c5-b692-a14e5a8fb91e · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.435161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.435161Z digest=sha256:82ac62f1fd8e7a18b2210baf3cf64f60b1cb1e2f71005c2e2a6eefca8d5affd1

Observation 90967185-4b8f-4448-80df-44aa4b22d059 · outbound

This paper cites Learning building extraction in aerial scenes with convolutional networks.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Learning building extraction in aerial scenes with convolutional networks

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.216943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.439134Z digest=sha256:07d09d7be2c6ad69d8ceac68fcf66a18867409a85e30cfc28ebe981d9f6611fc

Observation 184b47f0-cbea-4043-99d8-a666a269c6b0 · outbound

This paper cites Osprey: Pixel understanding with visual instruction tuning.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Osprey: Pixel understanding with visual instruction tuning

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.202555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.443193Z digest=sha256:705777e3e15d2032f7c657e7cc0da6cd155c5254b8f4849dbc891d4d80897d31

Observation 22a3060a-2a1b-45da-a5e6-2f7ab87ee0dc · outbound

This paper cites SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.447693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.447693Z digest=sha256:70ebfdeea63251fae17b2115433aca7f88131cd85cd0eb1f8b9e52cd24dcc0bc

Observation 73189925-8303-4776-8873-e0f744cf96dd · outbound

This paper cites HiT: Building Mapping with Hierarchical Transformers.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs HiT: Building Mapping with Hierarchical Transformers

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.188633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.452471Z digest=sha256:f3c025af092ccbc61ecbc764753620500b53cbdf26f18d780c314c9af93c9a44

Observation f294c6ba-585a-46ea-8119-03ae8b83043c · outbound

This paper cites Gpt4roi: Instruction tuning large language model on region-of-interest.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Gpt4roi: Instruction tuning large language model on region-of-interest

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.174907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.456326Z digest=sha256:7bf5b5e5251c40bd2b815b9f42e2b258cf59198dbfe9bfe1d757af69a0060f84

Observation d7f53afa-67ab-4129-8fb1-ac41a3f9eff3 · outbound

This paper cites OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs OMG-LLaVA: Bridging Image-level, Object-level, Pixel-level Reasoning and Understanding

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.460373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.460373Z digest=sha256:4eeb345a1fc6bdacb627c23a36cb274b6c5032bdeca10962aca07211164d14c3

Observation 5d347ad1-9a51-40ec-851f-1a87105d9ce2 · outbound

This paper cites Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-06T19:46:12.464562Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T19:46:12.464562Z digest=sha256:937b1b8e4064a700858ab832b065443a693ec9a3718e3ef3436974902aee1779

Observation f1eeaea0-1738-4a20-860f-427110480de5 · outbound

This paper cites E2ec: An end-to-end contour-based method for high-quality high-speed instance segmentation.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs E2ec: An end-to-end contour-based method for high-quality high-speed instance segmentation

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.160009Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.468585Z digest=sha256:74a8273d26d8bf16f09416ce639217ab5eca301ec750065f5e6145be73158abf

Observation 75124c75-c625-4f0c-9a5d-d9db4f6e4197 · outbound

This paper cites P2PFormer: A Primitive-to-polygon Method for Regular Building Contour Extraction from Remote Sensing Images.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs P2PFormer: A Primitive-to-polygon Method for Regular Building Contour Extraction from Remote Sensing Images

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.145458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.472846Z digest=sha256:40b1bdc15f2c23c3145e87f87377a6929638d3bb1a2a17b7fa231bf48d3b9452

Observation e38dff98-5106-4db9-b4a2-82426754e431 · outbound

This paper cites Popeye: A Unified Visual-Language Model for Multi-Source Ship Detection from Remote Sensing Imagery.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Popeye: A Unified Visual-Language Model for Multi-Source Ship Detection from Remote Sensing Imagery

Reference 96

Resolution
verified exact
local_arxiv, observed 2026-08-06T19:46:12.597719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.477081Z digest=sha256:08e55e032d6a0e84d9f49279b9fd8ea435623d85ae93e3e42beb1a9180e1cc44

Observation f2d2a953-4bb8-4b02-a514-c83c161a84e7 · outbound

This paper cites Earthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domain.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Earthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domain

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.130723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.481956Z digest=sha256:74883dbde102078683e81c3c83906468b73d6490ed9c8b2f69024b3346a03d9c

Observation a09fa0d4-3ad7-4110-88ca-09c776f96563 · outbound

This paper cites RS5M and GeoRSCLIP: A large scale vision-language dataset and a large vision-language model for remote sensing.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs RS5M and GeoRSCLIP: A large scale vision-language dataset and a large vision-language model for remote sensing

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.117110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.486032Z digest=sha256:6167c1692fcb5ff8b659127c3527cc442865c25b76a87ddf9deff01dd421181d

Observation 3a167d90-4e72-487e-a4cc-b679bf24355a · outbound

This paper cites Building extraction from satellite images using mask r-cnn with building boundary regularization.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Building extraction from satellite images using mask r-cnn with building boundary regularization

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.102277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.490111Z digest=sha256:130e18a4ccd32927426b57895cee38ef1cff01681f1ba826636ca46b69826237

Observation c48c7d0f-3e53-4584-8f84-413a13286d3f · outbound

This paper cites Building instance segmentation and boundary regularization from high-resolution remote sensing images.

VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs Building instance segmentation and boundary regularization from high-resolution remote sensing images

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T19:46:13.085675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T19:46:12.494249Z digest=sha256:bf29bc86cfe5d9df3a16ce2fe2def2c0376a4ffad7104f5bec7c4e1c7b3e1db2

Pith citing papers

Observation 03578315-1c94-4a01-a3a4-eb074b2e5e40 · inbound

HoliTracer: Holistic Vectorization of Geographic Objects from Large-Size Remote Sensing Imagery cites this paper.

HoliTracer: Holistic Vectorization of Geographic Objects from Large-Size Remote Sensing Imagery VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:19:20.268713Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:19:20.192163Z digest=sha256:fd0f5653e124f1fe1936b7cba93e4bef832ed72c2ba5d862e5fafdc0568968d5

Observation 87847722-2f3c-4d59-87fc-5c8b6bb0e045 · inbound

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO cites this paper.

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO VectorLLM: Human-like Extraction of Structured Building Contours vis Multimodal LLMs

Reference 71

Resolution
unresolved
no resolver link, observed 2026-07-14T04:38:05.237334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:38:05.237334Z digest=sha256:7719c0044f6b97f7f3b3b0b3e9dfef5395dda7178911a910fa471c6d648a619a