Pith. sign in

Paper Citation Record · LEDGER

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing

As of 15 August 2026, this Paper Citation Record lists 59 of 59 outbound references and 7 inbound Pith citation observations for arXiv:2501.06828.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.06828 v2

Coverage vector

measured 59 of 59 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T20:53:49.353753Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:09:55.854741Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-15T00:48:24.793323Z

Reference resolution

59 of 59 outbound references displayed

  • verified exact2
  • verified fuzzy21
  • unresolved34
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation b57f24db-3cc5-4524-b60a-e8a3273e966d · outbound

This paper cites RSGPT: A Remote Sensing Vision Language Model and Benchmark.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing RSGPT: A Remote Sensing Vision Language Model and Benchmark

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:47.960477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:47.960477Z digest=sha256:b20455aff193213250b0be6e5f7be64c0ee03ff249ccf5b5ca2eab55b38a6cd9

Observation 479916e7-27cf-4b1e-82f4-0518bc6bb86e · outbound

This paper cites Geochat: Grounded large vision- language model for remote sensing,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Geochat: Grounded large vision- language model for remote sensing,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.402256Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:53:48.024344Z digest=sha256:d75dddecc6739cf28cc372b7255f545c8680254ff8655f967ebc7f8cba4e42e8

Observation c4737594-d29f-47d6-881d-6c84d4b989a5 · outbound

This paper cites SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing SkyEyeGPT: Unifying Remote Sensing Vision-Language Tasks via Instruction Tuning with Large Language Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.029329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.029329Z digest=sha256:49afc724e3b283e61d1cc3b39f2206d445559d57dce2721731392eac1288c999

Observation e517ff45-b20f-4bf6-8563-74498db1a163 · outbound

This paper cites Earthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domain,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Earthgpt: A universal multi-modal large language model for multi-sensor image comprehension in remote sensing domain,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.392167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:53:48.034921Z digest=sha256:fff45f1ec7131b3f14da08f84489ae4a6c6485dcc58111cf3e67aac0f8487310

Observation 74407812-8a35-4aae-8fe8-744d5b6b82ba · outbound

This paper cites Bb-geogpt: A framework for learning a large language model for ge- ographic information science,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Bb-geogpt: A framework for learning a large language model for ge- ographic information science,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.382002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:53:48.039512Z digest=sha256:55df4b4e8dcbb73e558f593c7038e9d46df0132ba8c8efe0ee8959d7acce2531

Observation d87e0878-d02d-47c0-82e8-3265553242e3 · outbound

This paper cites LHRS-Bot: Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language Model.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing LHRS-Bot: Empowering Remote Sensing with VGI-Enhanced Large Multimodal Language Model

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.124779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.124779Z digest=sha256:b7efb4c533a03a5b56b4f128ac5d90b6f6196c05d70065fe3420e0fc9fe597df

Observation 2db5c505-469c-4d4a-8d02-c746e5aa3b60 · outbound

This paper cites Mtp: Advancing remote sensing foundation model via multi-task pre- training,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Mtp: Advancing remote sensing foundation model via multi-task pre- training,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.373021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:53:48.225842Z digest=sha256:9da0f6c9998542aad07fe6dc8d3d59b8ff0317955e0201b6939f3cc35826b5c7

Observation ed8fa80b-baed-4c4e-bb80-6eaef91a6e75 · outbound

This paper cites SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.229191Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.229191Z digest=sha256:29bdf74fa3e5398d0b9f8f32be021dff993ed1675551a58fe484abda733508bd

Observation b4ed543d-7b2f-49b6-8fd0-eaf5c210dcdf · outbound

This paper cites TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing TEOChat: A Large Vision-Language Assistant for Temporal Earth Observation Data

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.233807Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.233807Z digest=sha256:a6416e4c942a601197b947e74f9680df240b4be935decb487615568aeffe0bc9

Observation 44b65459-5320-461b-bf2a-5d9080062463 · outbound

This paper cites Earthmarker: A visual prompting multi-modal large language model for remote sensing,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Earthmarker: A visual prompting multi-modal large language model for remote sensing,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.365058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:53:48.239203Z digest=sha256:372e2990d27649b83d7ba89aabdbace15a9ad60cae643b212696f6ea532c9885

Observation fbd223b6-275f-4767-b043-163aaa67b744 · outbound

This paper cites Rsvg: Exploring data and models for visual grounding on remote sensing data,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Rsvg: Exploring data and models for visual grounding on remote sensing data,

Reference 11

Resolution
malformed identifier
no resolver link, observed 2026-08-10T20:53:48.243234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.243234Z digest=sha256:da6afe699467c4312d561a9461fa4745f8c1f0a66fbf79bba8f8fe6143b31bc6

Observation 74e8ace1-bf93-4ecc-9ce8-3337ffa99726 · outbound

This paper cites Samrs: Scaling- up remote sensing segmentation dataset with segment anything model,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Samrs: Scaling- up remote sensing segmentation dataset with segment anything model,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.355995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:53:48.247566Z digest=sha256:636c0940318ffffc24523be2a0391892c3e5b9d33c2c26b442bfa1cb297b88d9

Observation 676270de-e63e-4f45-8e34-f2115e1e0a37 · outbound

This paper cites Rrsis: Referring remote sensing image segmentation,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Rrsis: Referring remote sensing image segmentation,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.347108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:53:48.251055Z digest=sha256:05097acf8efa657324c7b0f709ae689718dbd210b1600b3e0cf3914f7dd455f8

Observation 0a0d5c78-5d9e-465f-a00d-8f79f4e83ca6 · outbound

This paper cites Rotated Multi-Scale Interaction Network for Referring Remote Sensing Image Segmentation.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Rotated Multi-Scale Interaction Network for Referring Remote Sensing Image Segmentation

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:53:50.015691Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:53:48.254286Z digest=sha256:654d62894c0497d0704788a5921a5e8c9b8749747638da9f398317cd014eeddb

Observation e286015a-c9ee-4903-94b0-ece529454917 · outbound

This paper cites VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing VRSBench: A Versatile Vision-Language Benchmark Dataset for Remote Sensing Image Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.302408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.302408Z digest=sha256:be8eff6116c1b9e4c4362ec84329b686f1e2645fe2a0fc6de4296210dc99490c

Observation e734e1a0-b2b5-494a-b871-5c8184f1f067 · outbound

This paper cites an unresolved cited work.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:53:50.336893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:53:48.371796Z digest=sha256:f6f3d5b8ceb580f044ad27f27033d9c5806705c7361db9603b51f53157f2fb58

Observation 23036da1-0560-4222-9955-73812f1a3d9f · outbound

This paper cites an unresolved cited work.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:53:50.325240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:53:48.375743Z digest=sha256:c2aaa731b8758dc00dfac5eaf392c9270dbe0e315fb7a8b09d869f1966d74f88

Observation 132a4d93-a91a-4b8d-bbd1-c496dcbb6acd · outbound

This paper cites an unresolved cited work.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:53:50.316077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:53:48.380182Z digest=sha256:4498e303859359f70369c3f8dac679f0b64c4e05f50244ac82bd9af6a49b7977

Observation 60e82997-5cfc-43bd-a2fe-72a5345fbd19 · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.385169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.385169Z digest=sha256:59bf14a14c2cf0f7b6566d21fbeb35832120500ed9682bbed626ccedba1dc2e9

Observation 72c10116-e0cb-43bf-af8a-f827f1acfd0d · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.389986Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.389986Z digest=sha256:bbac89f2ec10886a60ff208ea91c7e1a12689b41d87e2a9625cfe1a02b9bc378

Observation 9f5833d6-7dd0-4bd2-9e0e-1dafa976c291 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.394343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.394343Z digest=sha256:1debe7cdcf857d3d1b9d6ebe6e872e7b4fcc4874e9d6e45f324b9b206a88a28c

Observation 2ad8e6e5-a33a-4a77-a1b2-694f80493fcc · outbound

This paper cites CogVLM: Visual Expert for Pretrained Language Models.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing CogVLM: Visual Expert for Pretrained Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.427182Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.427182Z digest=sha256:8bd0c06f4b6718ae91f0565d1b7c5540e4bba7b7e484229f8d699510b6a729ac

Observation d2218fbd-ede2-4ba3-9b6f-0d325e5b5a67 · outbound

This paper cites InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing InstructBLIP: Towards General-purpose Vision-Language Models with Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.493042Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.493042Z digest=sha256:299b601060e0cc2d785821dce3d73dd02b58c7fba6f935e34517c841a867e9c4

Observation 4f6c3bab-f864-42f2-bede-2d92b567c5dd · outbound

This paper cites Otter: A Multi-Modal Model with In-Context Instruction Tuning.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Otter: A Multi-Modal Model with In-Context Instruction Tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.528819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.528819Z digest=sha256:c59839119f41515bbea9cb0adb247604bb00fcc3346d47baee497065c8a0d6d1

Observation ed754b60-189c-4733-bfe6-b64d3a37e722 · outbound

This paper cites Instruction Tuning with GPT-4.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Instruction Tuning with GPT-4

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.533532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.533532Z digest=sha256:cfac5b284ee8f6b9dd4fa536831d4929b034da44387493376f8fc6f2f1644e00

Observation ddc7b7af-d2c7-45dc-a28a-874ae8c0d7c1 · outbound

This paper cites InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing InternGPT: Solving Vision-Centric Tasks by Interacting with ChatGPT Beyond Language

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.539652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.539652Z digest=sha256:df1044fb9459525fae87a114a96d2cfb3adef7271c0e14e9b2f99b225329a2fc

Observation 6a71dd86-2518-4f1c-b76f-f687a36fbd61 · outbound

This paper cites LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing LLaMA-Adapter: Efficient Fine-tuning of Language Models with Zero-init Attention

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.544572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.544572Z digest=sha256:6e31d5cab5aa939d99ac0e7b28cf45a33bdc1562982eac71b17ebb94b3ccba46

Observation 57de04f7-f32d-42a0-82ae-abf953042751 · outbound

This paper cites LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing LLaMA-Adapter V2: Parameter-Efficient Visual Instruction Model

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.548802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.548802Z digest=sha256:6a59856e457cfc3670acd55b0b589d49404c8a92460c141252d4db0c7629f824

Observation b1f5dab7-0915-4a14-8676-0d8951c142fe · outbound

This paper cites LISA: Reasoning Segmentation via Large Language Model.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing LISA: Reasoning Segmentation via Large Language Model

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.552690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.552690Z digest=sha256:fdd105444df98486f206d5dec9f851c6b5a54d50bc72905b3bcc2486e7684c80

Observation ad57b2d2-6bb5-4359-90fd-37aa725ec24e · outbound

This paper cites Segment Anything.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Segment Anything

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.556327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.556327Z digest=sha256:d0eb975ffda6ebbf111bcf6d518be93716a949ece5d8289c96990309ddbc04b9

Observation 86681f54-61b8-4a88-a6f5-273f4ef7fc9f · outbound

This paper cites PixelLM: Pixel Reasoning with Large Multimodal Model.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing PixelLM: Pixel Reasoning with Large Multimodal Model

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.653877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.653877Z digest=sha256:d27fab7be20fbb04ae3218065fa4d2415142d85d191b94d7bb3ac1d2a03295ef

Observation fac7f3f5-a830-40fd-8a84-133276b70349 · outbound

This paper cites GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing GeoPixel: Pixel Grounding Large Multimodal Model in Remote Sensing

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.729794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.729794Z digest=sha256:b54da043fbc3101d10b6bce3ea7f88c515252a2a11223060cfc2abca6098845e

Observation a59bfb6a-f2d4-4226-9d05-6de0780ababb · outbound

This paper cites Object detection in optical remote sensing images: A 13 survey and a new benchmark,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Object detection in optical remote sensing images: A 13 survey and a new benchmark,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.306212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:53:48.733812Z digest=sha256:8c88697bcef28694109a2622ac5b13c68721fb1c7d374ec000a50fc5e0702d8f

Observation 69fb0bfc-0a14-469f-b565-175e31227b04 · outbound

This paper cites Object detection in aerial images: A large-scale benchmark and challenges,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Object detection in aerial images: A large-scale benchmark and challenges,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.295806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:53:48.844459Z digest=sha256:7a07ec6b80fd2152eb306b71aabe9be91d6bdf2bcd1530ed05be1a385b15cc71

Observation c07902bd-c878-46fc-8958-fed9ce27a3a1 · outbound

This paper cites Fair1m: A bench- mark dataset for fine-grained object recognition in high- resolution remote sensing imagery,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Fair1m: A bench- mark dataset for fine-grained object recognition in high- resolution remote sensing imagery,

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.915827Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.915827Z digest=sha256:343603c2628daf29f780b1f674bb88b8d33ac7f3f84bbad59cea739472e8e065

Observation 74699657-9a64-4fbe-815b-0bff53a95b28 · outbound

This paper cites Aid: A benchmark data set for performance evaluation of aerial scene classification,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Aid: A benchmark data set for performance evaluation of aerial scene classification,

Reference 36

Resolution
metadata mismatch
raw_fallback, observed 2026-08-10T20:53:49.755682Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:53:48.920809Z digest=sha256:018aad806b1973cc73826e448c4083c59b604e27a0978d629a9693d82c9bd729

Observation 276bb595-2037-4757-a1ea-05da2a137c85 · outbound

This paper cites Nwpu-crowd: A large-scale benchmark for crowd counting and lo- calization,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Nwpu-crowd: A large-scale benchmark for crowd counting and lo- calization,

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:48.924161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:48.924161Z digest=sha256:4c8d414a9e025cb988d7335db8c1b9e2d7b3b924f77a0d442ca79d0aea95163c

Observation 5cc974fa-04de-4841-94af-f937cc7041a8 · outbound

This paper cites Eu- rosat: A novel dataset and deep learning benchmark for land use and land cover classification,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Eu- rosat: A novel dataset and deep learning benchmark for land use and land cover classification,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.286416Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:53:48.929800Z digest=sha256:cfc821317709a2ce76e2cc25bcd4995ff6cb5954cc287c89697525313f51053b

Observation f6f8f5c1-212f-41f0-bf55-3d951d630960 · outbound

This paper cites Rsvqa: Visual question answering for remote sensing data,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Rsvqa: Visual question answering for remote sensing data,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.275459Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:53:48.934859Z digest=sha256:2164adf46a8d77147b5505f2d94a6060b5a0adb0b3b9227cb3d9f6044c04a6c7

Observation 0b1b3e55-2252-4aa4-99c2-8052b752ed3e · outbound

This paper cites Memory Matching Networks for Genomic Sequence Classification.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Memory Matching Networks for Genomic Sequence Classification

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-10T20:53:49.604414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:53:48.943707Z digest=sha256:177e644da9c64bed48b2c96e30b17b9e1f5e3511ea7d3b44d70b3988bd823e64

Observation f9f2c85f-b234-4ac9-8f1e-f914575d4ec1 · outbound

This paper cites Unsupervised learning using pretrained cnn and associative memory bank,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Unsupervised learning using pretrained cnn and associative memory bank,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.256559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:53:48.948787Z digest=sha256:fceb038dd84f665a81f3aacfff8884aa266456688cc5b6355ca62aee038f545f

Observation d64bc291-b7bb-4597-b758-eedc1136c231 · outbound

This paper cites Point cloud classi- fication via learnable memory bank,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Point cloud classi- fication via learnable memory bank,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.246020Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:53:48.952127Z digest=sha256:a40f5dbc1b034b58f57ca49a03426991f2784b96670d7d717ed87fd966885fb6

Observation 3e1a7e91-9290-4fcd-a985-29f759521b7d · outbound

This paper cites A new approach to automatic memory banking using trace-based address mining,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing A new approach to automatic memory banking using trace-based address mining,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.235618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:53:48.960698Z digest=sha256:89f1fa9a7fbf7be5c69796134578875535517b5401e28d751ac47ace95f0d331

Observation 2c9ba90b-a9b2-406f-8d28-2c936e290849 · outbound

This paper cites Visual anomaly detection via partition memory bank module and error estimation,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Visual anomaly detection via partition memory bank module and error estimation,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.222736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:53:49.040394Z digest=sha256:7fa656085de2073341daa473f2f2357e88d39af7cba9426cf3299a743f5369d8

Observation ef5425ed-c87c-4bb3-ba1c-9a302af34009 · outbound

This paper cites Mamba: Multi-level aggregation via memory bank for video ob- ject detection,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Mamba: Multi-level aggregation via memory bank for video ob- ject detection,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.210886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:53:49.134468Z digest=sha256:785c32042b97803127a13d80204aa8844a7e5d728876e3e90b8e223691223521

Observation 5bfbfdac-8b31-4c20-8aca-15c4f3c3fc2a · outbound

This paper cites Semi-supervised se- mantic segmentation using unreliable pseudo-labels,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Semi-supervised se- mantic segmentation using unreliable pseudo-labels,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.198699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:53:49.138237Z digest=sha256:3571bfcee1cb1c3419d0396a711dc08eb3680786c044fbf0fe82ed229662f344

Observation 415f4932-3d03-44c9-a949-187556a79f3b · outbound

This paper cites Weakly su- pervised semantic segmentation by pixel-to-prototype contrast,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Weakly su- pervised semantic segmentation by pixel-to-prototype contrast,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.187358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:53:49.142549Z digest=sha256:c264f6bf81fa9458fbeeaa60c79691fb6b83680ffef133e0b451bf688e6e5e94

Observation 684a14ed-537d-4607-87aa-0a3b54aafca4 · outbound

This paper cites Memory-based cross-image con- texts for weakly supervised semantic segmentation,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Memory-based cross-image con- texts for weakly supervised semantic segmentation,

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:49.145828Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:49.145828Z digest=sha256:e78e4c1a9b103912a0dad28a2e04b7da5123fc30310aabdb093a03757d3e592a

Observation 22fb036f-3f3b-49b7-9d3f-73c5c9f51bce · outbound

This paper cites SAM 2: Segment Anything in Images and Videos.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing SAM 2: Segment Anything in Images and Videos

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:49.150228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:49.150228Z digest=sha256:9cb214271a702036e639acc0d6ee63aace188a50398976e17265d182ebe98459

Observation 0c4f3726-bc66-403a-8186-5295f5245880 · outbound

This paper cites SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing SkySense: A Multi-Modal Remote Sensing Foundation Model Towards Universal Interpretation for Earth Observation Imagery

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:49.155213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:49.155213Z digest=sha256:efa629aef8dd8efe5b09768075d81e8f06d803b02793eb2e3943bef58c21fb58

Observation acc63c6e-dc1a-491c-9f95-9c707cdbb79b · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing LoRA: Low-Rank Adaptation of Large Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:49.159740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:49.159740Z digest=sha256:67220f3df72468ced4a2c716fb7f2ea73791dcf619108d793d0dd662c39d70e2

Observation fb343ad8-1d33-493f-939c-733a0419130b · outbound

This paper cites an unresolved cited work.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-10T20:53:50.176587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:53:49.163538Z digest=sha256:732fed0e583e94a9e6ef84c8badec00e42db5d3dd45b306ca4fe511498210f82

Observation 27cc5b5b-8844-4ab4-9905-14ce799e1e9c · outbound

This paper cites Bleu: A method for automatic evaluation of machine translation,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Bleu: A method for automatic evaluation of machine translation,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.164726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:53:49.175607Z digest=sha256:74b05cd8e900f6f81d7dad3b45f0a30dfd5eacc8c7f77fc372548a90697f7070

Observation d9a9a7f3-0cd3-45c5-bb48-edcc708c2fcb · outbound

This paper cites Rouge: A package for automatic evaluation of summaries,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Rouge: A package for automatic evaluation of summaries,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:49.236877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:49.236877Z digest=sha256:d65aa5cd8f7ef54d549d1cbc67b69ccf49bdfddc451a497cc40befe7f2b10975

Observation 2072881c-4e83-4d51-b916-1a177d2167bc · outbound

This paper cites Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Meteor: An automatic metric for mt evaluation with improved correlation with human judgments,

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.148712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:53:49.345436Z digest=sha256:debe234c7c7ff457a7f0825a236afb7aba7af6c80c61f6855132e6a8d6421ca4

Observation 6b402188-5f59-45d8-b7e8-43466cc73e3f · outbound

This paper cites Cider: Consensus-based image description evaluation,.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Cider: Consensus-based image description evaluation,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:49.349243Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:49.349243Z digest=sha256:7488ebd3447393c11bd5cd000a38328d0edb838cc81adeee91d12084a9366c9d

Observation e11a5396-91b5-45f2-a37e-e9b145a0dd63 · outbound

This paper cites CLAIR: Evaluating Image Captions with Large Language Models.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing CLAIR: Evaluating Image Captions with Large Language Models

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:49.353753Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:49.353753Z digest=sha256:d387216439a6f78974473f2f2fc8a658a87ab1aa530c6f9d86776d7dc28ac538

Observation 4d9b901b-18c8-4fb0-ad40-bd1ab864714a · outbound

This paper cites 1109 / tgrs.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing 1109 / tgrs

Reference 644

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T20:53:50.265991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-10T20:53:48.939219Z digest=sha256:ab86817415b92a74a6e5678ac4ca6d5d1a99293cc7ebc8f1a01be83bcbc0a057

Observation 718daccd-d34f-4936-b586-5d5d9b13f84f · outbound

This paper cites Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models.

GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing Mini-Gemini: Mining the Potential of Multi-modality Vision Language Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T20:53:49.166936Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T20:53:49.166936Z digest=sha256:d8783f0bdf6a152c916f429f0d4050feead1ba29bed4ecabfcfd9f8fd1d535db

Pith citing papers

Observation af11f5c5-d302-4dc7-a9f3-23d3a0e51e42 · inbound

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions cites this paper.

Chain-of-Talkers (CoTalk): Fast Human Annotation of Dense Image Captions GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T13:09:55.854741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:09:55.854741Z digest=sha256:937f489e956b3fed6a48c015485e7b039a9346598cd538557ee3fb323eb8d354

Observation 834d3338-f0eb-4b07-8389-c1ba5296fd3f · inbound

GeoMag: A Vision-Language Model for Pixel-level Fine-Grained Remote Sensing Image Parsing cites this paper.

GeoMag: A Vision-Language Model for Pixel-level Fine-Grained Remote Sensing Image Parsing GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T19:21:13.258836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T19:21:13.258836Z digest=sha256:d75940c6eb451cd8f31abbf85ba3b7f076b1cabaa3268519a841a6305a4623db

Observation 73294226-ce04-4cd9-bb65-ad36eefdba9f · inbound

SHARP: Spectrum-aware Highly-dynamic Adaptation for Resolution Promotion in Remote Sensing Synthesis cites this paper.

SHARP: Spectrum-aware Highly-dynamic Adaptation for Resolution Promotion in Remote Sensing Synthesis GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-15T00:48:24.795113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T00:46:34.777585Z digest=sha256:b79cf0739740cb6d7fb6aa9b2a23b4c0a4639185de20ea49723dd44cfd45e71f

Observation a421a2c2-4c4e-42fe-bacb-e066eecb8672 · inbound

SHARP: Spectrum-aware Highly-dynamic Adaptation for Resolution Promotion in Remote Sensing Synthesis cites this paper.

SHARP: Spectrum-aware Highly-dynamic Adaptation for Resolution Promotion in Remote Sensing Synthesis GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-14T20:17:28.243545Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T20:17:28.243545Z digest=sha256:dde02ff7a58510bed81a6244df080a2e17c966689346e1c02fbd8385ba0cd04e

Observation 6f3903d3-6baa-450f-87e1-bcfa7ce2f8bd · inbound

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation cites this paper.

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-10T09:23:37.112891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-10T09:20:54.635375Z digest=sha256:d5a1caa064c35d73ac0a91457b26b5f3b49ecf657524f685f616b7a320f36b29

Observation 6b1db3a5-69b2-4226-ab85-d9179c01d41b · inbound

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation cites this paper.

PixDLM: A Dual-Path Multimodal Language Model for UAV Reasoning Segmentation GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-02T16:10:02.323102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:10:02.323102Z digest=sha256:04dd5ccf0e12c3f5d5fb9ff4ba4b8d30d75b35fbf5ff79c1c576c6eac2ee8ae3

Observation 1b903af9-48f8-45b4-8a74-25804925781b · inbound

WeaveEarth: Structured Evidence Construction and Reasoning for Training-Free UHR Remote Sensing Understanding cites this paper.

WeaveEarth: Structured Evidence Construction and Reasoning for Training-Free UHR Remote Sensing Understanding GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T14:09:30.395518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T14:09:30.395518Z digest=sha256:2932d6cccefda1a54558ad271647692ffc9d94df93a42441474e15579ae090ec