Pith. sign in

Paper Citation Record · LEDGER

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features

As of 19 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2501.10144.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.10144 v1

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T19:26:44.966467Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved11
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c1838a2c-9fe3-4eb5-bfcd-0e56e5146ecd · outbound

This paper cites Canopy averaged chlorophyll content pre- diction using convolutional autoencoder on hyperspectral data,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Canopy averaged chlorophyll content pre- diction using convolutional autoencoder on hyperspectral data,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.545801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T19:26:44.847841Z digest=sha256:d77d7b5f2c39524894ee26cd9111437d04b9335a1e597620b3e9de56370f5de5

Observation 1694854d-5b7e-4220-8dee-6b573e39c490 · outbound

This paper cites Enhancing deforestation monitoring in the brazilian amazon: A semi-automatic approach leveraging uncer- tainty estimation,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Enhancing deforestation monitoring in the brazilian amazon: A semi-automatic approach leveraging uncer- tainty estimation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.532352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T19:26:44.852903Z digest=sha256:00a3c3a350c33e8c006ac5bb3ddca9978a737754d594a7b9fcbd19310fc82ed8

Observation bc850173-9504-4c33-9c38-0a9ff8a53add · outbound

This paper cites Sen1floods11: a georeferenced dataset to train and test deep learning flood algorithms for sentinel-1,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Sen1floods11: a georeferenced dataset to train and test deep learning flood algorithms for sentinel-1,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.517400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T19:26:44.857234Z digest=sha256:9cb02644cfcd320ecd05fdb57aacc5cf17ec3addbc5b5f3db09d7e1094ef58e3

Observation dbe34249-0774-4a29-bdf4-dcb65b466d23 · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Improved Baselines with Visual Instruction Tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-10T19:26:44.861667Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:26:44.861667Z digest=sha256:944262420d974ea94593aca70031eb7c179c8b1793b623acd42459904dc64487

Observation 4cf2e9f3-cdc3-4d4f-a818-d66cdceaaecd · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Llava-next: Improved reasoning, ocr, and world knowledge,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-10T19:26:44.866767Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:26:44.866767Z digest=sha256:57bdad10c92cac66cc2651faae4d74610b31d9580da933483e7ad7df17531393

Observation 4b4ef0d6-8697-41fc-9267-72625460c83d · outbound

This paper cites Visual instruction tuning,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Visual instruction tuning,

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T19:26:44.870831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:26:44.870831Z digest=sha256:40a837b61cc77e8908331e81e14a942b8fb0c0b96ccbbf715cc874b0314b1b88

Observation eb33002a-ee4f-46f5-ad82-f8456ef1c552 · outbound

This paper cites Blip- 3: A family of open large multimodal models,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Blip- 3: A family of open large multimodal models,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-10T19:26:44.875200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:26:44.875200Z digest=sha256:9c4240680df13e16fdfa1ae8610d3270101396c2e541ed5e5f27fc44d32793e9

Observation 2af9673b-755f-4fd6-ba7a-cd1e8b138c42 · outbound

This paper cites Geochat: Grounded large vision- language model for remote sensing,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Geochat: Grounded large vision- language model for remote sensing,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.488360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T19:26:44.879243Z digest=sha256:1a34683bbe8d823d86ea38fd42d36ddb81732843ce786eee3f60f310600a4da5

Observation e99ded90-b835-4b0e-adb9-5f79cfd646c5 · outbound

This paper cites Geollava: Efficient vision-language models for temporal change detection in remote sensing,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Geollava: Efficient vision-language models for temporal change detection in remote sensing,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.476126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T19:26:44.883227Z digest=sha256:1d7ca4b505a5c030d28c7e12e3ce05477cee642c254d96407a15d4160db239a5

Observation ca4e9245-c9bb-4459-b843-cc1052393104 · outbound

This paper cites SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features SkySenseGPT: A Fine-Grained Instruction Tuning Dataset and Model for Remote Sensing Vision-Language Understanding

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T19:26:44.887132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:26:44.887132Z digest=sha256:e197f7a3e2cbeaca3cfbe5a845bbbbb912d4f287404719d942e91c763178aea2

Observation fd2f19f5-4f7e-45d9-8907-80706657e817 · outbound

This paper cites reBEN: Refined BigEarthNet Dataset for Remote Sensing Image Analysis.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features reBEN: Refined BigEarthNet Dataset for Remote Sensing Image Analysis

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-10T19:26:44.891341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:26:44.891341Z digest=sha256:ec9df676f35433fe764eb1daa5e68caace7eac6a19a339c6349770fcf949d4d8

Observation 4863534a-e817-46a7-9594-825c2c60a759 · outbound

This paper cites First principles residual resistivity using locally self-consistent multiple scattering method.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features First principles residual resistivity using locally self-consistent multiple scattering method

Reference 12

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T19:26:45.085727Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T19:26:44.895190Z digest=sha256:8eb8b211c03d3e7f9d396cc658e60ef106c79693a45006fd69f5ba157619b0d1

Observation 58760e58-4351-44e1-a8ab-6582f79edd20 · outbound

This paper cites Spectral reconstruction from satellite multispectral imagery using convolution and transformer joint network,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Spectral reconstruction from satellite multispectral imagery using convolution and transformer joint network,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.463752Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T19:26:44.899023Z digest=sha256:cedd2601128d84efa3927a78303ed74b20f4455c0747ae47652fe0c89cc891ca

Observation 061acac1-f5ce-41fb-b4cb-fef4bbe9b770 · outbound

This paper cites Ringmo: A remote sensing foundation model with masked image modeling,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Ringmo: A remote sensing foundation model with masked image modeling,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.451307Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T19:26:44.902304Z digest=sha256:f26d5c3725d4bd6969bf344a643673a8d96aae754e25ef904d5a4a6b0483e0d9

Observation 2cbe90f8-8d2f-4d07-94e2-66cf27b0fd8a · outbound

This paper cites Masked autoencoders are scalable vision learners,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Masked autoencoders are scalable vision learners,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.438591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T19:26:44.906082Z digest=sha256:bc45cfb920f36f3ffbd6f153d0b55187e6498c85a0f10a94dc535a61dd75fb95

Observation bda0b774-b894-48e7-aaed-7e210aad771b · outbound

This paper cites Late-time Evolution and Instabilities of Tidal Disruption Disks.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Late-time Evolution and Instabilities of Tidal Disruption Disks

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-08-10T19:26:45.067195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T19:26:44.909901Z digest=sha256:c5ec2a1c92fd5d7de8c5dc8bfe139f8faa9db118bafe9d102b22b03bc50b756e

Observation 41ff7c79-c082-4b5c-9f4e-158b0e998fa9 · outbound

This paper cites Spectralgpt: Spectral remote sensing foundation model,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Spectralgpt: Spectral remote sensing foundation model,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.425343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T19:26:44.913939Z digest=sha256:258bca6d82b62360150dd4a4d6a8c22bfab202f8f99c879ecb15eda745bdbba0

Observation 21f93b25-e620-4a28-beb6-81179c3601da · outbound

This paper cites Rsgpt: A remote sensing vision-language model and benchmark,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Rsgpt: A remote sensing vision-language model and benchmark,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.412442Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T19:26:44.917729Z digest=sha256:0e5a62b2a6e907eadb5e1e9114a4cde1094106064130e6f6bae515883d1f658a

Observation 1937ae6e-5625-4d8a-b76b-53a8f2514bc6 · outbound

This paper cites Earthmarker: A visual prompting multi-modal large language model for remote sensing,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Earthmarker: A visual prompting multi-modal large language model for remote sensing,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.400090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T19:26:44.921340Z digest=sha256:687b33615238180e3596c7c126435338bb0ce7180c2765c5b44b7755290beef1

Observation afed98c1-59ec-4a35-877e-4be148390dd5 · outbound

This paper cites Combinatorics of generalized orthogonal polynomials of type $R_{II}$.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Combinatorics of generalized orthogonal polynomials of type $R_{II}$

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-10T19:26:44.925055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:26:44.925055Z digest=sha256:74401d5703ff3e3c2863ad58ec4bb899a70876d585cd7f8b89733fb8c49425e8

Observation fb283d58-5694-487e-bbb6-c1bfaa93ed9c · outbound

This paper cites LHRS-Bot-Nova: Improved Multimodal Large Language Model for Remote Sensing Vision-Language Interpretation.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features LHRS-Bot-Nova: Improved Multimodal Large Language Model for Remote Sensing Vision-Language Interpretation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-10T19:26:44.930836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:26:44.930836Z digest=sha256:d0bb6e10776086604635e13477ddd4c3dbba83fb738ef1c7f0f699f172f085ed

Observation 107ffc3c-f686-4b15-b1a2-65da33d38e76 · outbound

This paper cites GeoGround: A Unified Large Vision-Language Model for Remote Sensing Visual Grounding.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features GeoGround: A Unified Large Vision-Language Model for Remote Sensing Visual Grounding

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-10T19:26:44.935014Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:26:44.935014Z digest=sha256:41cc8454156cadf564b0fc4c7ecc477f8ed1030253b50ade7d089f36fe03a074

Observation d72f046e-e008-4199-937e-5e13bb5abb3f · outbound

This paper cites Skyeyegpt: Unifying remote sensing vision-language tasks via instruction tun- ing with large language model,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Skyeyegpt: Unifying remote sensing vision-language tasks via instruction tun- ing with large language model,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.281077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T19:26:44.939029Z digest=sha256:1dfc7c6b7faa5bcfccd7470095634349fc024a05062fc51495fd9745f10a37a6

Observation d51d1951-6a09-4315-942c-46bc30b80357 · outbound

This paper cites Ringmogpt: A unified remote sensing foundation model for vision, language, and grounded tasks,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Ringmogpt: A unified remote sensing foundation model for vision, language, and grounded tasks,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.266758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T19:26:44.942758Z digest=sha256:8c83b80f280a92ff455fff08c186c3d8bf18131e5e73757ee685142636bf2885

Observation f2d47d54-df15-40e7-9b89-6f41eef3f49c · outbound

This paper cites Teochat: A vision-language assistant for earth observation data,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Teochat: A vision-language assistant for earth observation data,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.252934Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T19:26:44.946669Z digest=sha256:18ec2ae44f2bd23b1bf62360537f04b5fe989cac531a03b3e07e1c11bb3c8e7b

Observation b576194a-f80a-4478-a2ef-8059ee3c040e · outbound

This paper cites VisionTrap: Vision-Augmented Trajectory Prediction Guided by Textual Descriptions.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features VisionTrap: Vision-Augmented Trajectory Prediction Guided by Textual Descriptions

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T19:26:44.950567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:26:44.950567Z digest=sha256:9665369daa51287507e905c1054ce7d0d30993d86f6fea9c549a0443f6d4fc19

Observation eeac5fb0-3a0c-4606-993f-0df4e692bf1e · outbound

This paper cites Functional map of the world,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Functional map of the world,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.239515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T19:26:44.954718Z digest=sha256:9ccb96f60308aa429fcbe96f8dd0e7e913e29bcd784c70579f98736ac5499d3b

Observation 5901d53d-5e2b-4278-b5c2-7c087f02401d · outbound

This paper cites BigEarthNet: A large-scale benchmark archive for re- mote sensing image understanding,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features BigEarthNet: A large-scale benchmark archive for re- mote sensing image understanding,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.227417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T19:26:44.958497Z digest=sha256:dcefecef3368949d10404cc6b031ed91ae494fa0dc69db14d326bf73283e7463

Observation 1d9f7c85-4c4c-4afb-8b45-e0d0ccc77b26 · outbound

This paper cites Distributionally Robust Receive Combining.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Distributionally Robust Receive Combining

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-10T19:26:44.962285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T19:26:44.962285Z digest=sha256:6050d556a648fef3c3d486255f5170871a3ce59e2ab274ea5e9831e019ad640f

Observation 8df9fe71-2f7c-403c-83cb-3bf198841104 · outbound

This paper cites Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification,.

A Vision-Language Framework for Multispectral Scene Representation Using Language-Grounded Features Eurosat: A novel dataset and deep learning benchmark for land use and land cover classification,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T19:26:45.214254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-10T19:26:44.966467Z digest=sha256:d0ebe4c823fe9c68cfb992e55fb26671a0833ae42f77682a4d08e9bacc60706e

Pith citing papers

No inbound Pith citation observations are available.