Pith. sign in

Paper Citation Record · LEDGER

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study

As of 10 August 2026, this Paper Citation Record lists 24 of 24 outbound references and 0 inbound Pith citation observations for arXiv:2508.20188.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.20188 v1

Coverage vector

measured 24 of 24 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-05T15:18:46.209624Z

measured 24 of 24 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

24 of 24 outbound references displayed

  • verified exact1
  • verified fuzzy14
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c91005ca-6987-4f4c-9641-6d12bc93fcf0 · outbound

This paper cites MedImageInsight: An Open-Source Embedding Model for General Domain Medical Imaging.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study MedImageInsight: An Open-Source Embedding Model for General Domain Medical Imaging

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:44.377415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:44.377415Z digest=sha256:8ca25c5765fd70e8c84bac7e2969d1e0886ca1eb86911441c9714f57efe36824

Observation b0b28e8d-3d76-440c-8f9e-27eeb4d521ec · outbound

This paper cites Validation of artificial intelligence prediction models for skin cancer diagnosis using dermoscopy images: the 2019 international skin imaging collaboration grand challenge.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Validation of artificial intelligence prediction models for skin cancer diagnosis using dermoscopy images: the 2019 international skin imaging collaboration grand challenge

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:49.080855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:18:44.439846Z digest=sha256:a39ab71cc1e9c11bf3470184c6b8ed154f8a12c76751543ab4c056ab64d33bf2

Observation e79d4228-dae4-4244-9455-8673d2e8745c · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study An image is worth 16x16 words: Transformers for image recognition at scale

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:44.544462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:44.544462Z digest=sha256:17d51dea2ce5068f711333d9214eb4df9804a85f55d0c7d9718d6ae5166f8d7d

Observation 5a16eecf-1504-447a-8fa7-1edfd031d9ce · outbound

This paper cites Analysis of trends in geographic distribution of us dermatology workforce density.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Analysis of trends in geographic distribution of us dermatology workforce density

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:48.831526Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:18:44.610713Z digest=sha256:db0870452334fcf85b00fe602ac64d083d7f2e623710a02e4029ee3f4d40382d

Observation 933d7108-5a32-45fa-a79d-ddecd706221b · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study LoRA: Low-rank adaptation of large language models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:44.684834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:44.684834Z digest=sha256:8f41b7d3984ed0dafbb92031a77a628aa2316b36ddfb559da2f4b3cc3f915eaa

Observation 1983dd57-45d5-400c-aedb-0c3c9b4f10e5 · outbound

This paper cites Transparent medical image ai via an image–text foundation model grounded in medical literature.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Transparent medical image ai via an image–text foundation model grounded in medical literature

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:48.588042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:18:44.760942Z digest=sha256:f2190e9277451d9c2c77063d261c2003cf52bcabf4b840f01ba3efc421b168a3

Observation 27d36465-a168-496c-8511-8e2f48e32d3b · outbound

This paper cites Human-ai interaction in skin cancer diagnosis: a systematic review and meta-analysis.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Human-ai interaction in skin cancer diagnosis: a systematic review and meta-analysis

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:48.428291Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:18:44.831224Z digest=sha256:25027d422bf75719dfdef2f145422200450b0c62d45422f4e9b59d3e44f5c201

Observation a3fee907-6528-46d1-a91f-5d3aac03d17a · outbound

This paper cites The slice-3d dataset: 400,000 skin lesion image crops extracted from 3d tbp for skin cancer detection.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study The slice-3d dataset: 400,000 skin lesion image crops extracted from 3d tbp for skin cancer detection

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:48.172329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:18:44.937867Z digest=sha256:fb3d7da1a7b84dc6eb9829399b7020226eef2486654ad9c8f8efe10c826800cd

Observation c931a68e-0be6-483a-b49d-4896618de5bf · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:45.014349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:45.014349Z digest=sha256:b6294e7fcaeba2fdb96bd364643d62d1e65cfde4f13fb2bdc2b22cafc8fcc269

Observation 8478ded3-3bee-4586-9f57-563c5a983952 · outbound

This paper cites 3d whole-body skin imaging for automated melanoma detection.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study 3d whole-body skin imaging for automated melanoma detection

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:48.022744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:18:45.098968Z digest=sha256:aca2bc7e2966270a4471cc9a434261e66f5af3c8e6832f611470112229d49b95

Observation fb5e098f-c316-43cf-8753-15d5ed2c38b1 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Learning transferable visual models from natural language supervision

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:47.894507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:18:45.198250Z digest=sha256:62207c1f5d9997ddab854c539dcbb8d0d05e7583f3fc579232cf389793795ad5

Observation e275ef87-4999-4481-954a-b740839a2537 · outbound

This paper cites Steering llama 2 via contrastive activation addition.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Steering llama 2 via contrastive activation addition

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:47.719743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:18:45.305119Z digest=sha256:20797d0d1c8d370929284614faecf85734fbcefa88ac0ea1951495c05398c663

Observation 17bb9f56-093f-4dcb-b880-cf22a6bfcc39 · outbound

This paper cites Grad-cam: Visual explanations from deep networks via gradient-based localization.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Grad-cam: Visual explanations from deep networks via gradient-based localization

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:45.416054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:45.416054Z digest=sha256:03b4fdd410e886c65c92e79297cbbd5551bf951f7956d477b4b7f455a16edc33

Observation d3e5c552-b236-4167-be21-716c29041771 · outbound

This paper cites Socioeconomic and geographic barriers to dermatology care in urban and rural us populations.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Socioeconomic and geographic barriers to dermatology care in urban and rural us populations

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:47.508759Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:18:45.489390Z digest=sha256:29e7db02f4b69b3cd79f44080a735cc0022ddb1126eff7decef841644029a053

Observation b55f806f-f8ff-45cb-b12e-1bb749c356bc · outbound

This paper cites Composing text and image for image retrieval-an empirical odyssey.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Composing text and image for image retrieval-an empirical odyssey

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:47.360891Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:18:45.558953Z digest=sha256:35903219e038539039319e652691fd79c5f09335c8d8db7d9a59f2ca1514a957

Observation 83d6baf8-9c44-4186-b411-733ac619071d · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:45.626372Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:45.626372Z digest=sha256:9ff82c0342f326f8db2aace0419b81cbb3ec618bf25c9c2f449813d1e1ac5028

Observation 49e3d848-5bca-4379-ba46-518088999649 · outbound

This paper cites VisNumBench: Evaluating Number Sense of Multimodal Large Language Models.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study VisNumBench: Evaluating Number Sense of Multimodal Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:45.699326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:45.699326Z digest=sha256:b298ba8e9517c9f01f8e938711059bb1d8688caacc3194099c73a06fc5e628c5

Observation 1b61aaa9-76cf-41ff-9db0-539acd98ffeb · outbound

This paper cites Derm1M: A Million-scale Vision-Language Dataset Aligned with Clinical Ontology Knowledge for Dermatology.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Derm1M: A Million-scale Vision-Language Dataset Aligned with Clinical Ontology Knowledge for Dermatology

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:45.769309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:45.769309Z digest=sha256:7c7ea0946a5c40691f82d1acb18f8630693c0320d24318a5c3d7d24cef020c50

Observation b776e1cc-ef69-4b29-8767-f9e37f939b3b · outbound

This paper cites A multimodal vision foundation model for clinical dermatology.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study A multimodal vision foundation model for clinical dermatology

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:47.196661Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:18:45.818436Z digest=sha256:3974eaa8652f1d039ef8a403794e36cadc08ed9eae57d618377f4ecbaa63b65c

Observation 6b2d822d-bf12-435f-bb7f-beb71bf363eb · outbound

This paper cites Qwen2 Technical Report.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Qwen2 Technical Report

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-05T15:18:45.884802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T15:18:45.884802Z digest=sha256:4e64d6f0bc8119c3f5e4ca57f3b3c34817378f0887e33e30996b64854491ed89

Observation 177cd5a3-b2db-4ad6-8b59-acf93633a55f · outbound

This paper cites Visrag: Vision-based retrieval-augmented generation on multi-modality documents.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Visrag: Vision-based retrieval-augmented generation on multi-modality documents

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:46.973569Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:18:45.992780Z digest=sha256:f2131b4d7daf68ca26fbbcf310d9a699971f4f53c303a54f53251807db2f6894

Observation 49f59181-aba5-4486-86a5-4641af924ae8 · outbound

This paper cites MM-Skin: Enhancing Dermatology Vision-Language Model with an Image-Text Dataset Derived from Textbooks.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study MM-Skin: Enhancing Dermatology Vision-Language Model with an Image-Text Dataset Derived from Textbooks

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-05T15:18:46.381344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:18:46.079905Z digest=sha256:640284551585548d86702b9041a0118ca4df27e4fed2f22cddcae702af33e480

Observation 2a98d987-e749-44ef-a093-65e209631c1c · outbound

This paper cites Revisiting the trustworthiness of saliency methods in radiology ai.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Revisiting the trustworthiness of saliency methods in radiology ai

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:46.821178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:18:46.140687Z digest=sha256:e7bd7d2692475eaea950788f06b399c036b1a923b47a687998183a41cad08905

Observation a19ca186-d62a-46ee-b73c-647881d6bf97 · outbound

This paper cites Pre-trained multimodal large language model enhances dermatological diagnosis using skingpt-4.

Grounding Multimodal Large Language Models with Quantitative Skin Attributes: A Retrieval Study Pre-trained multimodal large language model enhances dermatological diagnosis using skingpt-4

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-05T15:18:46.651647Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T15:18:46.209624Z digest=sha256:9794b6476e7fbaeec146ccb60916e6bc764a071e09647adce870fb073fa4f487

Pith citing papers

No inbound Pith citation observations are available.