Pith. sign in

Paper Citation Record · LEDGER

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text

As of 9 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2508.00447.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.00447 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T10:10:57.654052Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact2
  • verified fuzzy18
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 149c51b0-aecb-49d5-a262-ce0a338be6f0 · outbound

This paper cites Large Language Models: A Survey.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Large Language Models: A Survey

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T10:10:57.566045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:10:57.566045Z digest=sha256:587860df7c65cf6964cce3b7cbe4fcb95a337255943847fcb0d246e9812db46e

Observation dd2710f2-4dd6-469b-9396-5ac2e867eca4 · outbound

This paper cites Learning transferable visual models from natural language supervision.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Learning transferable visual models from natural language supervision

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T10:10:57.570247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:10:57.570247Z digest=sha256:821711cf093306cb617a56cf27043d7e508cf8c7c604258c84102736e88d00ed

Observation ab1a973f-5791-4f81-af60-63ef50eb0d9b · outbound

This paper cites Vlm agents generate their own memories: Distilling experience into embodied programs of thought.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Vlm agents generate their own memories: Distilling experience into embodied programs of thought

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.912864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:10:57.574095Z digest=sha256:5056ef520ab73746c512503be29b19d50daea5564cf7dae815f2fec290b13386

Observation ba806b59-659a-4982-92a8-73a7bd038118 · outbound

This paper cites Rethinking vlms and llms for image classification.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Rethinking vlms and llms for image classification

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.904957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:10:57.577567Z digest=sha256:c87b77a89fb97c02999d54d2270467ac338a1033ba0d0b48a83840766bd329be

Observation 84c0b498-b7a4-40ae-893a-5076fbf49f16 · outbound

This paper cites Regression in EO: Are VLMs Up to the Challenge?.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Regression in EO: Are VLMs Up to the Challenge?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T10:10:57.580373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:10:57.580373Z digest=sha256:349d17e2a8b3f102e8e896dbba08c9370f0ef12b762708634d00acb2a40e61b8

Observation 37c8c514-ea71-4243-888f-106e2da05f68 · outbound

This paper cites From Images to Signals: Are Large Vision Models Useful for Time Series Analysis?.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text From Images to Signals: Are Large Vision Models Useful for Time Series Analysis?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T10:10:57.584389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:10:57.584389Z digest=sha256:4fd09e8d7745bc3d64448ea0f6fa44340cc1e70f9d84ad26d2ff791ba1607add

Observation 06d4ebcc-338b-4ed5-8fcd-de514be79caa · outbound

This paper cites Multi-step time series forecasting with an ensemble of varied length mixture models.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Multi-step time series forecasting with an ensemble of varied length mixture models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.896346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:10:57.588226Z digest=sha256:7a0e19d7bc54d98c4abe44c8bf01e4aa73e48348050f22c0b7585add77023856

Observation 88a1cdb1-3f66-4c60-9aaf-442fd10cca4d · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text CogVLM2: Visual Language Models for Image and Video Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T10:10:57.591421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:10:57.591421Z digest=sha256:834a91759b1f8f8c1fef90ea018c6c9a1c600dff0294d91703f3541eae8ef312

Observation fefd3627-07d8-4bf1-b645-0d9866c49b91 · outbound

This paper cites Integrating llms with its: Recent advances, potentials, challenges, and future directions.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Integrating llms with its: Recent advances, potentials, challenges, and future directions

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.889206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:10:57.594638Z digest=sha256:68d9ff1427c6c375435addad65408e84ac63a3f5f57af5aa36354be151810c34

Observation ad8958b7-ae7b-4f20-9d7c-4532d0a30b87 · outbound

This paper cites iTransformer: Inverted Transformers Are Effective for Time Series Forecasting.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text iTransformer: Inverted Transformers Are Effective for Time Series Forecasting

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T10:10:57.597170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:10:57.597170Z digest=sha256:3bd8953063a59428f3d297a17d789af0e19cbc564013223eeb86b14311d18da2

Observation 85e09ceb-a398-446b-bfe7-f9155f22000c · outbound

This paper cites PV-VLM: A Multimodal Vision-Language Approach Incorporating Sky Images for Intra-Hour Photovoltaic Power Forecasting.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text PV-VLM: A Multimodal Vision-Language Approach Incorporating Sky Images for Intra-Hour Photovoltaic Power Forecasting

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T10:10:57.600421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:10:57.600421Z digest=sha256:c785f9295f1c5d54346466f95769b3422818531585a05b2cb2efd7db0251556c

Observation 708f493d-fa75-4949-879b-c5ed5b98eb1a · outbound

This paper cites Time-VLM: Exploring Multimodal Vision-Language Models for Augmented Time Series Forecasting.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Time-VLM: Exploring Multimodal Vision-Language Models for Augmented Time Series Forecasting

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T10:10:57.603925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:10:57.603925Z digest=sha256:3b2014a74ef25db65d477ac9dbd85dba246e36e5c5b0a98d3a6bc39e5354d53a

Observation 9a3baa57-ccf5-416c-b4aa-8ffdbb9b11d3 · outbound

This paper cites Leveraging temporal contextualization for video action recognition.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Leveraging temporal contextualization for video action recognition

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.880538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:10:57.607461Z digest=sha256:31f3058a56fe8b162d4bc49678c102979ca7c3cdd697f20a70b95ee2def21ba6

Observation 200ee217-6049-461d-97c3-20ac5323b43e · outbound

This paper cites C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature Dispersion.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature Dispersion

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T10:10:57.611080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:10:57.611080Z digest=sha256:2ed9a226276a6e9cea5d83d85caf4321a0332854021cc67c7502991e971f921a

Observation 3af5bf5a-1e2a-4926-ac49-757828e85738 · outbound

This paper cites Pattern recognition and prediction in time series data through retrieval- augmented techniques.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Pattern recognition and prediction in time series data through retrieval- augmented techniques

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.871564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:10:57.614083Z digest=sha256:f071439846e63d91e1301b48ff9da2111df6efa2b3a2a60bba656bcdaebf2d98

Observation 30b41a4b-27b0-49a2-8b9e-55a1c9f526df · outbound

This paper cites Driver intention prediction using text prompts with in-cabin and out-cabin cameras.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Driver intention prediction using text prompts with in-cabin and out-cabin cameras

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.862058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:10:57.617359Z digest=sha256:b45b884a38c86341753edcd8c1db2d7b7f929a2f9721012696e6957e725a25c1

Observation 02535a6d-e486-452a-b1d2-2551c8154a34 · outbound

This paper cites Diverse data augmentation with dif- fusions for effective test-time prompt tuning.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Diverse data augmentation with dif- fusions for effective test-time prompt tuning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.851547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:10:57.620833Z digest=sha256:5e1ad111580b736083a409c3b6725f5b650f3bdf11a298f1fe57940cbaece877

Observation 2a3e5744-844b-4eea-a6ae-4b2e99768ebb · outbound

This paper cites FungalZSL: Zero-Shot Fungal Classification with Image Captioning Using a Synthetic Data Approach.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text FungalZSL: Zero-Shot Fungal Classification with Image Captioning Using a Synthetic Data Approach

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-06T10:10:57.692926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:10:57.623645Z digest=sha256:18239bf7c7f973b98140dd2f32a5cfe10cdb0230f0b9c74989b8c886ea7a284a

Observation e39d859c-8326-4a61-ad81-65683ad4de9a · outbound

This paper cites Fungal biology.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Fungal biology

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.842797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:10:57.626549Z digest=sha256:49f27b57e5f23974aff12a1dcbd4cd07df78b3c518cf8f8ae1d04b0eb56ef3a1

Observation 76d98e21-6e5f-4a84-8f23-c50cecc9dc07 · outbound

This paper cites Assembling the fungal tree of life: progress, classification, and evolution of subcellular traits.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Assembling the fungal tree of life: progress, classification, and evolution of subcellular traits

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.834678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:10:57.629001Z digest=sha256:05868efd4bc171f931061ecdfabb421a114797babd9627ad4d7c20f6237fe37f

Observation 31ccd902-c5bc-4f0a-8d96-7db992310a9c · outbound

This paper cites Developments in fungal taxonomy.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Developments in fungal taxonomy

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.825328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:10:57.631540Z digest=sha256:af12f9532eb2872804c3528a3993346c7d3aa94d8e621ec3050f1ccb7666741c

Observation 26edf3e3-5f98-40be-b0df-33f4c42e9440 · outbound

This paper cites Assessment of fungal spores and spore-like diversity in environmental samples by targeted lysis.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Assessment of fungal spores and spore-like diversity in environmental samples by targeted lysis

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.815366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:10:57.634084Z digest=sha256:9141e43e62233fe1ca3c4c7bc81726c5937d2437cfc2ac8c65edffc88251966d

Observation 5e4e9dc0-eb84-4df6-b842-fdc7a8f03d09 · outbound

This paper cites Architecture and developmental dynamics of the external mycelium of the arbuscular mycorrhizal fungus glomus intraradices grown under monoxenic conditions.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Architecture and developmental dynamics of the external mycelium of the arbuscular mycorrhizal fungus glomus intraradices grown under monoxenic conditions

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.806844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:10:57.637069Z digest=sha256:2014999af3438937d8d8da901be18fd76b9e7e1cc4dec15c668672c55d06b99c

Observation 85fc7581-3e42-4ef4-851f-6102411aadc6 · outbound

This paper cites Study of kinetic model for fungal spore germination under dynamic conditions: Case study on germination of penicillium expansum spores.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Study of kinetic model for fungal spore germination under dynamic conditions: Case study on germination of penicillium expansum spores

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.796630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:10:57.640618Z digest=sha256:16065cf5fa93c8832e3e42cb085231f0b885b293525502524040dc3d797f353c

Observation 662c32db-1362-43e2-916a-dff97d1ddc09 · outbound

This paper cites Pathways of pathogenicity: transcriptional stages of germination in the fatal fungal pathogen rhizopus delemar.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Pathways of pathogenicity: transcriptional stages of germination in the fatal fungal pathogen rhizopus delemar

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.787085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:10:57.643506Z digest=sha256:92927f20f4cdd67e43680236d1a292d6441d3dc318924d54bb5ea5299b4d9770

Observation 1f2ddbcc-5c21-465a-a143-ca725ba058e7 · outbound

This paper cites A model for growth of a single fungal hypha based on well-mixed tanks in series: simulation of nutrient and vesicle transport in aerial reproductive hyphae.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text A model for growth of a single fungal hypha based on well-mixed tanks in series: simulation of nutrient and vesicle transport in aerial reproductive hyphae

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.777381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:10:57.645926Z digest=sha256:226cf95f8f5a1205c7b7effb966723d62a7042ad601ae01d01da9fd11a198456

Observation c75dbd80-d519-4699-9dc0-8a43782189e2 · outbound

This paper cites Neurospora crassa nadph oxidase nox-1 is localized in the vacuolar system and the plasma membrane.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Neurospora crassa nadph oxidase nox-1 is localized in the vacuolar system and the plasma membrane

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.766849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:10:57.648290Z digest=sha256:34f8fa663dab16d1bfc526cede5776e68bce281f66de079a26d1eb723eb453ec

Observation 1f1b6b55-5eea-4796-bb2b-4c01b29d5ef3 · outbound

This paper cites Synthetic Fungi Datasets: A Time-Aligned Approach.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Synthetic Fungi Datasets: A Time-Aligned Approach

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-06T10:10:57.680562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:10:57.650687Z digest=sha256:c6f2ac4dab1196b1e0a5ed21e3f8f6efc513075d8771c5cf64c9a1efe316f101

Observation 204e3eb2-1690-49f5-bbd4-d825d233d203 · outbound

This paper cites Fungalzsl: A fine-grained synthetic dataset for zero-shot fungi classification.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Fungalzsl: A fine-grained synthetic dataset for zero-shot fungi classification

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.756848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T10:10:57.654052Z digest=sha256:28ad5cb2040cc29167964649e764c21996dec668c493296e2b9389fb67e6bac3

Pith citing papers

No inbound Pith citation observations are available.