Pith. sign in

Paper Citation Record · LEDGER

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text

As of 14 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 0 inbound Pith citation observations for arXiv:2508.00447.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.00447 v1

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T10:10:57.654052Z

measured 29 of 29 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

29 of 29 outbound references displayed

  • verified exact2
  • verified fuzzy18
  • unresolved9
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 149c51b0-aecb-49d5-a262-ce0a338be6f0 · outbound

This paper cites Large Language Models: A Survey.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Large Language Models: A Survey

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T10:10:57.566045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:10:57.566045Z digest=sha256:74d1e7cf5d9cccdc9e92a57b5157ff185ed37a583c33a3a9fd4969407af5ebd9

Observation dd2710f2-4dd6-469b-9396-5ac2e867eca4 · outbound

This paper cites Learning transferable visual models from natural language supervision.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Learning transferable visual models from natural language supervision

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-06T10:10:57.570247Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:10:57.570247Z digest=sha256:f6d5e09ac66db26071512972eaf6e329846fe54bcf68e4a1ac571b583d0b9dfd

Observation ab1a973f-5791-4f81-af60-63ef50eb0d9b · outbound

This paper cites Vlm agents generate their own memories: Distilling experience into embodied programs of thought.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Vlm agents generate their own memories: Distilling experience into embodied programs of thought

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.912864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T10:10:57.574095Z digest=sha256:a640434c898dcb94c3d20402f4cb71a3c6e625e071020177bdc12e6dbda8a1f9

Observation ba806b59-659a-4982-92a8-73a7bd038118 · outbound

This paper cites Rethinking vlms and llms for image classification.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Rethinking vlms and llms for image classification

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.904957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T10:10:57.577567Z digest=sha256:bd683a713aaf5fb7c54ffbab209d75da7b1aa03b01330f40d33b73e49bb91779

Observation 84c0b498-b7a4-40ae-893a-5076fbf49f16 · outbound

This paper cites Regression in EO: Are VLMs Up to the Challenge?.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Regression in EO: Are VLMs Up to the Challenge?

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T10:10:57.580373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:10:57.580373Z digest=sha256:9533240d31cfb89c07f5069713748b8870a22f3ac36b8bda116f7ebef51744be

Observation 37c8c514-ea71-4243-888f-106e2da05f68 · outbound

This paper cites From Images to Signals: Are Large Vision Models Useful for Time Series Analysis?.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text From Images to Signals: Are Large Vision Models Useful for Time Series Analysis?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T10:10:57.584389Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:10:57.584389Z digest=sha256:0ae6c4c0fdb2eadcc07aea533a4909db44089d0f1db830da5e97cddd2023c8c4

Observation 06d4ebcc-338b-4ed5-8fcd-de514be79caa · outbound

This paper cites Multi-step time series forecasting with an ensemble of varied length mixture models.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Multi-step time series forecasting with an ensemble of varied length mixture models

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.896346Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T10:10:57.588226Z digest=sha256:6fbb62bf32662923cb55b9c95e1b76d073d985bffc0c10a5c1401cb2bc2d9d6e

Observation 88a1cdb1-3f66-4c60-9aaf-442fd10cca4d · outbound

This paper cites CogVLM2: Visual Language Models for Image and Video Understanding.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text CogVLM2: Visual Language Models for Image and Video Understanding

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T10:10:57.591421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:10:57.591421Z digest=sha256:457aa978149c627f5c08d6cc17a1f99fae951d734e5d8ea51c473354d1cfdcd6

Observation fefd3627-07d8-4bf1-b645-0d9866c49b91 · outbound

This paper cites Integrating llms with its: Recent advances, potentials, challenges, and future directions.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Integrating llms with its: Recent advances, potentials, challenges, and future directions

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.889206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T10:10:57.594638Z digest=sha256:467ae1abac149e2cfea89c2f1f409a549b452b408605ec251af67d7925e4dad6

Observation ad8958b7-ae7b-4f20-9d7c-4532d0a30b87 · outbound

This paper cites iTransformer: Inverted Transformers Are Effective for Time Series Forecasting.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text iTransformer: Inverted Transformers Are Effective for Time Series Forecasting

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T10:10:57.597170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:10:57.597170Z digest=sha256:1bd4703b61e606fd008b7a888d02a3dccb40abac1fc7e9730266bab718b10896

Observation 85e09ceb-a398-446b-bfe7-f9155f22000c · outbound

This paper cites PV-VLM: A Multimodal Vision-Language Approach Incorporating Sky Images for Intra-Hour Photovoltaic Power Forecasting.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text PV-VLM: A Multimodal Vision-Language Approach Incorporating Sky Images for Intra-Hour Photovoltaic Power Forecasting

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-06T10:10:57.600421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:10:57.600421Z digest=sha256:708dba72f5c082e877f5d46d7318a827b9a186a6191e64531d8be96be1b8d444

Observation 708f493d-fa75-4949-879b-c5ed5b98eb1a · outbound

This paper cites Time-VLM: Exploring Multimodal Vision-Language Models for Augmented Time Series Forecasting.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Time-VLM: Exploring Multimodal Vision-Language Models for Augmented Time Series Forecasting

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T10:10:57.603925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:10:57.603925Z digest=sha256:5befeaed6b0183ed9c66e8b014d682e85cdd61c99fcdedbbcfae7c778cf7f1d5

Observation 9a3baa57-ccf5-416c-b4aa-8ffdbb9b11d3 · outbound

This paper cites Leveraging temporal contextualization for video action recognition.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Leveraging temporal contextualization for video action recognition

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.880538Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T10:10:57.607461Z digest=sha256:98abbbffb922a967e49aab0a5f9904c72dbf08d0b14c06adc8d845518d68f6cc

Observation 200ee217-6049-461d-97c3-20ac5323b43e · outbound

This paper cites C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature Dispersion.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text C-TPT: Calibrated Test-Time Prompt Tuning for Vision-Language Models via Text Feature Dispersion

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T10:10:57.611080Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T10:10:57.611080Z digest=sha256:4c0350048962747952d4fadd84b4124a98e001ee5a699b6d3f569878aafb162e

Observation 3af5bf5a-1e2a-4926-ac49-757828e85738 · outbound

This paper cites Pattern recognition and prediction in time series data through retrieval- augmented techniques.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Pattern recognition and prediction in time series data through retrieval- augmented techniques

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.871564Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T10:10:57.614083Z digest=sha256:5fdcf22064abd5f14d3552d19fdabd43afe83fcdc3ce030c5f47070bb2db5de5

Observation 30b41a4b-27b0-49a2-8b9e-55a1c9f526df · outbound

This paper cites Driver intention prediction using text prompts with in-cabin and out-cabin cameras.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Driver intention prediction using text prompts with in-cabin and out-cabin cameras

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.862058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T10:10:57.617359Z digest=sha256:7fe12fd203c4dfb7ece5c4716041d1e46d54e156291e1ebb175c9a1a01fdd2e5

Observation 02535a6d-e486-452a-b1d2-2551c8154a34 · outbound

This paper cites Diverse data augmentation with dif- fusions for effective test-time prompt tuning.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Diverse data augmentation with dif- fusions for effective test-time prompt tuning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.851547Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T10:10:57.620833Z digest=sha256:6a94eb925971abc949abc966a5e6b466f80afd15cb57ed119affc18d98b06554

Observation 2a3e5744-844b-4eea-a6ae-4b2e99768ebb · outbound

This paper cites FungalZSL: Zero-Shot Fungal Classification with Image Captioning Using a Synthetic Data Approach.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text FungalZSL: Zero-Shot Fungal Classification with Image Captioning Using a Synthetic Data Approach

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-06T10:10:57.692926Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T10:10:57.623645Z digest=sha256:6a1930e340ac2b0fa2c345f855f50722dcecaac3734e67ff7f06dd53204a5738

Observation e39d859c-8326-4a61-ad81-65683ad4de9a · outbound

This paper cites Fungal biology.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Fungal biology

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.842797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T10:10:57.626549Z digest=sha256:1858e4e8f0515458ff30afc7d015ed669448b5a14aefd98bbde0e42c41359105

Observation 76d98e21-6e5f-4a84-8f23-c50cecc9dc07 · outbound

This paper cites Assembling the fungal tree of life: progress, classification, and evolution of subcellular traits.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Assembling the fungal tree of life: progress, classification, and evolution of subcellular traits

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.834678Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T10:10:57.629001Z digest=sha256:62f38f4d2e2ba418b81c12a132be1d31ef3469182968ac6761f5315deb0cee8f

Observation 31ccd902-c5bc-4f0a-8d96-7db992310a9c · outbound

This paper cites Developments in fungal taxonomy.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Developments in fungal taxonomy

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.825328Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T10:10:57.631540Z digest=sha256:0729e8ca4f84d01b1e7af9630e19826ef97a9241f381cb503a779e175013e500

Observation 26edf3e3-5f98-40be-b0df-33f4c42e9440 · outbound

This paper cites Assessment of fungal spores and spore-like diversity in environmental samples by targeted lysis.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Assessment of fungal spores and spore-like diversity in environmental samples by targeted lysis

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.815366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T10:10:57.634084Z digest=sha256:24d8fd22f48e3947435e19939cc0e31e71689cda8e0b1f9983d9c3f5e2083d87

Observation 5e4e9dc0-eb84-4df6-b842-fdc7a8f03d09 · outbound

This paper cites Architecture and developmental dynamics of the external mycelium of the arbuscular mycorrhizal fungus glomus intraradices grown under monoxenic conditions.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Architecture and developmental dynamics of the external mycelium of the arbuscular mycorrhizal fungus glomus intraradices grown under monoxenic conditions

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.806844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T10:10:57.637069Z digest=sha256:e9a1535a05f487b3c0368115636ef0940f6c80112de5ff8775185c8529361664

Observation 85fc7581-3e42-4ef4-851f-6102411aadc6 · outbound

This paper cites Study of kinetic model for fungal spore germination under dynamic conditions: Case study on germination of penicillium expansum spores.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Study of kinetic model for fungal spore germination under dynamic conditions: Case study on germination of penicillium expansum spores

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.796630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T10:10:57.640618Z digest=sha256:02255b9e554aa61c0a12f8df53d3ccd14966ce73da1b04a8c80ffc05241198af

Observation 662c32db-1362-43e2-916a-dff97d1ddc09 · outbound

This paper cites Pathways of pathogenicity: transcriptional stages of germination in the fatal fungal pathogen rhizopus delemar.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Pathways of pathogenicity: transcriptional stages of germination in the fatal fungal pathogen rhizopus delemar

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.787085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T10:10:57.643506Z digest=sha256:d67ed2510e9ab916661654d5a625739f9a417e91e28925f7afd142f202d92a75

Observation 1f2ddbcc-5c21-465a-a143-ca725ba058e7 · outbound

This paper cites A model for growth of a single fungal hypha based on well-mixed tanks in series: simulation of nutrient and vesicle transport in aerial reproductive hyphae.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text A model for growth of a single fungal hypha based on well-mixed tanks in series: simulation of nutrient and vesicle transport in aerial reproductive hyphae

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.777381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T10:10:57.645926Z digest=sha256:c7d7331cecc935542bd96416730a2de0fb49f1cabd1c6700094e73afa34a95cd

Observation c75dbd80-d519-4699-9dc0-8a43782189e2 · outbound

This paper cites Neurospora crassa nadph oxidase nox-1 is localized in the vacuolar system and the plasma membrane.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Neurospora crassa nadph oxidase nox-1 is localized in the vacuolar system and the plasma membrane

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.766849Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T10:10:57.648290Z digest=sha256:b39a06f66515476f8b913bc1d5dffe775c26fbed76644c83f95cbe8ff18e56b7

Observation 1f1b6b55-5eea-4796-bb2b-4c01b29d5ef3 · outbound

This paper cites Synthetic Fungi Datasets: A Time-Aligned Approach.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Synthetic Fungi Datasets: A Time-Aligned Approach

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-08-06T10:10:57.680562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T10:10:57.650687Z digest=sha256:f8b5d340ec51236cf0707216a309ae85b53d8a5714a17ebb176680eaf715be6f

Observation 204e3eb2-1690-49f5-bbd4-d825d233d203 · outbound

This paper cites Fungalzsl: A fine-grained synthetic dataset for zero-shot fungi classification.

CLIPTime: Time-Aware Multimodal Representation Learning from Images and Text Fungalzsl: A fine-grained synthetic dataset for zero-shot fungi classification

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T10:10:57.756848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-06T10:10:57.654052Z digest=sha256:277cb50f828bcbd0f3c4120a1109e4fe6c9431300c03c4d1b4486af1f4cfbc4a

Pith citing papers

No inbound Pith citation observations are available.