Pith. sign in

Paper Citation Record · LEDGER

WISE: A Multimodal Search Engine for Visual Scenes, Audio, Objects, Faces, Speech, and Metadata

As of 10 August 2026, this Paper Citation Record lists 15 of 15 outbound references and 0 inbound Pith citation observations for arXiv:2602.12819.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2602.12819 v2

Coverage vector

measured 15 of 15 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-02T23:43:46.800172Z

measured 15 of 15 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

15 of 15 outbound references displayed

  • verified exact1
  • verified fuzzy0
  • unresolved14
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 021ae899-4baa-4bc0-9168-34af7c51db8c · outbound

This paper cites an unresolved cited work.

WISE: A Multimodal Search Engine for Visual Scenes, Audio, Objects, Faces, Speech, and Metadata Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-02T23:43:45.257162Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:43:45.257162Z digest=sha256:84583ea917b5d418984a83db69ef589112bc0ea9449cc14b91959f07a0fcbeb8

Observation 5b93d119-1f88-4297-90cf-08b8aaa71f56 · outbound

This paper cites an unresolved cited work.

WISE: A Multimodal Search Engine for Visual Scenes, Audio, Objects, Faces, Speech, and Metadata Unresolved cited work

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-02T23:43:45.388597Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:43:45.388597Z digest=sha256:4a9cea1eaad420761a6415309910090d9d8029ab2bb493b69b18e2143254d215

Observation df9131d5-5b57-499f-9163-72b70f80914c · outbound

This paper cites an unresolved cited work.

WISE: A Multimodal Search Engine for Visual Scenes, Audio, Objects, Faces, Speech, and Metadata Unresolved cited work

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-02T23:43:45.458802Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:43:45.458802Z digest=sha256:6a93ce3ca1459bcdf59c66a375ee45513bc53e93b43092854a87d853eac8fb38

Observation 689d7270-0652-4b89-bf99-361c09e2150e · outbound

This paper cites The Faiss library.

WISE: A Multimodal Search Engine for Visual Scenes, Audio, Objects, Faces, Speech, and Metadata The Faiss library

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-02T23:43:45.497421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:43:45.497421Z digest=sha256:f761b3b380b44e19c8d4e3894468801ea3bb58f956248c60ebfb040384576399

Observation 8e4d268f-3ea7-473b-87a3-69f3946a7e2e · outbound

This paper cites an unresolved cited work.

WISE: A Multimodal Search Engine for Visual Scenes, Audio, Objects, Faces, Speech, and Metadata Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T23:43:45.649786Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:43:45.649786Z digest=sha256:4ed6fc2d368b68c030ea12d69ccdd60f74819446f2d641524094214664285e1e

Observation fd9847a4-2feb-4f2e-86be-18f58c7510ae · outbound

This paper cites Richard Hipp.

WISE: A Multimodal Search Engine for Visual Scenes, Audio, Objects, Faces, Speech, and Metadata Richard Hipp

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-02T23:43:45.876623Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:43:45.876623Z digest=sha256:42558df1720f01d369200253d26fab43a047c4cc6d26b5f6ae3b8a38929d857c

Observation 1feb67a9-c110-4db3-980e-4962e5d6f67f · outbound

This paper cites an unresolved cited work.

WISE: A Multimodal Search Engine for Visual Scenes, Audio, Objects, Faces, Speech, and Metadata Unresolved cited work

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-02T23:43:45.982426Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:43:45.982426Z digest=sha256:daa9155d7910da3fa426a8e4f3439a80db7bcb8d4faf5ccfa265c79cce9ccb06

Observation 49fe1455-d905-454b-9c9e-8158a324db88 · outbound

This paper cites Self-supervised Learning on Camera Trap Footage Yields a Strong Universal Face Embedder.

WISE: A Multimodal Search Engine for Visual Scenes, Audio, Objects, Faces, Speech, and Metadata Self-supervised Learning on Camera Trap Footage Yields a Strong Universal Face Embedder

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-02T23:43:46.074733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:43:46.074733Z digest=sha256:2f28a3f2f7c5a8dc63b5b9735ea096a9b1e9aeafccd567a3996201ff986b271d

Observation 23416c91-db46-4a41-9d82-ebed4f11efa8 · outbound

This paper cites 2025.OpenCLIP.

WISE: A Multimodal Search Engine for Visual Scenes, Audio, Objects, Faces, Speech, and Metadata 2025.OpenCLIP

Reference 9

Resolution
verified exact
doi, observed 2026-08-02T23:49:46.777873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-02T23:43:46.179127Z digest=sha256:33ae7587df9924b22a58968fb51540b247a01d875876ee19ffeedf3a3e0ef71e

Observation e3f9bcf6-e127-4ffe-a449-bd3635ff35be · outbound

This paper cites an unresolved cited work.

WISE: A Multimodal Search Engine for Visual Scenes, Audio, Objects, Faces, Speech, and Metadata Unresolved cited work

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-02T23:43:46.261990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:43:46.261990Z digest=sha256:5fe3c9c9ce0346afc5cf5d023e50e15f3063d91229195add355c73614f5b0ce3

Observation e58d91f4-bb44-46db-8054-c68109bf47d5 · outbound

This paper cites an unresolved cited work.

WISE: A Multimodal Search Engine for Visual Scenes, Audio, Objects, Faces, Speech, and Metadata Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-02T23:43:46.410470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:43:46.410470Z digest=sha256:d13bb3766984aac3ae33acfbfec7cdee75dd4d1d3578386f613da9e156c80c94

Observation 60857ee3-2044-40dd-bd74-31249b1dbabd · outbound

This paper cites an unresolved cited work.

WISE: A Multimodal Search Engine for Visual Scenes, Audio, Objects, Faces, Speech, and Metadata Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-02T23:43:46.560074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:43:46.560074Z digest=sha256:f6fafbe60496369c85b8209d57afd4a254329da02285e0b9df6eef6be174e97c

Observation 7a57b258-a45c-46b5-a8eb-008a09d74a5f · outbound

This paper cites an unresolved cited work.

WISE: A Multimodal Search Engine for Visual Scenes, Audio, Objects, Faces, Speech, and Metadata Unresolved cited work

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-02T23:43:46.800172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:43:46.800172Z digest=sha256:55dfe40fbe08379fccc5fb7b8239ebf13505504cbaa78996cd2fb2d1904affd5

Observation b7a94a39-aeba-4c33-93b0-0822f1cc5c54 · outbound

This paper cites In International conference on machine learning.

WISE: A Multimodal Search Engine for Visual Scenes, Audio, Objects, Faces, Speech, and Metadata In International conference on machine learning

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-02T23:43:46.688077Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:43:46.688077Z digest=sha256:8890d9dc6d56eb1958d88d22cad42cb72982c2adc1eaf788ef15451de939a3f4

Observation 65c977d5-2f2e-429e-9424-d067a194053d · outbound

This paper cites InICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP).

WISE: A Multimodal Search Engine for Visual Scenes, Audio, Objects, Faces, Speech, and Metadata InICASSP 2023-2023 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-02T23:43:45.733030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T23:43:45.733030Z digest=sha256:7c0f0a6a7abc9a6802137f4e468dffd2f304a8f83bc85904293380f2d8f3d35b

Pith citing papers

No inbound Pith citation observations are available.