Pith. sign in

Paper Citation Record · LEDGER

Open-set Cross Modal Generalization via Multimodal Unified Representation

As of 9 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 1 inbound Pith citation observation for arXiv:2507.14935.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.14935 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:48:33.963779Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:53:28.021948Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T21:53:34.084878Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact5
  • verified fuzzy45
  • unresolved14
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9a653eda-7ed8-461e-9c05-254125e640ed · outbound

This paper cites Robust cross-modal representation learning with progressive self- distillation.

Open-set Cross Modal Generalization via Multimodal Unified Representation Robust cross-modal representation learning with progressive self- distillation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:46.945276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:25.548996Z digest=sha256:44abb25c046f4e3de3acb15537316cba2ac157089efc2096cee20da453f0e204

Observation da722cd7-14fe-4fba-9270-de4615f6ea92 · outbound

This paper cites On the effectiveness of image rotation for open set domain adaptation.

Open-set Cross Modal Generalization via Multimodal Unified Representation On the effectiveness of image rotation for open set domain adaptation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:46.699431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:25.670170Z digest=sha256:4ad659e665979ea2fe5bbde95dec3260e6f46b47ca8950be1091c522d9010438

Observation f27956e9-7a49-4e87-ac8a-f86e3d175bdf · outbound

This paper cites Domain generalization by solving jigsaw puzzles.

Open-set Cross Modal Generalization via Multimodal Unified Representation Domain generalization by solving jigsaw puzzles

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:46.448429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:25.778914Z digest=sha256:5b35d97cdf7f79b194ee111ec61f9e5dfa52bba01c644c4335f4989f2abfc3ac

Observation dee88b25-f9c6-4d64-b7b1-a9ce145e3d7c · outbound

This paper cites Collecting highly paral- lel data for paraphrase evaluation.

Open-set Cross Modal Generalization via Multimodal Unified Representation Collecting highly paral- lel data for paraphrase evaluation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:25.886544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:25.886544Z digest=sha256:54607dc6512133b10ae8dd1d084e9ca876e397afacd73b6597adbde358c00d52

Observation 396e61a4-d780-417d-a576-1fbeb402b0c1 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

Open-set Cross Modal Generalization via Multimodal Unified Representation Vggsound: A large-scale audio-visual dataset

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:46.213437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:26.060304Z digest=sha256:273f1155ff576aabaf70320374d367769075a376fc88251df69b3daa2ca58cf0

Observation 40a90f4a-be71-4554-a7b7-24b51ec6dc28 · outbound

This paper cites Hts-at: A hierarchi- cal token-semantic audio transformer for sound classifica- tion and detection.

Open-set Cross Modal Generalization via Multimodal Unified Representation Hts-at: A hierarchi- cal token-semantic audio transformer for sound classifica- tion and detection

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:45.924381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:26.165853Z digest=sha256:a8373b4b19cd53f70aafe9f71052c04c9499d6be24c2df74951f9c5237a26fa7

Observation 9eacc5aa-1288-437f-a5f3-2e665af6c317 · outbound

This paper cites VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset.

Open-set Cross Modal Generalization via Multimodal Unified Representation VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:26.258804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:26.258804Z digest=sha256:61b608c5716664658b91abe42fdf54f16cdaf909465926e22322ab2aee47cbd9

Observation 4528c55e-2a93-4781-ae74-7c72c085ea0e · outbound

This paper cites Uniter: Universal image-text representation learning.

Open-set Cross Modal Generalization via Multimodal Unified Representation Uniter: Universal image-text representation learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:45.722400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:26.393490Z digest=sha256:9533e1b430e8de9ddc27b55a5ec16bf7c8b5265cf536212138400015d2e4c562

Observation 7406c049-2539-466b-aa97-8122938036f1 · outbound

This paper cites Sinkd: Sinkhorn distance minimization for knowledge distillation.

Open-set Cross Modal Generalization via Multimodal Unified Representation Sinkd: Sinkhorn distance minimization for knowledge distillation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:45.497935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:26.524128Z digest=sha256:34162bf6f476c4fd123aae99dd09d0cd166e130c2069e0e98a6355922a03cf89

Observation d694a12a-1115-452f-949a-aca3c1e76ad3 · outbound

This paper cites Sinkhorn distance minimization for knowledge distilla- tion.

Open-set Cross Modal Generalization via Multimodal Unified Representation Sinkhorn distance minimization for knowledge distilla- tion

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:45.205864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:26.655853Z digest=sha256:6dd2ff87af05e853996ad18fb1274bf72ab276ff516ae25eda72c0fda49aa700

Observation 1b229b04-2c2d-4e19-a83e-327b027cdebf · outbound

This paper cites Optical: Leveraging optimal trans- port for contribution allocation in dataset distillation.

Open-set Cross Modal Generalization via Multimodal Unified Representation Optical: Leveraging optimal trans- port for contribution allocation in dataset distillation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:44.922516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:26.771225Z digest=sha256:936cc973a7c946e46157badcee9758b764b80be12a2e34fe18878ff9cf5491a0

Observation 2f994dd1-b7c2-4e6d-8938-a2e94fdcde70 · outbound

This paper cites Layoutenc: Leveraging enhanced layout rep- resentations for transformer-based complex scene synthesis.

Open-set Cross Modal Generalization via Multimodal Unified Representation Layoutenc: Leveraging enhanced layout rep- resentations for transformer-based complex scene synthesis

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:44.599120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:26.885590Z digest=sha256:9778e17e283c5320ebbb800488f7a77c8eb2f89b5c665fa8e0bd814be41bc57d

Observation a3cf246e-ee9e-41e9-9ee8-07a0e685898f · outbound

This paper cites Streetsurfgs: Scalable ur- ban street surface reconstruction with planar-based gaussian splatting.

Open-set Cross Modal Generalization via Multimodal Unified Representation Streetsurfgs: Scalable ur- ban street surface reconstruction with planar-based gaussian splatting

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:44.384755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:27.037175Z digest=sha256:0fdcf9f3d2de6a6b1abbbda465332042e585b598abc5e87d98127ceec432d8fe

Observation af806a17-5526-4171-8036-6b1e99c5917e · outbound

This paper cites Towards Multimodal Open-Set Domain Generalization and Adaptation through Self-supervision.

Open-set Cross Modal Generalization via Multimodal Unified Representation Towards Multimodal Open-Set Domain Generalization and Adaptation through Self-supervision

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:48:35.385262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:27.140292Z digest=sha256:247a8ec8f348708cdd23dcc1a077c3bab97736d399b5f98b088daecb8bf7276b

Observation 33e87a74-1eb6-464c-b70a-86a9749d9f10 · outbound

This paper cites Simmmdg: A simple and effective framework for multi-modal domain generalization.

Open-set Cross Modal Generalization via Multimodal Unified Representation Simmmdg: A simple and effective framework for multi-modal domain generalization

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:44.159885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:27.241584Z digest=sha256:4f7c21c0f2e34e1113c5a57a4f23946affb55de82640dd939f85e6e0e17297bf

Observation 507023a4-0327-44b5-b8ea-a620ee805697 · outbound

This paper cites Clotho: An audio captioning dataset.

Open-set Cross Modal Generalization via Multimodal Unified Representation Clotho: An audio captioning dataset

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:43.936791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:27.379391Z digest=sha256:206a5bb71cd7f87e94e3289657aa4921f0e89256d29b4b560249f6d5f99c5459

Observation 09215f9d-9f71-494e-b01f-299865375760 · outbound

This paper cites Multi-modal align- ment using representation codebook.

Open-set Cross Modal Generalization via Multimodal Unified Representation Multi-modal align- ment using representation codebook

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:43.707248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:27.511828Z digest=sha256:de2b66ee29328c3416825b1f0f161d32f06656bc56bfdd6a9ac41a6945d88bfb

Observation 51531b0d-6a4b-410d-be22-2e8426a6893c · outbound

This paper cites Ace: A generative cross-modal retrieval framework with coarse-to-fine semantic modeling.

Open-set Cross Modal Generalization via Multimodal Unified Representation Ace: A generative cross-modal retrieval framework with coarse-to-fine semantic modeling

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-06T15:48:35.160122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:27.648562Z digest=sha256:0c507b2b15d18ab32c5a7d046eaf8a843cfd7a8a5d1031b6b94d71ad18152ee7

Observation 2f814264-7211-4be5-a0d5-027e32162906 · outbound

This paper cites Slowfast networks for video recognition.

Open-set Cross Modal Generalization via Multimodal Unified Representation Slowfast networks for video recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:27.779563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:27.779563Z digest=sha256:11712675af2c676405ebe8c1e52f5f6c15b6927da42609e5ba4951e780efe245

Observation fd407d55-5386-4d64-9c20-ccaf26e2e10f · outbound

This paper cites Domain-adversarial training of neural networks.

Open-set Cross Modal Generalization via Multimodal Unified Representation Domain-adversarial training of neural networks

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:43.460597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:27.947964Z digest=sha256:62fbad5bc12f239ff6d542eb9547e869daaef5aae3f048375507f603155a4149

Observation 4c64e29f-8989-4401-991e-dc5f1e2c3740 · outbound

This paper cites Imagebind: One embedding space to bind them all.

Open-set Cross Modal Generalization via Multimodal Unified Representation Imagebind: One embedding space to bind them all

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:28.115661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:28.115661Z digest=sha256:5d5b87b10230195a57cd04605fed39fe411ecf72bdef859e153c80dd48125c8f

Observation 5693eb52-ffb9-4a0e-a0b6-c3fd604f14b2 · outbound

This paper cites Enhancing Multimodal Unified Representations for Cross Modal Generalization.

Open-set Cross Modal Generalization via Multimodal Unified Representation Enhancing Multimodal Unified Representations for Cross Modal Generalization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:28.218101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:28.218101Z digest=sha256:e3f56e0a98e7760018e445ef58daff39fa97b31c98f050777c161e9ee2717ac2

Observation 302450ac-17aa-42a6-8c26-e1ed2f82c0bd · outbound

This paper cites Semantic residual for multimodal unified discrete representation.

Open-set Cross Modal Generalization via Multimodal Unified Representation Semantic residual for multimodal unified discrete representation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:43.239672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:28.320120Z digest=sha256:d967016816b992bd89710d99949d53b0fe2cd342bce8c2660a169cee3a5e4fae

Observation bdc993b6-96dd-4fdc-b61a-fd7b27d75e34 · outbound

This paper cites Bridging Domain Generalization to Multimodal Domain Generalization via Unified Representations.

Open-set Cross Modal Generalization via Multimodal Unified Representation Bridging Domain Generalization to Multimodal Domain Generalization via Unified Representations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:28.447490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:28.447490Z digest=sha256:4245e375d184fca8fdcdaf403c39a0a9b219e2bba22e23da359e52c98309b5ee

Observation b0f93d54-2536-44a3-91c0-881c016cfe4f · outbound

This paper cites WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling.

Open-set Cross Modal Generalization via Multimodal Unified Representation WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:28.600859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:28.600859Z digest=sha256:605e8e771b52e16fb8a077f37f702a46811384a26d3507db8711868db101198d

Observation 807fa901-b252-4444-8f09-06bbc1d8feb9 · outbound

This paper cites Learning to generalize: Meta-learning for do- main generalization.

Open-set Cross Modal Generalization via Multimodal Unified Representation Learning to generalize: Meta-learning for do- main generalization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:42.994612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:28.728539Z digest=sha256:456f94f62171cfc94c63dd9f999fd90f2e18f8834091ee165ffa5760db2ff916

Observation e4e07c83-b58c-4ac9-9336-480b1ebb98c7 · outbound

This paper cites Domain generalization for med- ical imaging classification with linear-dependency regular- ization.

Open-set Cross Modal Generalization via Multimodal Unified Representation Domain generalization for med- ical imaging classification with linear-dependency regular- ization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:42.783890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:28.883688Z digest=sha256:d8e134a4c3fe96a7cc946cc95738092e1cce9b98238b58367c80b7cbc5811c74

Observation 882ef81a-0a97-43bc-8971-19b5b257d701 · outbound

This paper cites Adjustment and alignment for unbiased open set domain adaptation.

Open-set Cross Modal Generalization via Multimodal Unified Representation Adjustment and alignment for unbiased open set domain adaptation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:42.499297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:29.019728Z digest=sha256:c8f9675c49607d8ade1e9e959ecc51f2e704613d01329c7e40ba207722b5481f

Observation b9e429f1-b3ba-46b4-82a2-594c73c93eff · outbound

This paper cites Cross-Modal Discrete Representation Learning.

Open-set Cross Modal Generalization via Multimodal Unified Representation Cross-Modal Discrete Representation Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:29.139488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:29.139488Z digest=sha256:996b51e58efa9054d2eaf828722b10d4d91d986e2758413cb5d61cf553c5caad

Observation 1c7a639d-2e41-4ca0-b4a7-ed46348417e7 · outbound

This paper cites Feddg: Federated domain generalization on medical image segmentation via episodic learning in continuous fre- quency space.

Open-set Cross Modal Generalization via Multimodal Unified Representation Feddg: Federated domain generalization on medical image segmentation via episodic learning in continuous fre- quency space

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:42.266176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:29.274536Z digest=sha256:ea0976b75c1145dcf7b5986f38802d7a1937deb43ddc7428f0cd18b654fe5259

Observation a53e4ee0-723a-424f-9877-952fddd56894 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Open-set Cross Modal Generalization via Multimodal Unified Representation Swin transformer: Hierarchical vision transformer using shifted windows

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:29.435332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:29.435332Z digest=sha256:b5436390c1f3f9f24255e0f0293eb472f8efeb4b1e0a855cc4e0eba6607202bc

Observation 3d48944d-23b5-4973-ace7-88db099cecdf · outbound

This paper cites Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks.

Open-set Cross Modal Generalization via Multimodal Unified Representation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:29.595479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:29.595479Z digest=sha256:64a996e18bb8ecda4619c680934b72f38b921e2910759178fd8668d304227045

Observation 604260d1-9db9-4e96-9617-9735f183c9ad · outbound

This paper cites Unsupervised learning of visual representations by solving jigsaw puzzles.

Open-set Cross Modal Generalization via Multimodal Unified Representation Unsupervised learning of visual representations by solving jigsaw puzzles

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:41.956515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:29.699219Z digest=sha256:22732644c87a419806583cff93134c94d2fb5e8a8f46c3caf9f08b87cd44357f

Observation a19c92f6-8dc0-4a89-b56e-f1530b4f90c6 · outbound

This paper cites Two at once: Enhancing learning and generalization capacities via ibn-net.

Open-set Cross Modal Generalization via Multimodal Unified Representation Two at once: Enhancing learning and generalization capacities via ibn-net

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:41.672767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:29.830914Z digest=sha256:1a9393967c7216511d40db96ee565d0226353e4651920045ac46a2b2a073216b

Observation 0a5add89-0bdf-4104-b171-e11c1a236067 · outbound

This paper cites Estimating Visual Information From Audio Through Manifold Learning.

Open-set Cross Modal Generalization via Multimodal Unified Representation Estimating Visual Information From Audio Through Manifold Learning

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:48:34.769435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:29.953749Z digest=sha256:3ef927a28c5f3c30c14b7a312819bb0919d1c2681df45b9f012ee3eb4b5b2e01

Observation b771fb0b-8b4f-4812-8c64-7d035d63b104 · outbound

This paper cites Audio-visual speech recognition with a hybrid ctc/attention architecture.

Open-set Cross Modal Generalization via Multimodal Unified Representation Audio-visual speech recognition with a hybrid ctc/attention architecture

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:41.464962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:30.116086Z digest=sha256:aeec2238e88cadf91f6c2100bbf19f8a16699f1c2fc6a54e6be062ea06f35882

Observation 0a61c64f-76dc-42b1-930c-c750e10ea19f · outbound

This paper cites Domain generalization through audio- visual relative norm alignment in first person action recog- nition.

Open-set Cross Modal Generalization via Multimodal Unified Representation Domain generalization through audio- visual relative norm alignment in first person action recog- nition

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:41.216157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:30.254045Z digest=sha256:655ea396e1412b3d7a328dad24b4fed3dc4dad7c7bf3c50664938968ed68a2dc

Observation eb751c58-b667-4e10-8fff-866a3022576c · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Open-set Cross Modal Generalization via Multimodal Unified Representation Learning transferable visual models from natural language supervi- sion

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:30.385147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:30.385147Z digest=sha256:53975a44956afc21f95b46a19ce958733e4f09a0b940c6491b2725a88b4caa91

Observation 0f742929-52d9-43e4-a3c9-5c0dc224b5e2 · outbound

This paper cites Mask2anomaly: Mask transformer for uni- versal open-set segmentation.

Open-set Cross Modal Generalization via Multimodal Unified Representation Mask2anomaly: Mask transformer for uni- versal open-set segmentation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:40.949706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:30.518039Z digest=sha256:7b82a6b6abd80bd26ab148e13eb4ac526b452d0b2c7a24707f878eacdd9283dd

Observation 62bdf01a-5e88-4d31-b262-8bf4980429fa · outbound

This paper cites XKD: Cross-modal Knowledge Distillation with Domain Alignment for Video Representation Learning.

Open-set Cross Modal Generalization via Multimodal Unified Representation XKD: Cross-modal Knowledge Distillation with Domain Alignment for Video Representation Learning

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:48:34.487611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:30.648250Z digest=sha256:470f596064b2e904b412e140e60c075de25c42bf15a51a263574b2740578042a

Observation 24c727c0-79b1-4434-95d5-26697d9abdf6 · outbound

This paper cites Open domain generalization with domain- augmented meta-learning.

Open-set Cross Modal Generalization via Multimodal Unified Representation Open domain generalization with domain- augmented meta-learning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:40.701072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:30.804084Z digest=sha256:1af105a1180a4c94ee5c8a821a24eeb9a5a8e7829495981712f4cd16aceb24fd

Observation 7281a904-b633-496b-bf6b-43836baae6e1 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Open-set Cross Modal Generalization via Multimodal Unified Representation UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:30.917220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:30.917220Z digest=sha256:113e814758c6e7152426ae037ab1c326ce1ce4c56d9865977d96d4783ce0e13f

Observation ef32786e-a737-49b6-8428-03a2dcfd6dff · outbound

This paper cites Audio-visual event localization in unconstrained videos.

Open-set Cross Modal Generalization via Multimodal Unified Representation Audio-visual event localization in unconstrained videos

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:40.399365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:31.035641Z digest=sha256:a90d4a1c2d903b3d87d8d9314933da570646e81399151d08336a74b6d2c471be

Observation 9651d3e8-bc5d-4b0a-bc27-b1d713ae757f · outbound

This paper cites Unified mul- tisensory perception: Weakly-supervised audio-visual video parsing.

Open-set Cross Modal Generalization via Multimodal Unified Representation Unified mul- tisensory perception: Weakly-supervised audio-visual video parsing

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:40.163541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:31.148002Z digest=sha256:4cac5d1de8b60fdd40b5e6af65cd5a111153a36647b1b75bb34c3403a848abd5

Observation 5de95d9f-85c0-426e-9575-b89836e6821f · outbound

This paper cites Domain randomization for transferring deep neural networks from simulation to the real world.

Open-set Cross Modal Generalization via Multimodal Unified Representation Domain randomization for transferring deep neural networks from simulation to the real world

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:39.944179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:31.310142Z digest=sha256:883c521e39f2a37d39e72ad4464216dd6f6463d76c0653f519c5ecc45f70f1d5

Observation ec0d5cd5-d07d-4e68-a099-f64740d5f903 · outbound

This paper cites Deep Domain Confusion: Maximizing for Domain Invariance.

Open-set Cross Modal Generalization via Multimodal Unified Representation Deep Domain Confusion: Maximizing for Domain Invariance

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:31.427887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:31.427887Z digest=sha256:0bf751ef44f0e14d9b057eca8c9ed1b756fecd43d9d0e776d67d4adffb8902e7

Observation 4f23e8e0-ffe8-4fd5-b1ec-ea7f4cc4fc3e · outbound

This paper cites IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models.

Open-set Cross Modal Generalization via Multimodal Unified Representation IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:31.574133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:31.574133Z digest=sha256:5a6ee1b65c5ccf1da10e19527cb8e18f9dbe1801340211061b37dd1511cb1656

Observation 02e59edd-c139-4d7d-a293-980bf430387a · outbound

This paper cites Towards Transformer-Based Aligned Generation with Self-Coherence Guidance.

Open-set Cross Modal Generalization via Multimodal Unified Representation Towards Transformer-Based Aligned Generation with Self-Coherence Guidance

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:48:34.215630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:31.694527Z digest=sha256:314a28ca2ef712878103776f59b59046f71446df0a0cfa8fdb82609f1ff7b7f6

Observation 449e3041-9ccf-43fd-b742-3b06a1d220cb · outbound

This paper cites Vlmixer: Unpaired vision-language pre-training via cross-modal cutmix.

Open-set Cross Modal Generalization via Multimodal Unified Representation Vlmixer: Unpaired vision-language pre-training via cross-modal cutmix

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:39.699412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:31.802571Z digest=sha256:b3b485922bfd9edd04ae259b978d6dc40bf7215b855e0864e0f6b546e509d328

Observation 4d17012a-086b-478e-98a8-7edd453a268c · outbound

This paper cites General- izable decision boundaries: Dualistic meta-learning for open set domain generalization.

Open-set Cross Modal Generalization via Multimodal Unified Representation General- izable decision boundaries: Dualistic meta-learning for open set domain generalization

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:39.416582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:31.971663Z digest=sha256:c58d6c98a9edcd13edd8997db12af528af8afec4b3441eecfa440572c3e88b6a

Observation 820ec6f5-c22a-461a-9ce0-6d02045a981c · outbound

This paper cites Achiev- ing cross modal generalization with multimodal unified rep- resentation.

Open-set Cross Modal Generalization via Multimodal Unified Representation Achiev- ing cross modal generalization with multimodal unified rep- resentation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:39.140480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:32.142674Z digest=sha256:c6a5bd96d146c4eabd6af1065fa2e5daf21de02d1bbb75ec9365f6e0b9d60095

Observation 19e3608f-e6cf-420e-aaaa-a654c73fda0a · outbound

This paper cites Hyper- spectral image classification based on unsupervised hetero- geneous domain adaptation cyclegan.

Open-set Cross Modal Generalization via Multimodal Unified Representation Hyper- spectral image classification based on unsupervised hetero- geneous domain adaptation cyclegan

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:38.914667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:32.246628Z digest=sha256:5ff0382843f4e4baa33e98f71481ffb50dd459940ed950706aecbdc873ab64dd

Observation 5d2da7bd-d3ac-45e0-a4d3-cf220f061a64 · outbound

This paper cites Class semantics modulation for open-set in- stance segmentation.

Open-set Cross Modal Generalization via Multimodal Unified Representation Class semantics modulation for open-set in- stance segmentation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:38.668841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:32.352440Z digest=sha256:3c3cc53f0bea04dd026642c2072b1fb67b8385a4d6f189efdb4863280b6f96f6

Observation 0c7c5007-6929-411b-8fac-ce29d6d3443e · outbound

This paper cites Learn- ing domain-invariant and discriminative features for homo- geneous unsupervised domain adaptation.

Open-set Cross Modal Generalization via Multimodal Unified Representation Learn- ing domain-invariant and discriminative features for homo- geneous unsupervised domain adaptation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:38.436124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:32.477375Z digest=sha256:ef0055356ccd7619db5b94ff1b68448e3e328737fad997aeca9cb9c2ba01680f

Observation c55857b4-5572-411e-8a7d-b081986d96b4 · outbound

This paper cites A du- ality based approach for realtime tv-l 1 optical flow.

Open-set Cross Modal Generalization via Multimodal Unified Representation A du- ality based approach for realtime tv-l 1 optical flow

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:38.192212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:32.607661Z digest=sha256:41cbf1abfac7a22301cb18caeccee1d8e8c4e2d351211d4e9650d72253a1015a

Observation 7f10b910-2872-4c0b-8b06-b8d5b1dc0448 · outbound

This paper cites mixup: Beyond empirical risk management.

Open-set Cross Modal Generalization via Multimodal Unified Representation mixup: Beyond empirical risk management

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:37.927785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:32.769497Z digest=sha256:6b24b58731cc257b2edfba9c4922da045e7e6563206a22c9058561b2e454f5a3

Observation 78b6b8c4-a9aa-4cd1-91a4-8751551b2670 · outbound

This paper cites Extending multi-modal contrastive rep- resentations.

Open-set Cross Modal Generalization via Multimodal Unified Representation Extending multi-modal contrastive rep- resentations

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:37.701085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:32.872091Z digest=sha256:1e1f56e5f14aa44306bb22ae5c331dc93c9b19a6f7cbf0be1ece8739011c8510

Observation 92ae4c94-4bfa-47f6-af31-dc2573259f2d · outbound

This paper cites Towards effective multi-modal interchanges in zero-resource sounding object localization.

Open-set Cross Modal Generalization via Multimodal Unified Representation Towards effective multi-modal interchanges in zero-resource sounding object localization

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:37.459916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:33.001801Z digest=sha256:61aefcf2db236d1d059346a73b846dc19d53cd3bda48f141b74f7385f1c84dad

Observation 55de53c6-1ece-419d-9a92-e46d4b6bf367 · outbound

This paper cites Positive sample propagation along the audio- visual event line.

Open-set Cross Modal Generalization via Multimodal Unified Representation Positive sample propagation along the audio- visual event line

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:37.179781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:33.180736Z digest=sha256:f2d70103e8953458acbdfe3f0a842439e5924f9424a45729161cb28f126b92b5

Observation 97ddbd13-ffc0-4fa6-b81a-5493df3479fa · outbound

This paper cites Contrastive pos- itive sample propagation along the audio-visual event line.

Open-set Cross Modal Generalization via Multimodal Unified Representation Contrastive pos- itive sample propagation along the audio-visual event line

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:36.975633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:33.294913Z digest=sha256:4a56c8b78ad01d556cac28727afe1d383b6c5efb807a4487ba69d04fcfdb988c

Observation 9e964db3-0843-4c9d-8e06-6d5cebaacee9 · outbound

This paper cites As shown in Figure 4, applying the same mask to paired multimodal samples helps improve model perfor- mance.

Open-set Cross Modal Generalization via Multimodal Unified Representation As shown in Figure 4, applying the same mask to paired multimodal samples helps improve model perfor- mance

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:36.711555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:33.446319Z digest=sha256:eb4f117ed962172776df98d6352922f6cc78505902be8beceff1aa1bd9e2052b

Observation ea7d8129-f526-46f0-a843-fcf27af389a9 · outbound

This paper cites As shown in Figure 5, we experimented with five different settings: 256, 400, 512, 800, and 1024.

Open-set Cross Modal Generalization via Multimodal Unified Representation As shown in Figure 5, we experimented with five different settings: 256, 400, 512, 800, and 1024

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:36.441196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:33.590702Z digest=sha256:e8795e83e117c6f4925f7e7a5e29b0436ed3710c7c74eb36a69117a69a6185de

Observation de07682c-87f8-4353-a6c1-e34fea5fdacb · outbound

This paper cites Lcoarse serves as the foundation of the model, while Lf ine and Lcujp further refine the unified representation space and enhance the model’s open-domain detection capabilities.

Open-set Cross Modal Generalization via Multimodal Unified Representation Lcoarse serves as the foundation of the model, while Lf ine and Lcujp further refine the unified representation space and enhance the model’s open-domain detection capabilities

Reference 63

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T15:48:36.199359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:33.713818Z digest=sha256:686fff3cbfb7bbcbdf1e05ef2d9334f9798ed4841f3b189be07a0ed7edce0a46

Observation 0ed9da58-6fc8-4401-ab6b-a4ca137ccec8 · outbound

This paper cites CUJP8, despite having more split block reorder- ing, optimizes memory usage and reduces training time compared to MMJP6 [14].

Open-set Cross Modal Generalization via Multimodal Unified Representation CUJP8, despite having more split block reorder- ing, optimizes memory usage and reduces training time compared to MMJP6 [14]

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:35.956451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:33.844344Z digest=sha256:fcf9c88a4cc13413d7a578f5a44b4dbea27bac4cc551c2aca31753e5f7d15ce2

Observation 4d271cca-30a9-434f-9f9a-a9835f3a91a5 · outbound

This paper cites The visualization maps audio-video-text triplets from the Valor32K dataset [7] into the unified rep- resentation space (codebook).

Open-set Cross Modal Generalization via Multimodal Unified Representation The visualization maps audio-video-text triplets from the Valor32K dataset [7] into the unified rep- resentation space (codebook)

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:35.616927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-06T15:48:33.963779Z digest=sha256:b623489f09e7c8c7d98e4f3626666eda95ce26741fe1caeb740d93649c8008de

Pith citing papers

Observation aa53c0d5-f2ba-4c21-9783-00fd73b7fa6f · inbound

TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather Removal cites this paper.

TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather Removal Open-set Cross Modal Generalization via Multimodal Unified Representation

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:53:34.172007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-08-05T21:53:28.021948Z digest=sha256:fe0c476124d9759568c3ed1582f62ca00fec2d325e10f88ca766b12317cae691