Pith. sign in

Paper Citation Record · LEDGER

Open-set Cross Modal Generalization via Multimodal Unified Representation

As of 10 August 2026, this Paper Citation Record lists 65 of 65 outbound references and 1 inbound Pith citation observation for arXiv:2507.14935.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.14935 v1

Coverage vector

measured 65 of 65 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:48:33.963779Z

measured 66 of 66 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-05T21:53:28.021948Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T21:53:34.084878Z

Reference resolution

65 of 65 outbound references displayed

  • verified exact5
  • verified fuzzy45
  • unresolved14
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9a653eda-7ed8-461e-9c05-254125e640ed · outbound

This paper cites Robust cross-modal representation learning with progressive self- distillation.

Open-set Cross Modal Generalization via Multimodal Unified Representation Robust cross-modal representation learning with progressive self- distillation

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:46.945276Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:25.548996Z digest=sha256:179ada65c3e03efca4f4a1d2608a7948b8ef8acf9d1044387ed5b1ebfbeb51a8

Observation da722cd7-14fe-4fba-9270-de4615f6ea92 · outbound

This paper cites On the effectiveness of image rotation for open set domain adaptation.

Open-set Cross Modal Generalization via Multimodal Unified Representation On the effectiveness of image rotation for open set domain adaptation

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:46.699431Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:25.670170Z digest=sha256:fae43f8701ae8aa8e134c02a2fe1b81555c8b7aaca352bfabbff97e69a2b8521

Observation f27956e9-7a49-4e87-ac8a-f86e3d175bdf · outbound

This paper cites Domain generalization by solving jigsaw puzzles.

Open-set Cross Modal Generalization via Multimodal Unified Representation Domain generalization by solving jigsaw puzzles

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:46.448429Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:25.778914Z digest=sha256:cc60368bee92f6256be054fcd3d46caa904390834f74eebbe7738fe0d390fac0

Observation dee88b25-f9c6-4d64-b7b1-a9ce145e3d7c · outbound

This paper cites Collecting highly paral- lel data for paraphrase evaluation.

Open-set Cross Modal Generalization via Multimodal Unified Representation Collecting highly paral- lel data for paraphrase evaluation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:25.886544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:25.886544Z digest=sha256:54607dc6512133b10ae8dd1d084e9ca876e397afacd73b6597adbde358c00d52

Observation 396e61a4-d780-417d-a576-1fbeb402b0c1 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

Open-set Cross Modal Generalization via Multimodal Unified Representation Vggsound: A large-scale audio-visual dataset

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:46.213437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:26.060304Z digest=sha256:edd6e430b058e9f818ca0a9025e750a70eaae6f67296726429511ab04e0a1ab2

Observation 40a90f4a-be71-4554-a7b7-24b51ec6dc28 · outbound

This paper cites Hts-at: A hierarchi- cal token-semantic audio transformer for sound classifica- tion and detection.

Open-set Cross Modal Generalization via Multimodal Unified Representation Hts-at: A hierarchi- cal token-semantic audio transformer for sound classifica- tion and detection

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:45.924381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:26.165853Z digest=sha256:c8c58bfd0d0136a5952b8954895a4966269fa783330c42419a1429f8e0011323

Observation 9eacc5aa-1288-437f-a5f3-2e665af6c317 · outbound

This paper cites VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset.

Open-set Cross Modal Generalization via Multimodal Unified Representation VALOR: Vision-Audio-Language Omni-Perception Pretraining Model and Dataset

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:26.258804Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:26.258804Z digest=sha256:61b608c5716664658b91abe42fdf54f16cdaf909465926e22322ab2aee47cbd9

Observation 4528c55e-2a93-4781-ae74-7c72c085ea0e · outbound

This paper cites Uniter: Universal image-text representation learning.

Open-set Cross Modal Generalization via Multimodal Unified Representation Uniter: Universal image-text representation learning

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:45.722400Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:26.393490Z digest=sha256:342d685551016381588b8c27d5210d19adfdea66c2fe453bb5c22b231b3c4011

Observation 7406c049-2539-466b-aa97-8122938036f1 · outbound

This paper cites Sinkd: Sinkhorn distance minimization for knowledge distillation.

Open-set Cross Modal Generalization via Multimodal Unified Representation Sinkd: Sinkhorn distance minimization for knowledge distillation

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:45.497935Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:26.524128Z digest=sha256:b58a6a9100a96f9480a522d2c12d319cac112ebbd636aab0110678e90f2ef0bc

Observation d694a12a-1115-452f-949a-aca3c1e76ad3 · outbound

This paper cites Sinkhorn distance minimization for knowledge distilla- tion.

Open-set Cross Modal Generalization via Multimodal Unified Representation Sinkhorn distance minimization for knowledge distilla- tion

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:45.205864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:26.655853Z digest=sha256:84817484c69db0fa6eb459b8e7afa4ec0ac8b5a3128ab7833b5893173bb39b3f

Observation 1b229b04-2c2d-4e19-a83e-327b027cdebf · outbound

This paper cites Optical: Leveraging optimal trans- port for contribution allocation in dataset distillation.

Open-set Cross Modal Generalization via Multimodal Unified Representation Optical: Leveraging optimal trans- port for contribution allocation in dataset distillation

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:44.922516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:26.771225Z digest=sha256:13ed4e164409a0babe800c2cd83e2f9b8eaaa6980e58e969afb253819620e8fe

Observation 2f994dd1-b7c2-4e6d-8938-a2e94fdcde70 · outbound

This paper cites Layoutenc: Leveraging enhanced layout rep- resentations for transformer-based complex scene synthesis.

Open-set Cross Modal Generalization via Multimodal Unified Representation Layoutenc: Leveraging enhanced layout rep- resentations for transformer-based complex scene synthesis

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:44.599120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:26.885590Z digest=sha256:2159022fe87a0cc07ae081ca78516f9677eebacb5098aec6cabcb2f672b75382

Observation a3cf246e-ee9e-41e9-9ee8-07a0e685898f · outbound

This paper cites Streetsurfgs: Scalable ur- ban street surface reconstruction with planar-based gaussian splatting.

Open-set Cross Modal Generalization via Multimodal Unified Representation Streetsurfgs: Scalable ur- ban street surface reconstruction with planar-based gaussian splatting

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:44.384755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:27.037175Z digest=sha256:f1a72b48d753683e2181f5449e8a1f7083f2da5eaa4312337ff66958e54341cd

Observation af806a17-5526-4171-8036-6b1e99c5917e · outbound

This paper cites Towards Multimodal Open-Set Domain Generalization and Adaptation through Self-supervision.

Open-set Cross Modal Generalization via Multimodal Unified Representation Towards Multimodal Open-Set Domain Generalization and Adaptation through Self-supervision

Reference 14

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:48:35.385262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:27.140292Z digest=sha256:c751c7281c29e1e75d478b79c3143eb0a4b41e916708efe89ff5f2e7703fa628

Observation 33e87a74-1eb6-464c-b70a-86a9749d9f10 · outbound

This paper cites Simmmdg: A simple and effective framework for multi-modal domain generalization.

Open-set Cross Modal Generalization via Multimodal Unified Representation Simmmdg: A simple and effective framework for multi-modal domain generalization

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:44.159885Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:27.241584Z digest=sha256:94d3a11708e311bc75d4c0123ba747d16c66cc9dc932969025f6d90345343657

Observation 507023a4-0327-44b5-b8ea-a620ee805697 · outbound

This paper cites Clotho: An audio captioning dataset.

Open-set Cross Modal Generalization via Multimodal Unified Representation Clotho: An audio captioning dataset

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:43.936791Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:27.379391Z digest=sha256:840a1f2c5a70e4ab81325bc7306519bd7c313c8d4bef93a992e5bfaac6dc5fdf

Observation 09215f9d-9f71-494e-b01f-299865375760 · outbound

This paper cites Multi-modal align- ment using representation codebook.

Open-set Cross Modal Generalization via Multimodal Unified Representation Multi-modal align- ment using representation codebook

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:43.707248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:27.511828Z digest=sha256:ac59f2b729f9930529040674fa48cea05b0dae7cad09a2133d6644c1c47c243e

Observation 51531b0d-6a4b-410d-be22-2e8426a6893c · outbound

This paper cites Ace: A generative cross-modal retrieval framework with coarse-to-fine semantic modeling.

Open-set Cross Modal Generalization via Multimodal Unified Representation Ace: A generative cross-modal retrieval framework with coarse-to-fine semantic modeling

Reference 18

Resolution
verified exact
raw_fallback, observed 2026-08-06T15:48:35.160122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:27.648562Z digest=sha256:667d7c39530b0fd22f18a67118b1e879e16f31d43a0045c698d1bc2a3a39dc02

Observation 2f814264-7211-4be5-a0d5-027e32162906 · outbound

This paper cites Slowfast networks for video recognition.

Open-set Cross Modal Generalization via Multimodal Unified Representation Slowfast networks for video recognition

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:27.779563Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:27.779563Z digest=sha256:11712675af2c676405ebe8c1e52f5f6c15b6927da42609e5ba4951e780efe245

Observation fd407d55-5386-4d64-9c20-ccaf26e2e10f · outbound

This paper cites Domain-adversarial training of neural networks.

Open-set Cross Modal Generalization via Multimodal Unified Representation Domain-adversarial training of neural networks

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:43.460597Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:27.947964Z digest=sha256:ffebad15293c546f2b2be342caae0ad8d63246309d5d7e45b154303c67a608a8

Observation 4c64e29f-8989-4401-991e-dc5f1e2c3740 · outbound

This paper cites Imagebind: One embedding space to bind them all.

Open-set Cross Modal Generalization via Multimodal Unified Representation Imagebind: One embedding space to bind them all

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:28.115661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:28.115661Z digest=sha256:5d5b87b10230195a57cd04605fed39fe411ecf72bdef859e153c80dd48125c8f

Observation 5693eb52-ffb9-4a0e-a0b6-c3fd604f14b2 · outbound

This paper cites Enhancing Multimodal Unified Representations for Cross Modal Generalization.

Open-set Cross Modal Generalization via Multimodal Unified Representation Enhancing Multimodal Unified Representations for Cross Modal Generalization

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:28.218101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:28.218101Z digest=sha256:e3f56e0a98e7760018e445ef58daff39fa97b31c98f050777c161e9ee2717ac2

Observation 302450ac-17aa-42a6-8c26-e1ed2f82c0bd · outbound

This paper cites Semantic residual for multimodal unified discrete representation.

Open-set Cross Modal Generalization via Multimodal Unified Representation Semantic residual for multimodal unified discrete representation

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:43.239672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:28.320120Z digest=sha256:0d61d59c52b92a1d633098a69f7a0a6b72e9b5891e7fc601e207bb2a66641498

Observation bdc993b6-96dd-4fdc-b61a-fd7b27d75e34 · outbound

This paper cites Bridging Domain Generalization to Multimodal Domain Generalization via Unified Representations.

Open-set Cross Modal Generalization via Multimodal Unified Representation Bridging Domain Generalization to Multimodal Domain Generalization via Unified Representations

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:28.447490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:28.447490Z digest=sha256:4245e375d184fca8fdcdaf403c39a0a9b219e2bba22e23da359e52c98309b5ee

Observation b0f93d54-2536-44a3-91c0-881c016cfe4f · outbound

This paper cites WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling.

Open-set Cross Modal Generalization via Multimodal Unified Representation WavTokenizer: an Efficient Acoustic Discrete Codec Tokenizer for Audio Language Modeling

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:28.600859Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:28.600859Z digest=sha256:605e8e771b52e16fb8a077f37f702a46811384a26d3507db8711868db101198d

Observation 807fa901-b252-4444-8f09-06bbc1d8feb9 · outbound

This paper cites Learning to generalize: Meta-learning for do- main generalization.

Open-set Cross Modal Generalization via Multimodal Unified Representation Learning to generalize: Meta-learning for do- main generalization

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:42.994612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:28.728539Z digest=sha256:060aacc53d9cd2da7ed0428f887bd8158fa3bfe19148614865357b63a73b5ec6

Observation e4e07c83-b58c-4ac9-9336-480b1ebb98c7 · outbound

This paper cites Domain generalization for med- ical imaging classification with linear-dependency regular- ization.

Open-set Cross Modal Generalization via Multimodal Unified Representation Domain generalization for med- ical imaging classification with linear-dependency regular- ization

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:42.783890Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:28.883688Z digest=sha256:c6df1ca7699110a1ca6c2735cb48489147bc2a01d0abcfa1640dd3b308e27973

Observation 882ef81a-0a97-43bc-8971-19b5b257d701 · outbound

This paper cites Adjustment and alignment for unbiased open set domain adaptation.

Open-set Cross Modal Generalization via Multimodal Unified Representation Adjustment and alignment for unbiased open set domain adaptation

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:42.499297Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:29.019728Z digest=sha256:75145ffdeda049367f370cc8279a6c25f361b5aa7f1af8c12d0a175febe4e430

Observation b9e429f1-b3ba-46b4-82a2-594c73c93eff · outbound

This paper cites Cross-Modal Discrete Representation Learning.

Open-set Cross Modal Generalization via Multimodal Unified Representation Cross-Modal Discrete Representation Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:29.139488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:29.139488Z digest=sha256:996b51e58efa9054d2eaf828722b10d4d91d986e2758413cb5d61cf553c5caad

Observation 1c7a639d-2e41-4ca0-b4a7-ed46348417e7 · outbound

This paper cites Feddg: Federated domain generalization on medical image segmentation via episodic learning in continuous fre- quency space.

Open-set Cross Modal Generalization via Multimodal Unified Representation Feddg: Federated domain generalization on medical image segmentation via episodic learning in continuous fre- quency space

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:42.266176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:29.274536Z digest=sha256:5eadf81f4467e4c03e380b8f34e2f389b5f5f61e2d3297787939d5c272326558

Observation a53e4ee0-723a-424f-9877-952fddd56894 · outbound

This paper cites Swin transformer: Hierarchical vision transformer using shifted windows.

Open-set Cross Modal Generalization via Multimodal Unified Representation Swin transformer: Hierarchical vision transformer using shifted windows

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:29.435332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:29.435332Z digest=sha256:b5436390c1f3f9f24255e0f0293eb472f8efeb4b1e0a855cc4e0eba6607202bc

Observation 3d48944d-23b5-4973-ace7-88db099cecdf · outbound

This paper cites Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks.

Open-set Cross Modal Generalization via Multimodal Unified Representation Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:29.595479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:29.595479Z digest=sha256:64a996e18bb8ecda4619c680934b72f38b921e2910759178fd8668d304227045

Observation 604260d1-9db9-4e96-9617-9735f183c9ad · outbound

This paper cites Unsupervised learning of visual representations by solving jigsaw puzzles.

Open-set Cross Modal Generalization via Multimodal Unified Representation Unsupervised learning of visual representations by solving jigsaw puzzles

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:41.956515Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:29.699219Z digest=sha256:4ee986162d67679300e8ba3a1d4404a466c40ec9532e24ef6ec61855e91a45fa

Observation a19c92f6-8dc0-4a89-b56e-f1530b4f90c6 · outbound

This paper cites Two at once: Enhancing learning and generalization capacities via ibn-net.

Open-set Cross Modal Generalization via Multimodal Unified Representation Two at once: Enhancing learning and generalization capacities via ibn-net

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:41.672767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:29.830914Z digest=sha256:262ffd6952b536cff93b87aec97ccf8eb1c8385a8c7a737371a524cd4612d985

Observation 0a5add89-0bdf-4104-b171-e11c1a236067 · outbound

This paper cites Estimating Visual Information From Audio Through Manifold Learning.

Open-set Cross Modal Generalization via Multimodal Unified Representation Estimating Visual Information From Audio Through Manifold Learning

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:48:34.769435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:29.953749Z digest=sha256:b890ded15bbeed5e43bfaa475d95acd8336ade54e21378d2f298780ed11b58b3

Observation b771fb0b-8b4f-4812-8c64-7d035d63b104 · outbound

This paper cites Audio-visual speech recognition with a hybrid ctc/attention architecture.

Open-set Cross Modal Generalization via Multimodal Unified Representation Audio-visual speech recognition with a hybrid ctc/attention architecture

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:41.464962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:30.116086Z digest=sha256:b2c5469010e35303391f3ad24d7d2d84f39c650abaeb6352aaf475e039dac748

Observation 0a61c64f-76dc-42b1-930c-c750e10ea19f · outbound

This paper cites Domain generalization through audio- visual relative norm alignment in first person action recog- nition.

Open-set Cross Modal Generalization via Multimodal Unified Representation Domain generalization through audio- visual relative norm alignment in first person action recog- nition

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:41.216157Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:30.254045Z digest=sha256:d94d88f9b69ea59f4e3ebdb9e1baf9484b7a5d38a040a863c08705be40895c10

Observation eb751c58-b667-4e10-8fff-866a3022576c · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Open-set Cross Modal Generalization via Multimodal Unified Representation Learning transferable visual models from natural language supervi- sion

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:30.385147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:30.385147Z digest=sha256:53975a44956afc21f95b46a19ce958733e4f09a0b940c6491b2725a88b4caa91

Observation 0f742929-52d9-43e4-a3c9-5c0dc224b5e2 · outbound

This paper cites Mask2anomaly: Mask transformer for uni- versal open-set segmentation.

Open-set Cross Modal Generalization via Multimodal Unified Representation Mask2anomaly: Mask transformer for uni- versal open-set segmentation

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:40.949706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:30.518039Z digest=sha256:7b0bd05638ab03ebf251ad4e22fdf3f766ecc7fb13730cd304dcaff0574c17f6

Observation 62bdf01a-5e88-4d31-b262-8bf4980429fa · outbound

This paper cites XKD: Cross-modal Knowledge Distillation with Domain Alignment for Video Representation Learning.

Open-set Cross Modal Generalization via Multimodal Unified Representation XKD: Cross-modal Knowledge Distillation with Domain Alignment for Video Representation Learning

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:48:34.487611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:30.648250Z digest=sha256:8ce31564a8e877a368dee24cdf45f18d70fb8d46103e874cae326ff2bd9d1b92

Observation 24c727c0-79b1-4434-95d5-26697d9abdf6 · outbound

This paper cites Open domain generalization with domain- augmented meta-learning.

Open-set Cross Modal Generalization via Multimodal Unified Representation Open domain generalization with domain- augmented meta-learning

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:40.701072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:30.804084Z digest=sha256:f07b77848f7d29b315b0e80d7fad0987f1241b0b6cfec1369e9d83f45dd6dcb8

Observation 7281a904-b633-496b-bf6b-43836baae6e1 · outbound

This paper cites UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild.

Open-set Cross Modal Generalization via Multimodal Unified Representation UCF101: A Dataset of 101 Human Actions Classes From Videos in The Wild

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:30.917220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:30.917220Z digest=sha256:113e814758c6e7152426ae037ab1c326ce1ce4c56d9865977d96d4783ce0e13f

Observation ef32786e-a737-49b6-8428-03a2dcfd6dff · outbound

This paper cites Audio-visual event localization in unconstrained videos.

Open-set Cross Modal Generalization via Multimodal Unified Representation Audio-visual event localization in unconstrained videos

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:40.399365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:31.035641Z digest=sha256:ec1ce483fae5ff6ea23ed7f750ff7974a4cb76230f3f7bc757fc9efa3c12ff2c

Observation 9651d3e8-bc5d-4b0a-bc27-b1d713ae757f · outbound

This paper cites Unified mul- tisensory perception: Weakly-supervised audio-visual video parsing.

Open-set Cross Modal Generalization via Multimodal Unified Representation Unified mul- tisensory perception: Weakly-supervised audio-visual video parsing

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:40.163541Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:31.148002Z digest=sha256:3f8ce3d4cffdeefd26dc3fb83e2618bcbfb257e0e6b5aa363167aa932ab06c9a

Observation 5de95d9f-85c0-426e-9575-b89836e6821f · outbound

This paper cites Domain randomization for transferring deep neural networks from simulation to the real world.

Open-set Cross Modal Generalization via Multimodal Unified Representation Domain randomization for transferring deep neural networks from simulation to the real world

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:39.944179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:31.310142Z digest=sha256:8412bd2c7691dc61d5e42aa3efdf329cb4e0d3490e152a6b964283f3ff1b3613

Observation ec0d5cd5-d07d-4e68-a099-f64740d5f903 · outbound

This paper cites Deep Domain Confusion: Maximizing for Domain Invariance.

Open-set Cross Modal Generalization via Multimodal Unified Representation Deep Domain Confusion: Maximizing for Domain Invariance

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:31.427887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:31.427887Z digest=sha256:0bf751ef44f0e14d9b057eca8c9ed1b756fecd43d9d0e776d67d4adffb8902e7

Observation 4f23e8e0-ffe8-4fd5-b1ec-ea7f4cc4fc3e · outbound

This paper cites IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models.

Open-set Cross Modal Generalization via Multimodal Unified Representation IRBridge: Solving Image Restoration Bridge with Pre-trained Generative Diffusion Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-06T15:48:31.574133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:48:31.574133Z digest=sha256:5a6ee1b65c5ccf1da10e19527cb8e18f9dbe1801340211061b37dd1511cb1656

Observation 02e59edd-c139-4d7d-a293-980bf430387a · outbound

This paper cites Towards Transformer-Based Aligned Generation with Self-Coherence Guidance.

Open-set Cross Modal Generalization via Multimodal Unified Representation Towards Transformer-Based Aligned Generation with Self-Coherence Guidance

Reference 48

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:48:34.215630Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:31.694527Z digest=sha256:ce041e38631c85fc5867990303e12ffa9a41c67f7fede973e5e9be1bfb81fb98

Observation 449e3041-9ccf-43fd-b742-3b06a1d220cb · outbound

This paper cites Vlmixer: Unpaired vision-language pre-training via cross-modal cutmix.

Open-set Cross Modal Generalization via Multimodal Unified Representation Vlmixer: Unpaired vision-language pre-training via cross-modal cutmix

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:39.699412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:31.802571Z digest=sha256:37d73abc8f2d849d10bb158619a5aa95c6d0ee744c299c34c35562ad35480134

Observation 4d17012a-086b-478e-98a8-7edd453a268c · outbound

This paper cites General- izable decision boundaries: Dualistic meta-learning for open set domain generalization.

Open-set Cross Modal Generalization via Multimodal Unified Representation General- izable decision boundaries: Dualistic meta-learning for open set domain generalization

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:39.416582Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:31.971663Z digest=sha256:1798569e4815f954f19e4f4e63deaea41de756b9dd4a6a080e2e34057ca8b725

Observation 820ec6f5-c22a-461a-9ce0-6d02045a981c · outbound

This paper cites Achiev- ing cross modal generalization with multimodal unified rep- resentation.

Open-set Cross Modal Generalization via Multimodal Unified Representation Achiev- ing cross modal generalization with multimodal unified rep- resentation

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:39.140480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:32.142674Z digest=sha256:8ba2dc7fbf2212e71cf2b12bc98f9aeba1fe6f9e6fc5139c387ee450c1fbafb4

Observation 19e3608f-e6cf-420e-aaaa-a654c73fda0a · outbound

This paper cites Hyper- spectral image classification based on unsupervised hetero- geneous domain adaptation cyclegan.

Open-set Cross Modal Generalization via Multimodal Unified Representation Hyper- spectral image classification based on unsupervised hetero- geneous domain adaptation cyclegan

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:38.914667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:32.246628Z digest=sha256:6d26d3d979673a4f35fc3d793a54bb2f7c447283efcc86ac4a08b195407b00dd

Observation 5d2da7bd-d3ac-45e0-a4d3-cf220f061a64 · outbound

This paper cites Class semantics modulation for open-set in- stance segmentation.

Open-set Cross Modal Generalization via Multimodal Unified Representation Class semantics modulation for open-set in- stance segmentation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:38.668841Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:32.352440Z digest=sha256:c8e64cb11f6a6d3792e36aecfe508673adfae09f94d79ca68e628037f0112832

Observation 0c7c5007-6929-411b-8fac-ce29d6d3443e · outbound

This paper cites Learn- ing domain-invariant and discriminative features for homo- geneous unsupervised domain adaptation.

Open-set Cross Modal Generalization via Multimodal Unified Representation Learn- ing domain-invariant and discriminative features for homo- geneous unsupervised domain adaptation

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:38.436124Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:32.477375Z digest=sha256:03019562da253b9db2ad1bce28af7f44f43c293840a2e525112f373157999607

Observation c55857b4-5572-411e-8a7d-b081986d96b4 · outbound

This paper cites A du- ality based approach for realtime tv-l 1 optical flow.

Open-set Cross Modal Generalization via Multimodal Unified Representation A du- ality based approach for realtime tv-l 1 optical flow

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:38.192212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:32.607661Z digest=sha256:1bb50dac230912c29db0b9919fd836fc01c229e45877f8360bda2948172451a1

Observation 7f10b910-2872-4c0b-8b06-b8d5b1dc0448 · outbound

This paper cites mixup: Beyond empirical risk management.

Open-set Cross Modal Generalization via Multimodal Unified Representation mixup: Beyond empirical risk management

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:37.927785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:32.769497Z digest=sha256:0c0e9b7b8298125d40f94c72e09aef79285142b6e7368b2064145a16f3864f30

Observation 78b6b8c4-a9aa-4cd1-91a4-8751551b2670 · outbound

This paper cites Extending multi-modal contrastive rep- resentations.

Open-set Cross Modal Generalization via Multimodal Unified Representation Extending multi-modal contrastive rep- resentations

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:37.701085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:32.872091Z digest=sha256:2d6b218aadcda610c8db809494369aee0d0db6e052f8588c61abab878a2f8bd0

Observation 92ae4c94-4bfa-47f6-af31-dc2573259f2d · outbound

This paper cites Towards effective multi-modal interchanges in zero-resource sounding object localization.

Open-set Cross Modal Generalization via Multimodal Unified Representation Towards effective multi-modal interchanges in zero-resource sounding object localization

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:37.459916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:33.001801Z digest=sha256:2faa77767fc6904fd13f995eb2098253fcb857e53519d867da9a854b36a23d0b

Observation 55de53c6-1ece-419d-9a92-e46d4b6bf367 · outbound

This paper cites Positive sample propagation along the audio- visual event line.

Open-set Cross Modal Generalization via Multimodal Unified Representation Positive sample propagation along the audio- visual event line

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:37.179781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:33.180736Z digest=sha256:000c23f1a271a63f89ad5c2261bc0c02657b4124cbbea306b430014e093b3beb

Observation 97ddbd13-ffc0-4fa6-b81a-5493df3479fa · outbound

This paper cites Contrastive pos- itive sample propagation along the audio-visual event line.

Open-set Cross Modal Generalization via Multimodal Unified Representation Contrastive pos- itive sample propagation along the audio-visual event line

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:36.975633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:33.294913Z digest=sha256:0e7ecd6f90b4cd670a8411b2cba5dd0f02fce6c9913d9a5cae31195b9386506a

Observation 9e964db3-0843-4c9d-8e06-6d5cebaacee9 · outbound

This paper cites As shown in Figure 4, applying the same mask to paired multimodal samples helps improve model perfor- mance.

Open-set Cross Modal Generalization via Multimodal Unified Representation As shown in Figure 4, applying the same mask to paired multimodal samples helps improve model perfor- mance

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:36.711555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:33.446319Z digest=sha256:bce5bcd639a56b7ae7b15db0a171d0eb4966d4a0695612922c2f4c48fac44248

Observation ea7d8129-f526-46f0-a843-fcf27af389a9 · outbound

This paper cites As shown in Figure 5, we experimented with five different settings: 256, 400, 512, 800, and 1024.

Open-set Cross Modal Generalization via Multimodal Unified Representation As shown in Figure 5, we experimented with five different settings: 256, 400, 512, 800, and 1024

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:36.441196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:33.590702Z digest=sha256:a0b6027617a6592f02990f328165539781b0325653c91401b2151a95b46405f5

Observation de07682c-87f8-4353-a6c1-e34fea5fdacb · outbound

This paper cites Lcoarse serves as the foundation of the model, while Lf ine and Lcujp further refine the unified representation space and enhance the model’s open-domain detection capabilities.

Open-set Cross Modal Generalization via Multimodal Unified Representation Lcoarse serves as the foundation of the model, while Lf ine and Lcujp further refine the unified representation space and enhance the model’s open-domain detection capabilities

Reference 63

Resolution
malformed identifier
raw_fallback, observed 2026-08-06T15:48:36.199359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:33.713818Z digest=sha256:36dbce2b31b039beaacfe18baf2a6a6c4253fc85ea938e8ee247c2d249dd9393

Observation 0ed9da58-6fc8-4401-ab6b-a4ca137ccec8 · outbound

This paper cites CUJP8, despite having more split block reorder- ing, optimizes memory usage and reduces training time compared to MMJP6 [14].

Open-set Cross Modal Generalization via Multimodal Unified Representation CUJP8, despite having more split block reorder- ing, optimizes memory usage and reduces training time compared to MMJP6 [14]

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:35.956451Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:33.844344Z digest=sha256:55deeb8fb49940402150e56d4f8176d5153a59ef7d5e4e36813ed025f8f1a4dc

Observation 4d271cca-30a9-434f-9f9a-a9835f3a91a5 · outbound

This paper cites The visualization maps audio-video-text triplets from the Valor32K dataset [7] into the unified rep- resentation space (codebook).

Open-set Cross Modal Generalization via Multimodal Unified Representation The visualization maps audio-video-text triplets from the Valor32K dataset [7] into the unified rep- resentation space (codebook)

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:48:35.616927Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T15:48:33.963779Z digest=sha256:664ae6ad542a34e6aacd6137397c9fcb295ef36bea476630edb9d04f6e06674d

Pith citing papers

Observation aa53c0d5-f2ba-4c21-9783-00fd73b7fa6f · inbound

TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather Removal cites this paper.

TAP: Parameter-efficient Task-Aware Prompting for Adverse Weather Removal Open-set Cross Modal Generalization via Multimodal Unified Representation

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-08-05T21:53:34.172007Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-05T21:53:28.021948Z digest=sha256:757b754e8920bfc36e6b64f5f045cacee90c9f5acc11cd340bc55932dc9ee23e