Pith. sign in

Paper Citation Record · LEDGER

GeoMM: On Geodesic Perspective for Multi-modal Learning

As of 16 August 2026, this Paper Citation Record lists 100 of 111 outbound references and 0 inbound Pith citation observations for arXiv:2505.11216.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.11216 v1

Coverage vector

measured 100 of 111 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T21:01:38.878311Z

measured 100 of 100 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

100 of 111 outbound references displayed

  • verified exact3
  • verified fuzzy44
  • unresolved53
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 66cfa46f-a0bd-4a78-a9c2-e6ee31dbede3 · outbound

This paper cites Geometry of oblique projections.

GeoMM: On Geodesic Perspective for Multi-modal Learning Geometry of oblique projections

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:01:39.347568Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.413972Z digest=sha256:d90406ea72332cf13eb9bf997ac3791a0f0543fa12c990fdcd8886c91e251254

Observation 8659d41e-7dc6-4ffd-93b0-1cb8e5a0b684 · outbound

This paper cites Vqa: Visual question an- swering.

GeoMM: On Geodesic Perspective for Multi-modal Learning Vqa: Visual question an- swering

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.420078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.420078Z digest=sha256:d74c214171bf47053db8ffcb625567b633c23e2f49db691844232cdb8b1c74a1

Observation 55c1edf9-c5ee-4ba4-9e2f-9b139486cc82 · outbound

This paper cites Geodesic matting: A framework for fast interactive image and video seg- mentation and matting.

GeoMM: On Geodesic Perspective for Multi-modal Learning Geodesic matting: A framework for fast interactive image and video seg- mentation and matting

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.424841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.424841Z digest=sha256:38cb9281afd82dc440213ad1f87cb222985359b55515633aa885c3460fd53344

Observation b7dec225-807d-4d7d-82d5-77a26f2fff1c · outbound

This paper cites Vlmo: Uni- fied vision-language pre-training with mixture-of- modality-experts.

GeoMM: On Geodesic Perspective for Multi-modal Learning Vlmo: Uni- fied vision-language pre-training with mixture-of- modality-experts

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.429348Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.429348Z digest=sha256:24f1e413fb3453d5a1697fe0a56258a92bc37a6080a6a2218e4b66b2f6cb05d1

Observation 21f0c57b-7496-4784-8aa4-7068f74eefbe · outbound

This paper cites Grit-vlp: Grouped mini-batch sam- pling for efficient vision and language pre-training.

GeoMM: On Geodesic Perspective for Multi-modal Learning Grit-vlp: Grouped mini-batch sam- pling for efficient vision and language pre-training

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.434050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.434050Z digest=sha256:ca7faea43acbd5ff8612e1bb8ec13b258563c908c84822ff1834529fe27b7e37

Observation 7a3ce4a2-f859-4cbe-af60-3c16ac592db2 · outbound

This paper cites Mafa: Managing false negatives for vision-language pre-training.

GeoMM: On Geodesic Perspective for Multi-modal Learning Mafa: Managing false negatives for vision-language pre-training

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.438839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.438839Z digest=sha256:bef39a82585a93f091e7630e20e592adb445697b590ee685454874c8fee3712d

Observation a48256e5-9822-428d-9acf-0f65b344a5e5 · outbound

This paper cites End-to-end object detection with trans- formers.

GeoMM: On Geodesic Perspective for Multi-modal Learning End-to-end object detection with trans- formers

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.443858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.443858Z digest=sha256:0c9ed6035ce1735b7564175f099954844b4b3f92d952b94cb1d1f93c7504e8a6

Observation db3622f6-0593-4b8f-87af-fff6fcf727bb · outbound

This paper cites Un- supervised learning of visual features by contrasting cluster assignments.

GeoMM: On Geodesic Perspective for Multi-modal Learning Un- supervised learning of visual features by contrasting cluster assignments

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.448550Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.448550Z digest=sha256:9464e659635a8a83861f9572bbdce56accdac0e638f9a4bb2d9a5d400fd0951c

Observation fb5ad8ad-31d2-4fa2-a7ab-c9bb930b2de1 · outbound

This paper cites Emerging properties in self-supervised vi- sion transformers.

GeoMM: On Geodesic Perspective for Multi-modal Learning Emerging properties in self-supervised vi- sion transformers

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.452997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.452997Z digest=sha256:f2ad946fa18123478f91d2194926cfced172d41e1dc067ab208a1299e7228921

Observation 99d13011-6d78-4073-b0c4-48eab8b4d2d9 · outbound

This paper cites Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts.

GeoMM: On Geodesic Perspective for Multi-modal Learning Conceptual 12m: Pushing web-scale image-text pre-training to recognize long-tail visual concepts

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.457508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.457508Z digest=sha256:4b8e70b0d25186138301bd995300de183177320e75f71eb2bcde9d2d3a4e2ec7

Observation 274db597-f706-466c-94af-bff1aab1ab82 · outbound

This paper cites STAIR: Learning Sparse Text and Image Representation in Grounded Tokens.

GeoMM: On Geodesic Perspective for Multi-modal Learning STAIR: Learning Sparse Text and Image Representation in Grounded Tokens

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.462143Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.462143Z digest=sha256:1a0d6909b94fd0d8f5134d9f2cf0cb413ab5156c04ae852a5049036b0103f761

Observation 65cd4779-0669-4769-8fb5-cf7d80842a14 · outbound

This paper cites Vlp: A survey on vision-language pre-training.

GeoMM: On Geodesic Perspective for Multi-modal Learning Vlp: A survey on vision-language pre-training

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.466898Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.466898Z digest=sha256:fdde3a1cb829dd308f144cda1517ec72c0f3b105dffce534e405ff28e95729b1

Observation 2fe0d320-69d6-4761-914f-af8965a929be · outbound

This paper cites A simple framework for con- trastive learning of visual representations.

GeoMM: On Geodesic Perspective for Multi-modal Learning A simple framework for con- trastive learning of visual representations

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.471735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.471735Z digest=sha256:d7c897038ef0ae824c1218dae612a1c3182431556579f4748ecde61b377ba086

Observation 004bccc6-a759-477a-b537-6ce78ea73233 · outbound

This paper cites Exploring simple siamese representation learning.

GeoMM: On Geodesic Perspective for Multi-modal Learning Exploring simple siamese representation learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.475985Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.475985Z digest=sha256:119777c74a09464672bad2f9a52771dbeecdaf5807367d4ba3bdf38c22eb3719

Observation 4196cc9e-1fce-4152-ae6c-f046f859e2b4 · outbound

This paper cites Improved Baselines with Momentum Contrastive Learning.

GeoMM: On Geodesic Perspective for Multi-modal Learning Improved Baselines with Momentum Contrastive Learning

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.480255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.480255Z digest=sha256:d62f89d082db1772c3dced68a0805b099a268b59b37fddeb6af69617d82892bb

Observation 8adae875-cc7c-4b7c-959d-a92eeedff320 · outbound

This paper cites X-volution: On the unification of convolution and self-attention.

GeoMM: On Geodesic Perspective for Multi-modal Learning X-volution: On the unification of convolution and self-attention

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:01:39.294667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.484805Z digest=sha256:9e5c7e49256c08662ce837f2f891cbae1b03fa789d7be8e39e65802405437dd1

Observation ad3267df-c538-4ee7-b467-0337bcb55ec6 · outbound

This paper cites Uniter: Universal image-text represen- tation learning.

GeoMM: On Geodesic Perspective for Multi-modal Learning Uniter: Universal image-text represen- tation learning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.489680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.489680Z digest=sha256:f6a921906998b2114650264944657f3dff0e32ad8b1d3417b0dd60180c1b71d3

Observation b5b4d7b4-0948-4db1-bff5-bb6998010a34 · outbound

This paper cites Unsupervised Opinion Summarization Using Approximate Geodesics.

GeoMM: On Geodesic Perspective for Multi-modal Learning Unsupervised Opinion Summarization Using Approximate Geodesics

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-08-15T21:01:39.267081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.494361Z digest=sha256:737b9c19eb4fcdeb5ffd6e1a7d058397b53c15f4fcf41b5ebf3be2707deb170c

Observation 54f20193-659f-4878-8e05-b9d15bf37185 · outbound

This paper cites Geodesics in heat: A new approach to computing distance based on heat flow.

GeoMM: On Geodesic Perspective for Multi-modal Learning Geodesics in heat: A new approach to computing distance based on heat flow

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.499121Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.499121Z digest=sha256:4403577d98731fe61d792605639c2fbb0c0d7ca57b3cbf13e588f5fea9aa9f75

Observation 3478c4dc-1800-4b8a-8fa1-150a8fac611b · outbound

This paper cites Imagenet: A large-scale hierar- chical image database.

GeoMM: On Geodesic Perspective for Multi-modal Learning Imagenet: A large-scale hierar- chical image database

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.503671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.503671Z digest=sha256:399d66c6279000c630e89ea55ccc2e7a8469e6158a75f4caa2da61f894951309

Observation b6a0e3ef-a335-403a-afe0-dea09844afcb · outbound

This paper cites BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding.

GeoMM: On Geodesic Perspective for Multi-modal Learning BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.508186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.508186Z digest=sha256:60e4e40a07565efb30475af4d332506c39a91767a0f52b259ea07e40f4dfef05

Observation 563e7ed6-9e23-49c7-8249-556a7a4e2d39 · outbound

This paper cites Similarity reasoning and filtration for image-text matching.

GeoMM: On Geodesic Perspective for Multi-modal Learning Similarity reasoning and filtration for image-text matching

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.512748Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.512748Z digest=sha256:6f22761692d52025e090209dad29f8f53fd4db2ee564a2508aa7b223460268a4

Observation 8ede5920-4d59-4b00-996e-06fa677e62b4 · outbound

This paper cites A note on two problems in con- nexion with graphs.

GeoMM: On Geodesic Perspective for Multi-modal Learning A note on two problems in con- nexion with graphs

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.517378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.517378Z digest=sha256:3a4df792373c2285023ad7597669214755ae90918c6fbdc1b53f60f8260693fa

Observation 2bf51a31-6b6b-42d0-a3bf-253e0f9ca610 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

GeoMM: On Geodesic Perspective for Multi-modal Learning An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.522500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.522500Z digest=sha256:a8202ef343f214db33ba1db43c63c3d5015f0946d2514201aa552cc37a456060

Observation 31a91c94-7b3c-49b9-b09a-29f0d3911a19 · outbound

This paper cites Algorithm 97: shortest path.

GeoMM: On Geodesic Perspective for Multi-modal Learning Algorithm 97: shortest path

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.528050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.528050Z digest=sha256:8d703ab94cdece48981706624b40e60ef6bde85e4b1f696774edbdf518dbab8a

Observation 3ef1901e-d388-47be-9797-5288ce17d716 · outbound

This paper cites Large-scale adversar- ial training for vision-and-language representation learning.

GeoMM: On Geodesic Perspective for Multi-modal Learning Large-scale adversar- ial training for vision-and-language representation learning

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.532702Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.532702Z digest=sha256:fd2c64efe169dff3bbc1f1c148f0552c763f0638a8fb1ca785ba4933ed5a082c

Observation 94e9c6b0-76e3-4cfc-825b-3ce545cc4019 · outbound

This paper cites Imagebind: One embedding space to bind them all.

GeoMM: On Geodesic Perspective for Multi-modal Learning Imagebind: One embedding space to bind them all

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.537300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.537300Z digest=sha256:3a1253e44f3e8ee57819c0c49ad40df45b9801771b132e8bfd3fadd299bdf90d

Observation 53a7fab5-efb6-4ff8-bbec-c2242f7cbfc4 · outbound

This paper cites Momentum contrast for unsupervised visual representation learning.

GeoMM: On Geodesic Perspective for Multi-modal Learning Momentum contrast for unsupervised visual representation learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.541805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.541805Z digest=sha256:3cee7fe691c0ae0fca4bb35e24d5aeca397c1ab2e85ab1fb56aab13f8939e6b0

Observation a3f9074b-fbab-445c-a06a-d73412d91706 · outbound

This paper cites Masked autoen- coders are scalable vision learners.

GeoMM: On Geodesic Perspective for Multi-modal Learning Masked autoen- coders are scalable vision learners

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.546457Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.546457Z digest=sha256:3533fedfcd4d10b014be745d42a18f138f337ed29b92b0c7ccabd0209d3329ee

Observation fc0aedf2-1406-4814-93fc-e27dac959f1f · outbound

This paper cites Geonet: Deep geodesic networks for point cloud analysis.

GeoMM: On Geodesic Perspective for Multi-modal Learning Geonet: Deep geodesic networks for point cloud analysis

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.551765Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.551765Z digest=sha256:e05204d44f5a24ee71a9854af93d4e058a68a147c25a906d707ae1ff3bb3083e

Observation 276fe121-e048-4455-a1f0-4be44f6352e2 · outbound

This paper cites Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers.

GeoMM: On Geodesic Perspective for Multi-modal Learning Pixel-BERT: Aligning Image Pixels with Text by Deep Multi-Modal Transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.561792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.561792Z digest=sha256:ac30ef219ae826d18f639c5d81546bfbd09aef7bf27c3c3acb0bb6523a66399d

Observation 8ab911b9-f86b-4331-9caf-a74d29e77a59 · outbound

This paper cites Scaling up visual and vision-language representation learning with noisy text supervision.

GeoMM: On Geodesic Perspective for Multi-modal Learning Scaling up visual and vision-language representation learning with noisy text supervision

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.566170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.566170Z digest=sha256:2a20a1df60eac78df2656c9a0411ac89ad5330ce3dabb8236250867c2bd68b96

Observation 4da22129-6ec5-4438-9d0f-38cac98f3077 · outbound

This paper cites Vilt: Vision-and-language transformer without convolu- tion or region supervision.

GeoMM: On Geodesic Perspective for Multi-modal Learning Vilt: Vision-and-language transformer without convolu- tion or region supervision

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.570688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.570688Z digest=sha256:bdd0af7acc580750b51421c4bda11ca90bfbd12dcc28b880f7a0e30890fc7e31

Observation 43dd0a4b-2483-4364-90ad-ed7f8b51290b · outbound

This paper cites Computing geodesic paths on manifolds.

GeoMM: On Geodesic Perspective for Multi-modal Learning Computing geodesic paths on manifolds

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.574725Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.574725Z digest=sha256:f2bfbd539ed826e760eb10e538eba76964fa8e72063e63e2c7ed46e47a10e944

Observation 27c0734b-23d0-43c1-9612-5e6207f46290 · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annota- tions.

GeoMM: On Geodesic Perspective for Multi-modal Learning Visual genome: Connecting language and vision using crowdsourced dense image annota- tions

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.579261Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.579261Z digest=sha256:42e5121e271b7b2908663e026846f6b7712a12bba6d8708506150db2389c2e1e

Observation 7e43bedb-4f44-468c-bcf5-b3c469f4d285 · outbound

This paper cites Numba: a llvm-based python JIT compiler.

GeoMM: On Geodesic Perspective for Multi-modal Learning Numba: a llvm-based python JIT compiler

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.583852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.583852Z digest=sha256:8a8b58b48644eab16a82c05e1164e0f9a523c0e9da524a347d8936173e5eb6c4

Observation aa33debf-79a4-4188-b5a2-4d393a1c1ee8 · outbound

This paper cites Le, Vu Nguyen, Chen-Ping Yu, and Dimitris Samaras.

GeoMM: On Geodesic Perspective for Multi-modal Learning Le, Vu Nguyen, Chen-Ping Yu, and Dimitris Samaras

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.588469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.588469Z digest=sha256:45b64dabc90b8dd482af6fc57ac7397a5f97ea6ebc7be3613f98fe3c258a222a

Observation 46ea7ca2-464f-42cf-bda9-8f1d75c8caee · outbound

This paper cites Stacked cross attention for image- text matching.

GeoMM: On Geodesic Perspective for Multi-modal Learning Stacked cross attention for image- text matching

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.592895Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.592895Z digest=sha256:d4a0702f170be98fe20c974f3a65b5af92bc3f56ca1d2f51fcdf396b99bcd8c4

Observation 23f519ed-53c6-45ae-a6e6-f7b0681e5b9d · outbound

This paper cites Multimodal Foundation Models: From Specialists to General-Purpose Assistants.

GeoMM: On Geodesic Perspective for Multi-modal Learning Multimodal Foundation Models: From Specialists to General-Purpose Assistants

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.597292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.597292Z digest=sha256:e5829572c092d615ee68395301635e636272611d9d1ba31ccb77d6b09b118f31

Observation 71eddf66-8905-4e23-8355-6cd0d419bc8e · outbound

This paper cites Align before fuse: Vision and lan- guage representation learning with momentum dis- tillation.

GeoMM: On Geodesic Perspective for Multi-modal Learning Align before fuse: Vision and lan- guage representation learning with momentum dis- tillation

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.602455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.602455Z digest=sha256:1fbe3998a32a01b3c64dd49217787273c50fae6bd3edc83bfbb3776e8bc7b4be

Observation f3f7b385-b0a0-4dbb-88c2-1a9fd92507e0 · outbound

This paper cites Blip: Bootstrapping language-image pre- training for unified vision-language understanding and generation.

GeoMM: On Geodesic Perspective for Multi-modal Learning Blip: Bootstrapping language-image pre- training for unified vision-language understanding and generation

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:40.191284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.607159Z digest=sha256:292468ea1bb3f140648fa3205d4116981102b5aacaa6923115bbb7ccdf349f99

Observation 365af228-59ce-46e0-ae1c-7e91497c43c6 · outbound

This paper cites HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training.

GeoMM: On Geodesic Perspective for Multi-modal Learning HERO: Hierarchical Encoder for Video+Language Omni-representation Pre-training

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.611677Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.611677Z digest=sha256:5d0558c6239f0df955e8badf9e029b6aa26f52fbc6d99eb81f3225a11be47d0b

Observation 4cb98fbf-d047-4e6d-bfcd-bb8854344691 · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

GeoMM: On Geodesic Perspective for Multi-modal Learning VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.616515Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.616515Z digest=sha256:87139d772a391044bbbb7ece4e5d6af71e7f26def89934d74b90afe38728faf3

Observation 4d8b67ea-7b08-4c7a-89fb-c7d019169cfe · outbound

This paper cites Oscar: Object-semantics aligned pre-training for vision-language tasks.

GeoMM: On Geodesic Perspective for Multi-modal Learning Oscar: Object-semantics aligned pre-training for vision-language tasks

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:40.177109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.621318Z digest=sha256:556601cd5fe3197979324516de7755d1216ad2674c8a18f72dda41f0dc362c91

Observation b18e0303-3dcd-4727-b2b8-c843df5c473e · outbound

This paper cites Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm.

GeoMM: On Geodesic Perspective for Multi-modal Learning Supervision Exists Everywhere: A Data Efficient Contrastive Language-Image Pre-training Paradigm

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.625976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.625976Z digest=sha256:aa4a23e58fa46f358e6cf8f2c761a92e82b4f64acc56958612e16c8b569798ea

Observation c4cc866d-c672-4f5d-9e80-c0fa2e1bc243 · outbound

This paper cites Scaling language- image pre-training via masking.

GeoMM: On Geodesic Perspective for Multi-modal Learning Scaling language- image pre-training via masking

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:40.163177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.630755Z digest=sha256:d105982a782d132e3b006868e8b51b161667c32c753f68115f3d5d1b0fff85ce

Observation 9254e8d4-3858-4246-a97a-f3fa3252d5e7 · outbound

This paper cites Geodesic self- attention for 3d point clouds.

GeoMM: On Geodesic Perspective for Multi-modal Learning Geodesic self- attention for 3d point clouds

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:40.148944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.635022Z digest=sha256:f498c41e63ca2054279fba58e0d9c2e5557df015095ada9f60657da92af3d90c

Observation 06178822-4c25-4d71-a750-3cdcaa1e655d · outbound

This paper cites Microsoft coco: Common objects in context.

GeoMM: On Geodesic Perspective for Multi-modal Learning Microsoft coco: Common objects in context

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:40.134318Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.639316Z digest=sha256:eba9f3210f57d04f81c06ef9cd0050a772c7df91f12495b5f46e64be097deee7

Observation 60a9f332-f09a-4b03-abb3-693a13838ba4 · outbound

This paper cites an unresolved cited work.

GeoMM: On Geodesic Perspective for Multi-modal Learning Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:01:40.118132Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.644092Z digest=sha256:cc7914616834fd3e5ad632d982b757729dc24e9fe9004ff4a09ae4d15cc2f3da

Observation 39a9358a-fce3-4c28-a839-3c9deeedb799 · outbound

This paper cites Adap- tive reconstruction network for weakly supervised re- ferring expression grounding.

GeoMM: On Geodesic Perspective for Multi-modal Learning Adap- tive reconstruction network for weakly supervised re- ferring expression grounding

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:40.102936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.648566Z digest=sha256:ed456cce5d297b6988776b3f15ac70c4be8739a037ff0e7d4fac9809087941e0

Observation 2351d9c4-6ef2-45fb-8709-8d500a8fbb70 · outbound

This paper cites Algorithm as 136: A k-means clustering algorithm.

GeoMM: On Geodesic Perspective for Multi-modal Learning Algorithm as 136: A k-means clustering algorithm

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:40.088456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.652851Z digest=sha256:d9f764c886af3d7ba50a4bfb471e380bb9bce5f41a1e294c895d9da90c717878

Observation 4d014c8a-5557-4533-8c80-b64489b474eb · outbound

This paper cites Decoupled weight decay regularization.

GeoMM: On Geodesic Perspective for Multi-modal Learning Decoupled weight decay regularization

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:40.074084Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.657146Z digest=sha256:45c02cf8d51635dc83dac09d2a6ca8b1a583bc4d0423d8af14d7debc3719b8ab

Observation fe50d3e3-97e6-4ba1-989f-c0f4b8c4e66e · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic rep- resentations for vision-and-language tasks.Advances in Neural Information Processing Systems, 32, 2019.

GeoMM: On Geodesic Perspective for Multi-modal Learning Vilbert: Pretraining task-agnostic visiolinguistic rep- resentations for vision-and-language tasks.Advances in Neural Information Processing Systems, 32, 2019

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:40.059524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.661372Z digest=sha256:3546c6003fed06ab5d496efcddab352fbe710c76469066a255076d16aaf0a5a0

Observation 5246da00-04c3-45f3-8e9a-2f135549850b · outbound

This paper cites Computing geodesics on triangular meshes.Comput- ers & Graphics, 29(5):667–675, 2005.

GeoMM: On Geodesic Perspective for Multi-modal Learning Computing geodesics on triangular meshes.Comput- ers & Graphics, 29(5):667–675, 2005

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:40.042868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.665984Z digest=sha256:d97552bacea3ab92e380f1ef87785a0654c39c1beafdb8df8026b6f9c003d47e

Observation b5cd8895-692e-4312-9a36-eb02c979757c · outbound

This paper cites Bron- stein, and Pierre Vandergheynst.

GeoMM: On Geodesic Perspective for Multi-modal Learning Bron- stein, and Pierre Vandergheynst

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:40.027471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.670412Z digest=sha256:894f4b4ba7ba9d536e26bedcda93969b0102d010eef4debd9970f31557cc2893

Observation db0dae6f-1b92-491a-9eec-622b4be87742 · outbound

This paper cites Jensen’s inequality.

GeoMM: On Geodesic Perspective for Multi-modal Learning Jensen’s inequality

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:40.011409Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.674622Z digest=sha256:06f421e459bd3a757cee8d7b2d8f1ed43b2cc016466fdc37f1513eb0e6447d2d

Observation 53b6afc6-3b03-4001-99e5-aa01f5ddb3f2 · outbound

This paper cites Towards bridging sample complexity and model capacity.

GeoMM: On Geodesic Perspective for Multi-modal Learning Towards bridging sample complexity and model capacity

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.995610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.679359Z digest=sha256:83ec44cb6d955e742d68e82860b047b219fb281d0d3767097d0abe7847b768a6

Observation 037bcfe0-7300-451f-8c0e-90acd50e9d26 · outbound

This paper cites Towards interpreting and utiliz- ing symmetry property in adversarial examples.

GeoMM: On Geodesic Perspective for Multi-modal Learning Towards interpreting and utiliz- ing symmetry property in adversarial examples

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.980284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.684246Z digest=sha256:26b17a6ed330a7135b5e3f2d504aea57917ff53289cdbcc518e247d4fbd8af98

Observation eb72ad7e-e4f0-4cb7-ac55-d7d92f501aee · outbound

This paper cites Exploring and utilizing pattern imbal- ance.

GeoMM: On Geodesic Perspective for Multi-modal Learning Exploring and utilizing pattern imbal- ance

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.965557Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.688557Z digest=sha256:9f0d2643d8e68396143d79f730036c5c1e06ff7e443a152a1e47baebc6080360

Observation fc249f86-bb9e-4169-b426-df1f88ff3b4b · outbound

This paper cites MSSIDD: A Benchmark for Multi-Sensor Denoising.

GeoMM: On Geodesic Perspective for Multi-modal Learning MSSIDD: A Benchmark for Multi-Sensor Denoising

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.693131Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.693131Z digest=sha256:cdcf77aa2c5a957ededb9f9c1da530e87c3a2ffb7cea9967efbd41a4ecb86199

Observation 8667fbc0-cf70-4172-a709-7e11fe756667 · outbound

This paper cites Object- oriented anchoring and modal alignment in multi- modal learning.

GeoMM: On Geodesic Perspective for Multi-modal Learning Object- oriented anchoring and modal alignment in multi- modal learning

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.950715Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.697887Z digest=sha256:515ca4d612da8e816b5f2a28ce7b664921f42fccbe159555757d9874bde29c18

Observation ab70750a-7f52-41d8-b35d-9c13d9cf1988 · outbound

This paper cites an unresolved cited work.

GeoMM: On Geodesic Perspective for Multi-modal Learning Unresolved cited work

Reference 63

Resolution
unresolved
raw_fallback, observed 2026-08-15T21:01:39.936600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.702453Z digest=sha256:3b788ba7f40f953a6fe5e21f18398acbb516a12014b5594a3b9b2fa3751cf5d8

Observation dfa2bc74-bcff-4279-9090-48d22c9a014a · outbound

This paper cites Analytic inequalities.

GeoMM: On Geodesic Perspective for Multi-modal Learning Analytic inequalities

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.921950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.706836Z digest=sha256:87fca3c0185fa7d51ad19c3ba7baa2030b072dad8eebb7f5a6f02026a4d69912

Observation 779b1cb0-4418-4c9a-b5ee-c5bbc4d44e24 · outbound

This paper cites Slip: Self-supervision meets language- image pre-training.

GeoMM: On Geodesic Perspective for Multi-modal Learning Slip: Self-supervision meets language- image pre-training

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.907114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.711315Z digest=sha256:e24d9a6bb9db6a755a89f080d71d7a012ca12ea89d03ea1dea2332a755cdb65d

Observation fa827551-edcc-4c9e-9912-3714a5411439 · outbound

This paper cites Geodesic-former: A geodesic-guided few-shot 3d point cloud instance segmenter.

GeoMM: On Geodesic Perspective for Multi-modal Learning Geodesic-former: A geodesic-guided few-shot 3d point cloud instance segmenter

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.891302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.715810Z digest=sha256:d3a57438d6facb92cf9d698425845be14437d5086cd13fd21c37eef16e304b6e

Observation 365b3a40-91df-420d-8e36-a9d33cf78743 · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

GeoMM: On Geodesic Perspective for Multi-modal Learning Representation Learning with Contrastive Predictive Coding

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.720127Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.720127Z digest=sha256:77549e557bfc525ebe2531030aa17db5201b9ad59935092fe9f12f9f451a5eb6

Observation 02cfc387-40af-4776-aaca-1ce14fa83ce9 · outbound

This paper cites Im2text: Describing images using 1 million cap- tioned photographs.

GeoMM: On Geodesic Perspective for Multi-modal Learning Im2text: Describing images using 1 million cap- tioned photographs

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.874796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.724243Z digest=sha256:1dfec8056d2db08c16a120dbfdc0017b260a444df8face93f1140d57a6738278

Observation c050648a-b3e4-47df-b98c-50d09c10e2f9 · outbound

This paper cites Automatic differentiation in pytorch.

GeoMM: On Geodesic Perspective for Multi-modal Learning Automatic differentiation in pytorch

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.860351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.728504Z digest=sha256:17d4dfc6399ffdefc7214e1f49db0416c07790a527266d1cd7805e6e76763858

Observation cdca1922-b4c5-44fe-b249-9b550c72b70f · outbound

This paper cites BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers.

GeoMM: On Geodesic Perspective for Multi-modal Learning BEiT v2: Masked Image Modeling with Vector-Quantized Visual Tokenizers

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.732801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.732801Z digest=sha256:d583ff04cbf1b778c284dddaac6ca2fe626cc378afa2b7eb61aad28903655d7c

Observation 963c7e38-774e-441b-a3ab-26fa8a413f82 · outbound

This paper cites Computational optimal transport: With applications to data science.

GeoMM: On Geodesic Perspective for Multi-modal Learning Computational optimal transport: With applications to data science

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.845941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.737328Z digest=sha256:ea1e707bda6b9155f8edba2ceee6b29febb18074ae9de1da71c055f397669ff9

Observation b1c8fd83-b920-4b38-bd53-203dd2f7e424 · outbound

This paper cites Combined scaling for zero-shot transfer learn- ing.

GeoMM: On Geodesic Perspective for Multi-modal Learning Combined scaling for zero-shot transfer learn- ing

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.831296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.741852Z digest=sha256:19122c63630bf7400eac69035c0909b6dad9adf8b2a1074df5ac58d586948b85

Observation a45f9e51-ef3b-4c23-9477-b6dc223754f4 · outbound

This paper cites Flickr30k entities: Collecting region-to- phrase correspondences for richer image-to-sentence models.

GeoMM: On Geodesic Perspective for Multi-modal Learning Flickr30k entities: Collecting region-to- phrase correspondences for richer image-to-sentence models

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.816611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.746508Z digest=sha256:3098e71b94ff7b2b1254f48f202833c05ae8a98efb7620127aeb5cac8d212e83

Observation 3aa8b2b5-f6c8-432a-bfad-9389940629e9 · outbound

This paper cites Straightest geodesics on polyhedral surfaces.

GeoMM: On Geodesic Perspective for Multi-modal Learning Straightest geodesics on polyhedral surfaces

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.800960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.751231Z digest=sha256:758f46003a5c408548bcea155b07661e6f540d2cef5519a8fef82f3bc72266c3

Observation 69d4cb8f-1a1c-4d7f-beaf-ea39ec9da25a · outbound

This paper cites Graphwalks: Efficient shape agnostic geodesic shortest path estimation.

GeoMM: On Geodesic Perspective for Multi-modal Learning Graphwalks: Efficient shape agnostic geodesic shortest path estimation

Reference 75

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.786494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.755787Z digest=sha256:ec375307ffeb31cf69f5089f52e6ded414d08b6fbb9d296d0e015b2845866335

Observation 83ece59c-1713-4f08-9db8-d05488a96e43 · outbound

This paper cites ImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data.

GeoMM: On Geodesic Perspective for Multi-modal Learning ImageBERT: Cross-modal Pre-training with Large-scale Weak-supervised Image-Text Data

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.760337Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.760337Z digest=sha256:d72d6398c13ac2a2a8cf0ce6e0b8968f73e4926ec91ec493c873c20d39c13378

Observation b91a6b2b-8976-4bc4-9df4-028ed954f459 · outbound

This paper cites Improving language under- standing by generative pre-training.

GeoMM: On Geodesic Perspective for Multi-modal Learning Improving language under- standing by generative pre-training

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.771893Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.764854Z digest=sha256:560134a6db22f2e4cc9319c1342ec2fddca735893fd3851c840dc1169c293445

Observation cf238571-3c66-4790-a28d-b0db7bb9eb3c · outbound

This paper cites Learning transferable visual models from natural language supervision.

GeoMM: On Geodesic Perspective for Multi-modal Learning Learning transferable visual models from natural language supervision

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.756264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.770102Z digest=sha256:5379dbea12c7280a9b6422b8940569b5b19c2dfa945f728fbf75e8c7e236ca5f

Observation 34b842d2-b60f-445a-a601-22e99fb888e9 · outbound

This paper cites Faster r-cnn: Towards real-time object detection with region proposal networks.

GeoMM: On Geodesic Perspective for Multi-modal Learning Faster r-cnn: Towards real-time object detection with region proposal networks

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.739955Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.774550Z digest=sha256:5b255f7ea3fc1efd47b2748666c090e09f46d8e8d1e08e5a4896e558e85bc95c

Observation 8da1e150-ae73-490d-a9c7-3157e38dbdf1 · outbound

This paper cites Grad-cam: Visual explanations from deep networks via gradient-based localization.

GeoMM: On Geodesic Perspective for Multi-modal Learning Grad-cam: Visual explanations from deep networks via gradient-based localization

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.724753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.779246Z digest=sha256:88e6b1d1de6043e0f352645f0daf16a23e9359f06e6ac84bae33023ab1de0420

Observation 71a53bd5-3456-4d4e-aa00-1ceb46daee9c · outbound

This paper cites Conceptual captions: A cleaned, hy- pernymed, image alt-text dataset for automatic image captioning.

GeoMM: On Geodesic Perspective for Multi-modal Learning Conceptual captions: A cleaned, hy- pernymed, image alt-text dataset for automatic image captioning

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.709325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.783626Z digest=sha256:a5f128e592eb8d6437988d719a31121d6ffaa09f163ee78ab05d112aeecb4d59

Observation ef6e53a7-c2d1-4362-8650-182f8a060ede · outbound

This paper cites VL-BERT: Pre-training of Generic Visual-Linguistic Representations.

GeoMM: On Geodesic Perspective for Multi-modal Learning VL-BERT: Pre-training of Generic Visual-Linguistic Representations

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.788840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.788840Z digest=sha256:d324d40752ac6afc8932ddba361e0bd04924fa9945761dcf01e4ffd07eb5b87c

Observation 37bb562b-fa2e-443f-b0de-ce00af6ef2dd · outbound

This paper cites PandaGPT: One Model To Instruction-Follow Them All.

GeoMM: On Geodesic Perspective for Multi-modal Learning PandaGPT: One Model To Instruction-Follow Them All

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.793508Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.793508Z digest=sha256:48ccb0cb1234de8006cd513a0ac308c34689124557879349515a3a4f75526870

Observation 2d304298-9594-4627-bfbf-0b7daf9418fe · outbound

This paper cites A Corpus for Reasoning About Natural Language Grounded in Photographs.

GeoMM: On Geodesic Perspective for Multi-modal Learning A Corpus for Reasoning About Natural Language Grounded in Photographs

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.798297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.798297Z digest=sha256:dc6656fa2a998663cd082ce60aaaa10c00db7f9d21393621ab121657ac84c907

Observation 140a7beb-1be5-4a53-ad80-278df7e9ca50 · outbound

This paper cites Revisiting unreasonable effective- ness of data in deep learning era.

GeoMM: On Geodesic Perspective for Multi-modal Learning Revisiting unreasonable effective- ness of data in deep learning era

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.694704Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.802864Z digest=sha256:f64709c7b1534d97ef3602e3bd85e047a17ddc1637bdc5a33d629958767615fa

Observation 80e6a150-979d-4e34-96b5-8e289fb4060b · outbound

This paper cites Gortler, and Hugues Hoppe.

GeoMM: On Geodesic Perspective for Multi-modal Learning Gortler, and Hugues Hoppe

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.680072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.807626Z digest=sha256:c4251227d814f232431cf9c28ad0181fda5bb0a1c8194374d3696cbfee4a470d

Observation 71871314-e8da-4eb1-821a-1a4f8785f83e · outbound

This paper cites LXMERT: learning cross-modality encoder representations from trans- formers.

GeoMM: On Geodesic Perspective for Multi-modal Learning LXMERT: learning cross-modality encoder representations from trans- formers

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.664840Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.812216Z digest=sha256:04f54d44496fae01a493ae3eac92ddee57e6d9638953e049532e07e556495e8c

Observation 06bf8c74-6ea4-4cce-afc8-ea3c2782f5ee · outbound

This paper cites Tenenbaum, Vin de Silva, and John C.

GeoMM: On Geodesic Perspective for Multi-modal Learning Tenenbaum, Vin de Silva, and John C

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.650272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.816540Z digest=sha256:292f2a9aad42580dc09958f78d5214e8deb17e2507e84274fa0e88a89520ef01

Observation 65913de9-b905-4a15-910a-51ccf44688da · outbound

This paper cites Pigeon hole principle.

GeoMM: On Geodesic Perspective for Multi-modal Learning Pigeon hole principle

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.635830Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.821250Z digest=sha256:b637dd7384be8e39464f2e562dacd4006ee4918b866e4e03aa21e7f341459a06

Observation 403d64ff-a537-4552-8f49-4041e217dbac · outbound

This paper cites Attention is all you need.

GeoMM: On Geodesic Perspective for Multi-modal Learning Attention is all you need

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.621277Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.826010Z digest=sha256:17e2d29fe8ee9399ac9bb873fa5b3aa592d24b93b1f94609850a47ead1e74104

Observation db607275-2142-4af8-8dd8-14cb766e6385 · outbound

This paper cites Optimal transport: old and new.

GeoMM: On Geodesic Perspective for Multi-modal Learning Optimal transport: old and new

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.606340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.831030Z digest=sha256:100a8203ffd20429222708b7ea6e7ade621456102530716324c1ea03f3cf55b2

Observation 2dcf3c28-e99f-4ec3-ae7c-e8c33c20b304 · outbound

This paper cites Learning to combine: Knowledge aggrega- tion for multi-source domain adaptation.

GeoMM: On Geodesic Perspective for Multi-modal Learning Learning to combine: Knowledge aggrega- tion for multi-source domain adaptation

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.591726Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.835846Z digest=sha256:2731c97786e2ad9bd85ca5d0d36529f1dec2605bb588d206b48a1767b8413df5

Observation 7c0b0f61-e2b7-46d9-bb19-459480b1a9e1 · outbound

This paper cites Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks.

GeoMM: On Geodesic Perspective for Multi-modal Learning Image as a Foreign Language: BEiT Pretraining for All Vision and Vision-Language Tasks

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.840420Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.840420Z digest=sha256:ca9bbce3e6fce2c6edc16cf4b3be66aadb784b2b032723ae39fe89118f27a3a7

Observation 55b012a0-b38b-49f9-b321-51c4f2c794b8 · outbound

This paper cites Mvp: Multimodality-guided visual pre-training.

GeoMM: On Geodesic Perspective for Multi-modal Learning Mvp: Multimodality-guided visual pre-training

Reference 94

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.577128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.845368Z digest=sha256:da82aee7b7144ee2a4105fc116ae064827881ba4f0d6c6c5fd72e2bd384d9ae8

Observation deaa2c2d-f7ac-4aec-a500-d1e3e843a593 · outbound

This paper cites Visual Entailment: A Novel Task for Fine-Grained Image Understanding.

GeoMM: On Geodesic Perspective for Multi-modal Learning Visual Entailment: A Novel Task for Fine-Grained Image Understanding

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.849934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.849934Z digest=sha256:8118e8245f10ac6c26ed571ec27764201ba304e3c7b72655cc964e42897a62b7

Observation 35dd7e54-89fe-4315-8fb4-62c1fc13119e · outbound

This paper cites A fast proximal point method for computing exact wasserstein distance.

GeoMM: On Geodesic Perspective for Multi-modal Learning A fast proximal point method for computing exact wasserstein distance

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.562192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.854916Z digest=sha256:d167705f2dac38df0f90c2a383afe210f1c3065a8a907f21dd40fe6fbc26a567

Observation 9f1ffb73-0c4f-45eb-8907-460c0cd33c99 · outbound

This paper cites Vision-language pre- training with triple contrastive learning.

GeoMM: On Geodesic Perspective for Multi-modal Learning Vision-language pre- training with triple contrastive learning

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.547744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.859498Z digest=sha256:027bb59eabec4d52022b8986004146d226f2b88313731bb0cc4cdc05fec2a2fd

Observation 89478e07-0d26-4164-881d-72924a471f13 · outbound

This paper cites Unified contrastive learning in image-text-label space.

GeoMM: On Geodesic Perspective for Multi-modal Learning Unified contrastive learning in image-text-label space

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.533274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.864421Z digest=sha256:258e0374d0f9c06615c11af7cefcb2067ae221c5ed687500a304d0a66141ae51

Observation 957b594d-6f13-47fe-96d9-93b44a7a7eb9 · outbound

This paper cites FILIP: Fine-grained Interactive Language-Image Pre-Training.

GeoMM: On Geodesic Perspective for Multi-modal Learning FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.868708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.868708Z digest=sha256:706956e0616af90f87cc009af5271d9d2960b9616392442857cbcb9c82956c24

Observation 5c558f01-be97-4bf8-aa03-68cd271fc20d · outbound

This paper cites FILIP: fine-grained interactive language-image pre-training.

GeoMM: On Geodesic Perspective for Multi-modal Learning FILIP: fine-grained interactive language-image pre-training

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T21:01:39.518572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T21:01:38.873663Z digest=sha256:4b1ba5b70484dfb39457e3d7cae3829629ba819c07ce33fc5b27bdd0b19eeba6

Observation 132e1aec-03d9-42ba-a78b-d52284e19a45 · outbound

This paper cites CoCa: Contrastive Captioners are Image-Text Foundation Models.

GeoMM: On Geodesic Perspective for Multi-modal Learning CoCa: Contrastive Captioners are Image-Text Foundation Models

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-15T21:01:38.878311Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T21:01:38.878311Z digest=sha256:9c1fda9eb09334ee3b4d35e15b3235673836bdddd0c259301e54e8efe6c6cb8c

Pith citing papers

No inbound Pith citation observations are available.