Pith. sign in

Paper Citation Record · LEDGER

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space

As of 18 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2507.23188.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.23188 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:07:20.190660Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy48
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c3d0d4c0-abe6-40c1-825e-40047a55204f · outbound

This paper cites Dual stream relation learning network for image-text retrieval,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Dual stream relation learning network for image-text retrieval,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:21.068747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:19.917710Z digest=sha256:3a82b303862d9b3583a5bd1f818d9aee92247897a800881f18290c638c7dce23

Observation 6d2d2a13-d0a9-4b2e-b8d1-c2da8f8b414c · outbound

This paper cites One-shot human motion transfer via occlusion-robust flow prediction and neural texturing,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space One-shot human motion transfer via occlusion-robust flow prediction and neural texturing,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:21.053234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:19.922972Z digest=sha256:b06de49e2b759beb74cc2bfc1009dcf84b8c13900c8ae9867324e5d40b3cecc6

Observation 7ceb4654-f82d-4be2-9590-087d0ae274dc · outbound

This paper cites Ta2v: Text-audio guided video generation,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Ta2v: Text-audio guided video generation,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:21.037886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:19.927797Z digest=sha256:49a1e7b51f27e25621339b4b375a21e9dae0090bf405e1efbcc36ec21dab900f

Observation ac0208f8-bb09-4692-a9a2-d32d9ba81178 · outbound

This paper cites Cross-modal quantization for co-speech gesture generation,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Cross-modal quantization for co-speech gesture generation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:21.022470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:19.932734Z digest=sha256:740dd90fc32b20ef9dd538676e29bcec850419379ebefd6d553633ddc653e14e

Observation d7739b76-5e27-4f66-bc15-a47e715def06 · outbound

This paper cites Generative adversarial graph convolutional networks for human action synthesis,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Generative adversarial graph convolutional networks for human action synthesis,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:21.007627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:19.937380Z digest=sha256:aca1b97824a1cba478e5fed32e185bc79da3ef143970734883acd1d7d876710d

Observation 522a19a5-af03-4a68-9383-3fd6bc0db884 · outbound

This paper cites Action-conditioned 3d human motion synthesis with transformer vae,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Action-conditioned 3d human motion synthesis with transformer vae,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.992982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:19.942055Z digest=sha256:603c9592b21899aade04cdbce3b9fa3ce7d557eb62764d3e53e99dfca2484527

Observation a865e48a-ab43-4969-8c7d-b0be2125ced4 · outbound

This paper cites Multiact: Long-term 3d human motion generation from multiple action labels,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Multiact: Long-term 3d human motion generation from multiple action labels,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.977904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:19.947230Z digest=sha256:034a9bd07c0b5f789a253426a0c48332eb6d0d3883108e9837d868eb62040b3d

Observation 49b08ef6-91b7-48c6-ad88-c94983e8c4f1 · outbound

This paper cites Executing your commands via motion diffusion in latent space,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Executing your commands via motion diffusion in latent space,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.963097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:19.951851Z digest=sha256:1798fc0a4a9007c6fb728d01c2a1c2a3ee2a34933aced1dd80d91b7298a9ce58

Observation 01d81e7a-032c-4844-8915-e6e7ae299068 · outbound

This paper cites The kit motion-language dataset,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space The kit motion-language dataset,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.947896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:19.956741Z digest=sha256:af9ccc6bd3876325790f2d5e5fc07f17777b31f752019a49a91a7ae02942c494

Observation 095dfe23-4e48-47f8-9e90-2b6f8acdda8e · outbound

This paper cites Generating diverse and natural 3d human motions from text,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Generating diverse and natural 3d human motions from text,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.932758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:19.961155Z digest=sha256:66009d78095869f0b18a2116abeed25c4d68f25f22c976ad112717b2336f7059

Observation 227eaee2-97e9-449c-a7c2-4ab8269e7e2f · outbound

This paper cites Human motion diffusion model,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Human motion diffusion model,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.918269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:19.965631Z digest=sha256:df9ffc04c677047a160ce169e667cd8504b5b328617092498e171c9a66de06aa

Observation e0e44298-732d-4cc5-b0ba-bed362a7e0a1 · outbound

This paper cites Generating human motion from textual descriptions with discrete representations,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Generating human motion from textual descriptions with discrete representations,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.902943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:19.970336Z digest=sha256:5468d6b3456233ea1377af639457922846e4e8379f056a0d0619b4b58e944681

Observation 36c58672-c52c-481e-a62c-dc7cb2ad78ac · outbound

This paper cites Groupdancer: Music to multi-people dance synthesis with style collaboration,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Groupdancer: Music to multi-people dance synthesis with style collaboration,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.887444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:19.974905Z digest=sha256:e1311f286af1bfc2a6272300648bf9d399494385c1382e2ed208d5ac8ccb94c6

Observation 743646df-9965-40b7-9484-63b14329913c · outbound

This paper cites Music- driven group choreography,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Music- driven group choreography,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.872221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:19.979443Z digest=sha256:9bc8d4fd8d1c6c835c5d96e5684a280a1ddfa56a55ab181634ae6a11e76147a2

Observation c150045f-27fa-4875-a74a-f66c9598bc86 · outbound

This paper cites Edge: Editable dance generation from music,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Edge: Editable dance generation from music,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.856766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:19.984111Z digest=sha256:0b3072b81597d160fc5658c17689b8b6f2ea47256fb4dbf3b9eab40d1de99d04

Observation d56ec5b9-7cc4-4c18-b93d-4e9ab369e531 · outbound

This paper cites Pc-dance: Posture- controllable music-driven dance synthesis,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Pc-dance: Posture- controllable music-driven dance synthesis,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.841247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:19.988790Z digest=sha256:15fd808e5a6fc5c70b0e00e2226a603f0cff978941c2949108bbdf669eecb6dc

Observation 0abead2e-3c61-4f25-aff1-009e29565489 · outbound

This paper cites Couch: Towards controllable human-chair interactions,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Couch: Towards controllable human-chair interactions,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.825215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:19.993269Z digest=sha256:efb14d70e1dcd217e354db7df122377d9bee78e4a87cd1da454d78461423983f

Observation a509508f-d436-49c5-983f-9de104d75d41 · outbound

This paper cites Goal: Generating 4d whole-body motion for hand-object grasping,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Goal: Generating 4d whole-body motion for hand-object grasping,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.809687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:19.997850Z digest=sha256:962dc41446b30896a1092158f7ea30906bce73cd6aacefc58a3135f9f99d2dae

Observation 86592244-b176-4903-be80-4943ac117cde · outbound

This paper cites Human motion generation: A survey,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Human motion generation: A survey,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.795219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:20.003279Z digest=sha256:7a86eabd7f7dd2aecd1c0f5b93545efc5cdc2039228edafa8bd2765346603b0c

Observation 6e0bc659-6b8e-48e9-878f-c6b549566402 · outbound

This paper cites Phase-functioned neural networks for character control,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Phase-functioned neural networks for character control,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.780450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:20.009756Z digest=sha256:8519dcaa3a8d6bcb616590cb3c198ee3561418a098af4546e82e416df32f7a1b

Observation 0b5a84df-0c83-4c98-bc86-5774b7ba02bb · outbound

This paper cites Learned motion matching,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Learned motion matching,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.765128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:20.014500Z digest=sha256:b7652fc34d3b0d23955c439c360d320f50f2b1e24f9b1b7d52e1e2f5b509c5d9

Observation d0429251-a731-407a-bc5d-e9b12eb55dce · outbound

This paper cites Tmr: Text-to-motion retrieval using contrastive 3d human motion synthesis,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Tmr: Text-to-motion retrieval using contrastive 3d human motion synthesis,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.750358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:20.023014Z digest=sha256:5bb9b64f2354d3e48b63b5f35389b098a3aee24565504f312296482f1d0712a0

Observation 46e64a69-2467-4b76-8040-8d643eca10fd · outbound

This paper cites Tri-modal motion retrieval by learning a joint embedding space,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Tri-modal motion retrieval by learning a joint embedding space,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.734968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:20.028116Z digest=sha256:a08c34e41eda6669d336ef0daed2653af0c53686bd48046717b790464973ac35

Observation af304b00-2e5b-41c4-a3b6-7ff024a44696 · outbound

This paper cites GPT-4 Technical Report.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space GPT-4 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T11:07:20.032813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:07:20.032813Z digest=sha256:62e1691b749c7794a5cc4f916909352b2b0a6132b7466c97efb57920e0cabab6

Observation cbaf0f20-03b5-457b-ac6b-7415c8d21989 · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.716915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:20.039109Z digest=sha256:d9cba4323949dd6bc0f83761baf7a87414cc6ab9d633414329396a7e33195550

Observation 980f1468-0f6e-48d7-bd95-59f812a4ea20 · outbound

This paper cites Better speech synthesis through scaling.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Better speech synthesis through scaling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T11:07:20.044506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:07:20.044506Z digest=sha256:ad6fb5c4ab5ba89274c8810a57244ded46949c10f8f3620ec3272a96f7d320bc

Observation 0e4cc8ae-f314-40d2-8d38-0f7fa9a0b957 · outbound

This paper cites Motionclip: Exposing human motion generation to clip space,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Motionclip: Exposing human motion generation to clip space,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.699757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:20.051712Z digest=sha256:747c1052549430f11abb4df80540d6742d0e561b3bd8dd7fc1aa232367c63e9e

Observation 5d148ccc-0ace-447d-a638-ecab485f4080 · outbound

This paper cites Auto-encoding variational bayes,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Auto-encoding variational bayes,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.684212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:20.059703Z digest=sha256:90dc248ed1feff453df80760846ad7e4fadb53c094e5675a605880b356c7b71e

Observation 9f74b422-15a8-4075-b93f-212138c3eea6 · outbound

This paper cites Temos: Generating diverse human motions from textual descriptions,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Temos: Generating diverse human motions from textual descriptions,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.669620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:20.064485Z digest=sha256:e7a47bc6b722ff418f4c8c5d4f54caea2ea035ee4abb513b7503508e9849c8db

Observation a1d28c1c-9ff6-4dae-8c36-96221fdfd196 · outbound

This paper cites Tm2t: Stochastic and tokenized modeling for the reciprocal generation of 3d human motions and texts,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Tm2t: Stochastic and tokenized modeling for the reciprocal generation of 3d human motions and texts,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.655016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:20.069062Z digest=sha256:23a7ee1f07c475be0647b2099351392d26f0b28d44b960d2ad075eef8276d5bc

Observation cc001b1f-5078-4bf3-8b27-61411de93679 · outbound

This paper cites Neural discrete representation learning,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Neural discrete representation learning,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.639950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:20.073904Z digest=sha256:1c4a6cd69f4174ae77bcca062c098b1b1a9d36dd727ce626ee4c058c06ecc01f

Observation 7d1a1a47-e705-484a-aa53-034314937976 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T11:07:20.079034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:07:20.079034Z digest=sha256:1c33f6f7e332eeed9cb2a335963394835f6869ca9ff21e7981c6e3db19d74d41

Observation 6f22fbf2-7921-43ad-98df-5f01da09ceb6 · outbound

This paper cites Coca: Contrastive captioners are image-text foundation models,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Coca: Contrastive captioners are image-text foundation models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.615160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:20.084405Z digest=sha256:8266e27b9d0953dd24edaf78b098e121c431ae42d368d0a807578c8cd27fcf83

Observation 801e400f-3c78-4cc2-8d11-48671b98f329 · outbound

This paper cites Global meets local: Dual activation hashing network for large-scale fine-grained image retrieval,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Global meets local: Dual activation hashing network for large-scale fine-grained image retrieval,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.600313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:20.089655Z digest=sha256:f4f33953b73809013aa1e4966e285ce8c5b46a8392ac2f2ca43a22a237fb6872

Observation 085ce84c-cda7-4bf3-a71f-3ed7e673b97b · outbound

This paper cites Dvf: Advancing robust and accurate fine-grained image retrieval with retrieval guidelines,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Dvf: Advancing robust and accurate fine-grained image retrieval with retrieval guidelines,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.585697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:20.094537Z digest=sha256:e68e05e367c43c4eab5feaa81c63e1b9139b8b483472c404ad54512db8c869da

Observation 3ecb2ba1-991b-4672-ae3d-88e04a36b54b · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Bert: Pre-training of deep bidirectional transformers for language understanding,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.570573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:20.099676Z digest=sha256:b3f556b390d9427b7cd04b98bd82eacc31ed05d1ba8de8de16cd9c430c797be1

Observation 6d3d0b5d-d03c-4756-a1dc-c42c77076b7a · outbound

This paper cites Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.554167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:20.103973Z digest=sha256:368da984d435a255b1be7081bc3ffdecb1b57c456ae1ace62817ba6abaf5b8c0

Observation 982ee88c-1a82-46ff-b814-f8c92ab31047 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.539447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:20.108193Z digest=sha256:dbe291c2fc0edfb04e1699bac6f05d8cab2729031871c8c8ed51c7ef8f384e6f

Observation c0871aeb-5f63-4cab-af47-1f51f2579950 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space LLaMA: Open and Efficient Foundation Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T11:07:20.113044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:07:20.113044Z digest=sha256:a6eeb682c6b52f48ac020a5566cf65b2689a2dd6133e3bf99e992722fe012539

Observation b59dbc53-86c7-4293-8276-ac490fdaf30c · outbound

This paper cites Audiolm: a language modeling approach to audio generation,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Audiolm: a language modeling approach to audio generation,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.524041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:20.117704Z digest=sha256:05b7b544c8da279a6da60b74251ffff87aeab378db2e56875a3c42e042d1c20b

Observation 4f80dfef-9939-46f6-9a24-a32d0ec6e2e1 · outbound

This paper cites Robust speech recognition via large-scale weak super- vision,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Robust speech recognition via large-scale weak super- vision,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.508969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:20.122266Z digest=sha256:10adcd0c2e46c1a116a73a11380c7553229d4d7c81a20e4225d0ad34891af0e8

Observation 4660fb3d-fdf7-40e6-89ae-fe74d8854c43 · outbound

This paper cites SUPERB: Speech processing Universal PERformance Benchmark.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space SUPERB: Speech processing Universal PERformance Benchmark

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T11:07:20.126818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:07:20.126818Z digest=sha256:1965c8889e793f549a15d17e599e157ddae6ae3847e060a7405ff0743b396a91

Observation dc30a259-434d-4211-a537-8b4b5270b804 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.494265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:20.131594Z digest=sha256:444fda1fedffd1ee9617e1034e55ff1a93d626ccd9320f0f725a75d12a8f343f

Observation d449bd73-c834-4a0d-80f1-de423635cec1 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Learning transferable visual models from natural language supervision,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T11:07:20.136841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:07:20.136841Z digest=sha256:06ecbbf8984bcae5b5e8b7f3dd13009b83009d5e7fca6c71622598eedb255a26

Observation ae742115-187d-4977-aa0a-de9c99ed3517 · outbound

This paper cites Videopoet: A large language model for zero-shot video generation,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Videopoet: A large language model for zero-shot video generation,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.468393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:20.141428Z digest=sha256:3f876bf71906b3c269201e1aafe9f3b1eca119e0753eed8b8707887a714cf3a5

Observation 1b317a51-829d-4de2-b82a-963fb5ab2e10 · outbound

This paper cites Label independent memory for semi-supervised few-shot video classification,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Label independent memory for semi-supervised few-shot video classification,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.453186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:20.146030Z digest=sha256:3a04020016859cd616875469ffb25f1cc1b70ed2cba4b50bb6e2e43bf659d297

Observation 801c818e-be16-4958-8cb8-7cc96a45afc4 · outbound

This paper cites Memory-enhanced transformer for representation learning on temporal heterogeneous graphs,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Memory-enhanced transformer for representation learning on temporal heterogeneous graphs,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.438474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:20.150446Z digest=sha256:68df13cc20222607b25864ff7c7beadcc702cb42aeabe86aee724330420fb113

Observation 392e1a5d-4b2b-4ed8-b666-15d909101a2c · outbound

This paper cites An efficient memory module for graph few-shot class-incremental learning,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space An efficient memory module for graph few-shot class-incremental learning,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.424127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:20.155284Z digest=sha256:d4447ed18d700c9bd3019e4428fee4b6bb405c33a21834f11c0b0b944dd5178e

Observation 5b0982a6-d2e9-414b-bb65-99388e515b22 · outbound

This paper cites Imagebind: One embedding space to bind them all,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Imagebind: One embedding space to bind them all,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.408819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:20.159521Z digest=sha256:7d288ba058ae2841c9be7efa60ce140905460e72de7dcb832b14f402a2b824ec

Observation a4b8a03c-d931-481a-88e3-c217fb40e66d · outbound

This paper cites Grounded language-image pre- training,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Grounded language-image pre- training,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.392196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:20.163885Z digest=sha256:e25755b39dbfa6863a9bf787e8b85feafb47b7a9b9626c2e68c62cedfdd525a9

Observation 0e899e0e-25bf-4e3e-ab1b-d6b312477551 · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space ActionCLIP: A New Paradigm for Video Action Recognition

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T11:07:20.168158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:07:20.168158Z digest=sha256:9a747aa5c2ecba8fb47794c5cfaa36b89a9e68cb09c673a21ef294fe2068742f

Observation 026a9884-f216-416a-b701-5d015e57fc77 · outbound

This paper cites Delving into multimodal prompting for fine-grained visual classifica- tion,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Delving into multimodal prompting for fine-grained visual classifica- tion,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.377514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:20.173210Z digest=sha256:6128fc26b24807ce285dcf604785c3aa9ec717dcdd73bf52b64185b9c7e26042

Observation 426b2e26-f8d3-4a76-b4a3-3f56eee89156 · outbound

This paper cites Amass: Archive of motion capture as surface shapes,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Amass: Archive of motion capture as surface shapes,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.361998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:20.177499Z digest=sha256:6d1b7ae2a21bb8c63060bb09d7358ecff654cebccd205d400f9f89ff49f69d19

Observation 7ef95020-8e13-4f9c-9011-b7434957e4bd · outbound

This paper cites Action2motion: Conditioned generation of 3d human motions,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Action2motion: Conditioned generation of 3d human motions,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.347117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:20.181731Z digest=sha256:a619a264d441ac65547d7595d9abab1a54ebb0cd0bff37d11d18f91def4bbe3f

Observation 0f038371-bae1-4e37-9541-80b9b3c4ff3d · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Representation Learning with Contrastive Predictive Coding

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T11:07:20.186032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:07:20.186032Z digest=sha256:9b292c169edd6257a9a1e598e985d265a192f0b8ed28144dda7fa44a3dee0063

Observation b1a739eb-9df3-4685-bf1f-f5100fa96c4b · outbound

This paper cites Randaugment: Practical automated data augmentation with a reduced search space,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Randaugment: Practical automated data augmentation with a reduced search space,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.331762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.

source=pdf_text observed=2026-08-06T11:07:20.190660Z digest=sha256:5a51a3caae556493eee4fa74b1cef98eb0430fb1ce3346091387541c04b5df63

Pith citing papers

No inbound Pith citation observations are available.