Pith. sign in

Paper Citation Record · LEDGER

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space

As of 8 August 2026, this Paper Citation Record lists 56 of 56 outbound references and 0 inbound Pith citation observations for arXiv:2507.23188.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.23188 v1

Coverage vector

measured 56 of 56 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T11:07:20.190660Z

measured 56 of 56 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

56 of 56 outbound references displayed

  • verified exact0
  • verified fuzzy48
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c3d0d4c0-abe6-40c1-825e-40047a55204f · outbound

This paper cites Dual stream relation learning network for image-text retrieval,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Dual stream relation learning network for image-text retrieval,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:21.068747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:19.917710Z digest=sha256:83fe41d287fdad5e767dbbbd4584def475cbf4309ab9882201773f4d89689eaa

Observation 6d2d2a13-d0a9-4b2e-b8d1-c2da8f8b414c · outbound

This paper cites One-shot human motion transfer via occlusion-robust flow prediction and neural texturing,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space One-shot human motion transfer via occlusion-robust flow prediction and neural texturing,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:21.053234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:19.922972Z digest=sha256:1eaa60d68e2c2911f6eec60b837153719ba7a876dc61123bb3432de4053ed16d

Observation 7ceb4654-f82d-4be2-9590-087d0ae274dc · outbound

This paper cites Ta2v: Text-audio guided video generation,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Ta2v: Text-audio guided video generation,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:21.037886Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:19.927797Z digest=sha256:af314dc98976120a785dde6c49c2ee51abf1fc2bec5f9d1828149e44b5635501

Observation ac0208f8-bb09-4692-a9a2-d32d9ba81178 · outbound

This paper cites Cross-modal quantization for co-speech gesture generation,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Cross-modal quantization for co-speech gesture generation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:21.022470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:19.932734Z digest=sha256:a2fc4d19bc6c377e65e69fdc5cdaf98871a82a4b9778aee1094e6028ea7b54c4

Observation d7739b76-5e27-4f66-bc15-a47e715def06 · outbound

This paper cites Generative adversarial graph convolutional networks for human action synthesis,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Generative adversarial graph convolutional networks for human action synthesis,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:21.007627Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:19.937380Z digest=sha256:2f9c8e2ab88595bccbb8ad3c31e3f98ef9d395a9f2f33dd41cbad045f424f6ef

Observation 522a19a5-af03-4a68-9383-3fd6bc0db884 · outbound

This paper cites Action-conditioned 3d human motion synthesis with transformer vae,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Action-conditioned 3d human motion synthesis with transformer vae,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.992982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:19.942055Z digest=sha256:5c81a3e5fc6ce9696d219491371ad30a69cdbde2c3a08c3e748d757a61734a6d

Observation a865e48a-ab43-4969-8c7d-b0be2125ced4 · outbound

This paper cites Multiact: Long-term 3d human motion generation from multiple action labels,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Multiact: Long-term 3d human motion generation from multiple action labels,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.977904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:19.947230Z digest=sha256:232cc9d4a43e3f1c0f9f87eb930dfebb3573c821f489673eb33e03bb28b7a3b0

Observation 49b08ef6-91b7-48c6-ad88-c94983e8c4f1 · outbound

This paper cites Executing your commands via motion diffusion in latent space,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Executing your commands via motion diffusion in latent space,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.963097Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:19.951851Z digest=sha256:e1f000c8cb4603ec64fe9c8a28c310a5f152c00737b883c756d488e630935656

Observation 01d81e7a-032c-4844-8915-e6e7ae299068 · outbound

This paper cites The kit motion-language dataset,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space The kit motion-language dataset,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.947896Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:19.956741Z digest=sha256:3bb32d9919ad856ad03ac2c5c09d8d7431e97c116589d16dd1c6535b1c268902

Observation 095dfe23-4e48-47f8-9e90-2b6f8acdda8e · outbound

This paper cites Generating diverse and natural 3d human motions from text,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Generating diverse and natural 3d human motions from text,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.932758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:19.961155Z digest=sha256:95abd67ea822b74ec5e0a3e1ed5287edd1f2266172e6498fb17d4570311aa67f

Observation 227eaee2-97e9-449c-a7c2-4ab8269e7e2f · outbound

This paper cites Human motion diffusion model,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Human motion diffusion model,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.918269Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:19.965631Z digest=sha256:5b9da941a6c4d1b345af6464a58b383920f57c2f209406b492499121e6657580

Observation e0e44298-732d-4cc5-b0ba-bed362a7e0a1 · outbound

This paper cites Generating human motion from textual descriptions with discrete representations,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Generating human motion from textual descriptions with discrete representations,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.902943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:19.970336Z digest=sha256:7479e470a133754bef68bb1e74ac2da12ea85d6f6e25ad31100997cf6d3964f2

Observation 36c58672-c52c-481e-a62c-dc7cb2ad78ac · outbound

This paper cites Groupdancer: Music to multi-people dance synthesis with style collaboration,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Groupdancer: Music to multi-people dance synthesis with style collaboration,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.887444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:19.974905Z digest=sha256:906cf8c4b204492d4a44cd16a8f2807adec96dc03bd3e69e773c8c785464f6b4

Observation 743646df-9965-40b7-9484-63b14329913c · outbound

This paper cites Music- driven group choreography,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Music- driven group choreography,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.872221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:19.979443Z digest=sha256:f43184e51573c590e435f207b150e7e922038012f2c23f38e3158824ce4d0b8a

Observation c150045f-27fa-4875-a74a-f66c9598bc86 · outbound

This paper cites Edge: Editable dance generation from music,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Edge: Editable dance generation from music,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.856766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:19.984111Z digest=sha256:136ff2ab538ce982713810f2b07d625a53d5056a30ac7fa24a818239f1f59f7f

Observation d56ec5b9-7cc4-4c18-b93d-4e9ab369e531 · outbound

This paper cites Pc-dance: Posture- controllable music-driven dance synthesis,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Pc-dance: Posture- controllable music-driven dance synthesis,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.841247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:19.988790Z digest=sha256:6b4bc5452e6d591e451c10a3e753b44faf30ec5d8e055832e0ad4d0a0fdb8e13

Observation 0abead2e-3c61-4f25-aff1-009e29565489 · outbound

This paper cites Couch: Towards controllable human-chair interactions,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Couch: Towards controllable human-chair interactions,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.825215Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:19.993269Z digest=sha256:76b05d16c7fcc3f4cf1119414ea815cd1e6f43f5d5f7c459910155835b055761

Observation a509508f-d436-49c5-983f-9de104d75d41 · outbound

This paper cites Goal: Generating 4d whole-body motion for hand-object grasping,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Goal: Generating 4d whole-body motion for hand-object grasping,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.809687Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:19.997850Z digest=sha256:30915f8e3e60910fe0f91a4c141183b8abc544413f768cf0457f580823e0d89f

Observation 86592244-b176-4903-be80-4943ac117cde · outbound

This paper cites Human motion generation: A survey,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Human motion generation: A survey,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.795219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:20.003279Z digest=sha256:0db1a5f3262f662348df6916e68b656af6dc1a1f8fd95ab22f1015401f5dcb66

Observation 6e0bc659-6b8e-48e9-878f-c6b549566402 · outbound

This paper cites Phase-functioned neural networks for character control,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Phase-functioned neural networks for character control,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.780450Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:20.009756Z digest=sha256:77aa6769921d51f4c4fbae433d2fe64a6eb0309d0fdf58297c87c95cfbcf6b41

Observation 0b5a84df-0c83-4c98-bc86-5774b7ba02bb · outbound

This paper cites Learned motion matching,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Learned motion matching,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.765128Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:20.014500Z digest=sha256:af6a555ee2e17a2727d5c0ed68e160122454a19ce60db5da50bf02117c563d0b

Observation d0429251-a731-407a-bc5d-e9b12eb55dce · outbound

This paper cites Tmr: Text-to-motion retrieval using contrastive 3d human motion synthesis,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Tmr: Text-to-motion retrieval using contrastive 3d human motion synthesis,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.750358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:20.023014Z digest=sha256:f428abce6a93dda23873e0c1b12942c811cd2291fd2c643d371bf5e0b7251683

Observation 46e64a69-2467-4b76-8040-8d643eca10fd · outbound

This paper cites Tri-modal motion retrieval by learning a joint embedding space,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Tri-modal motion retrieval by learning a joint embedding space,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.734968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:20.028116Z digest=sha256:78655f6b8fa7ffd488a8aaf00e01648328ef7fb3fd31e34e3594537d618823b4

Observation af304b00-2e5b-41c4-a3b6-7ff024a44696 · outbound

This paper cites GPT-4 Technical Report.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space GPT-4 Technical Report

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T11:07:20.032813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:07:20.032813Z digest=sha256:27fe7a67c9f70783acd52c279515242e0603a081b17e8ba473d0177241cef265

Observation cbaf0f20-03b5-457b-ac6b-7415c8d21989 · outbound

This paper cites Wavlm: Large-scale self-supervised pre- training for full stack speech processing,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Wavlm: Large-scale self-supervised pre- training for full stack speech processing,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.716915Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:20.039109Z digest=sha256:391a9df6115d7187c30c8cecf3dc4a0c5ac29980f9ae632d4aeca016c82135a1

Observation 980f1468-0f6e-48d7-bd95-59f812a4ea20 · outbound

This paper cites Better speech synthesis through scaling.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Better speech synthesis through scaling

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T11:07:20.044506Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:07:20.044506Z digest=sha256:1de2c8a11bf00299d2f63603bc46a4f0344b390f82ed815832fea768dcee9b66

Observation 0e4cc8ae-f314-40d2-8d38-0f7fa9a0b957 · outbound

This paper cites Motionclip: Exposing human motion generation to clip space,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Motionclip: Exposing human motion generation to clip space,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.699757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:20.051712Z digest=sha256:828b82720bb16143e72bd1c1bff598d4e8b70f59aba3dee5eee2172d6cbd8d8b

Observation 5d148ccc-0ace-447d-a638-ecab485f4080 · outbound

This paper cites Auto-encoding variational bayes,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Auto-encoding variational bayes,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.684212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:20.059703Z digest=sha256:670821d64b33d009f532328c1f5f027883267e2df56cb0e05708c4dbc97438ec

Observation 9f74b422-15a8-4075-b93f-212138c3eea6 · outbound

This paper cites Temos: Generating diverse human motions from textual descriptions,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Temos: Generating diverse human motions from textual descriptions,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.669620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:20.064485Z digest=sha256:9c6e13b65ddf8f7ebe331be37caf36e242931d3f1645f2319e2e003ecdb151af

Observation a1d28c1c-9ff6-4dae-8c36-96221fdfd196 · outbound

This paper cites Tm2t: Stochastic and tokenized modeling for the reciprocal generation of 3d human motions and texts,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Tm2t: Stochastic and tokenized modeling for the reciprocal generation of 3d human motions and texts,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.655016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:20.069062Z digest=sha256:9b20766aefceadef25f54f1c99dfa1e70626a0d94cdd57575098729aabbeaf88

Observation cc001b1f-5078-4bf3-8b27-61411de93679 · outbound

This paper cites Neural discrete representation learning,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Neural discrete representation learning,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.639950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:20.073904Z digest=sha256:c21a83095716ca4916d0df0b5237004f041cbf6bffc65604060aec23a6ed0be0

Observation 7d1a1a47-e705-484a-aa53-034314937976 · outbound

This paper cites Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T11:07:20.079034Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:07:20.079034Z digest=sha256:694ec345783dea7cfaf0eccbbd83982d446a2ced41b5c9d7678681d39d1f03c4

Observation 6f22fbf2-7921-43ad-98df-5f01da09ceb6 · outbound

This paper cites Coca: Contrastive captioners are image-text foundation models,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Coca: Contrastive captioners are image-text foundation models,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.615160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:20.084405Z digest=sha256:01d4d8a8ddb70a88225a3084c73bf9286f9b9865b057b6861235508180299fb8

Observation 801e400f-3c78-4cc2-8d11-48671b98f329 · outbound

This paper cites Global meets local: Dual activation hashing network for large-scale fine-grained image retrieval,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Global meets local: Dual activation hashing network for large-scale fine-grained image retrieval,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.600313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:20.089655Z digest=sha256:98c1b9023e30a17a2aee1bd472f90a0c5668e974ac724856ab3e411091b0aa54

Observation 085ce84c-cda7-4bf3-a71f-3ed7e673b97b · outbound

This paper cites Dvf: Advancing robust and accurate fine-grained image retrieval with retrieval guidelines,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Dvf: Advancing robust and accurate fine-grained image retrieval with retrieval guidelines,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.585697Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:20.094537Z digest=sha256:3e28d273fb52ad94b2841d84d68caba6794325e1885c817079da8de4b444c71e

Observation 3ecb2ba1-991b-4672-ae3d-88e04a36b54b · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Bert: Pre-training of deep bidirectional transformers for language understanding,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.570573Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:20.099676Z digest=sha256:c2f5cc709fa668aa64252bb7a6e93e883cc86fe92cc5325c22a0152d20c69726

Observation 6d3d0b5d-d03c-4756-a1dc-c42c77076b7a · outbound

This paper cites Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.554167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:20.103973Z digest=sha256:283fba8abb565c209aa09aa9b88949545b631c2a1fd3c7033dca5e82c912ac75

Observation 982ee88c-1a82-46ff-b814-f8c92ab31047 · outbound

This paper cites Exploring the limits of transfer learning with a unified text-to-text transformer,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Exploring the limits of transfer learning with a unified text-to-text transformer,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.539447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:20.108193Z digest=sha256:5d4b01c8db59ad002561d0e71180abe96ed201e69cb1d27187e71c7c6f12daf7

Observation c0871aeb-5f63-4cab-af47-1f51f2579950 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space LLaMA: Open and Efficient Foundation Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T11:07:20.113044Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:07:20.113044Z digest=sha256:4efdc0d208e51f56e9f74b49a7b8d9dfdc8a5dc4d587fa349a842099bf9324c9

Observation b59dbc53-86c7-4293-8276-ac490fdaf30c · outbound

This paper cites Audiolm: a language modeling approach to audio generation,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Audiolm: a language modeling approach to audio generation,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.524041Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:20.117704Z digest=sha256:f2c44e5fe589663a675e9f930843c127f381711e541672cb9723139b7acaf12e

Observation 4f80dfef-9939-46f6-9a24-a32d0ec6e2e1 · outbound

This paper cites Robust speech recognition via large-scale weak super- vision,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Robust speech recognition via large-scale weak super- vision,

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.508969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:20.122266Z digest=sha256:58b6b2cef5ff85a90b3a3ba860bf6fc4442d7cd08265ff123c2015c569db38fe

Observation 4660fb3d-fdf7-40e6-89ae-fe74d8854c43 · outbound

This paper cites SUPERB: Speech processing Universal PERformance Benchmark.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space SUPERB: Speech processing Universal PERformance Benchmark

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-06T11:07:20.126818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:07:20.126818Z digest=sha256:c74ba9a571859d6c5c08ea8dceb121ed2d83c76df34565b5676389db8cec7ea9

Observation dc30a259-434d-4211-a537-8b4b5270b804 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.494265Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:20.131594Z digest=sha256:131b13052df3c6e8a75c63d33e7e925fa66786f80464186357f22cf62a79f3b5

Observation d449bd73-c834-4a0d-80f1-de423635cec1 · outbound

This paper cites Learning transferable visual models from natural language supervision,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Learning transferable visual models from natural language supervision,

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-06T11:07:20.136841Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:07:20.136841Z digest=sha256:ca24ea219eeb9ce365681e74fdcf973dbad8481442e174fb27728b53313735a9

Observation ae742115-187d-4977-aa0a-de9c99ed3517 · outbound

This paper cites Videopoet: A large language model for zero-shot video generation,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Videopoet: A large language model for zero-shot video generation,

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.468393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:20.141428Z digest=sha256:bc0cc37a15b3a105ab95590de264a66f5aaf6b708e2bfc01bba05a0ac8017e5f

Observation 1b317a51-829d-4de2-b82a-963fb5ab2e10 · outbound

This paper cites Label independent memory for semi-supervised few-shot video classification,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Label independent memory for semi-supervised few-shot video classification,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.453186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:20.146030Z digest=sha256:1086055f56eefcbe257e64c9bdd29f0c251b482a2806a0554df6026a5984e244

Observation 801c818e-be16-4958-8cb8-7cc96a45afc4 · outbound

This paper cites Memory-enhanced transformer for representation learning on temporal heterogeneous graphs,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Memory-enhanced transformer for representation learning on temporal heterogeneous graphs,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.438474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:20.150446Z digest=sha256:b771b0f1f27c07d21a9ad093edca86e65b65ff94a45b1c683fd8d480d3222473

Observation 392e1a5d-4b2b-4ed8-b666-15d909101a2c · outbound

This paper cites An efficient memory module for graph few-shot class-incremental learning,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space An efficient memory module for graph few-shot class-incremental learning,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.424127Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:20.155284Z digest=sha256:1db955db32ebdbb00fd63e2b953bd84ecd5b614650b936ae1cea1e3b23b55e9a

Observation 5b0982a6-d2e9-414b-bb65-99388e515b22 · outbound

This paper cites Imagebind: One embedding space to bind them all,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Imagebind: One embedding space to bind them all,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.408819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:20.159521Z digest=sha256:58dfb39942bc5ba675b765d5cc553e826cafd1c291672c9218468d202e8e10dc

Observation a4b8a03c-d931-481a-88e3-c217fb40e66d · outbound

This paper cites Grounded language-image pre- training,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Grounded language-image pre- training,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.392196Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:20.163885Z digest=sha256:07a2dfb77ca95fe4fa749154af1a89ce39726f928a73955b273f52a83929676a

Observation 0e899e0e-25bf-4e3e-ab1b-d6b312477551 · outbound

This paper cites ActionCLIP: A New Paradigm for Video Action Recognition.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space ActionCLIP: A New Paradigm for Video Action Recognition

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-06T11:07:20.168158Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:07:20.168158Z digest=sha256:0ddc3155cf4c88b226cb3eb3f8d47c208fa21a3cf3afed8f2c72dc2a35876d94

Observation 026a9884-f216-416a-b701-5d015e57fc77 · outbound

This paper cites Delving into multimodal prompting for fine-grained visual classifica- tion,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Delving into multimodal prompting for fine-grained visual classifica- tion,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.377514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:20.173210Z digest=sha256:a6ff0adea2fc96ae489bd29296e0136884a4e0a284d2d74886bc7e1699d9a4a7

Observation 426b2e26-f8d3-4a76-b4a3-3f56eee89156 · outbound

This paper cites Amass: Archive of motion capture as surface shapes,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Amass: Archive of motion capture as surface shapes,

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.361998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:20.177499Z digest=sha256:72014276d8dfe0701f25e5e8c392767bda81c135de98013950277d3866d90f7b

Observation 7ef95020-8e13-4f9c-9011-b7434957e4bd · outbound

This paper cites Action2motion: Conditioned generation of 3d human motions,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Action2motion: Conditioned generation of 3d human motions,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.347117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:20.181731Z digest=sha256:18ac5d14c3032ee93e4e1408f546c1f2eaff2b512ea60eb4eae718734b092bd6

Observation 0f038371-bae1-4e37-9541-80b9b3c4ff3d · outbound

This paper cites Representation Learning with Contrastive Predictive Coding.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Representation Learning with Contrastive Predictive Coding

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T11:07:20.186032Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T11:07:20.186032Z digest=sha256:0a707706e21a7009ef07b544c873cc3cd2f47a71687842b40e6fb95ed346c724

Observation b1a739eb-9df3-4685-bf1f-f5100fa96c4b · outbound

This paper cites Randaugment: Practical automated data augmentation with a reduced search space,.

Multi-Modal Motion Retrieval by Learning a Fine-Grained Joint Embedding Space Randaugment: Practical automated data augmentation with a reduced search space,

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T11:07:20.331762Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-06T11:07:20.190660Z digest=sha256:3ff9c1a2bce0478c50d928f9dfc956cd1ba299ba8a408a151d97b48739b4cd02

Pith citing papers

No inbound Pith citation observations are available.