Pith. sign in

Paper Citation Record · LEDGER

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision

As of 10 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 0 inbound Pith citation observations for arXiv:2608.01392.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2608.01392 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T00:18:12.550250Z

measured 45 of 45 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

45 of 45 outbound references displayed

  • verified exact0
  • verified fuzzy28
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1d8eb615-bb3a-4df3-ba8c-b07aaedb640b · outbound

This paper cites Vivit: A video vision transformer.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Vivit: A video vision transformer

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:09.624279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:09.624279Z digest=sha256:416a12d21e085840c2e20f2a437390a1202f4ae8b6e5978d7687cb56d7db7629

Observation 1e3ab1af-bc8b-4f09-ae93-eefb8c2ba092 · outbound

This paper cites Seamless human motion composition with blended posi- tional encodings.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Seamless human motion composition with blended posi- tional encodings

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:18.524741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:18:09.694111Z digest=sha256:828d4d67a3a32a7dbde95c57fd8f26667612d8c0b1bc4bbf20682c7e32e708be

Observation 729720f7-9dcb-4e37-bb4f-45cc441c4bbc · outbound

This paper cites A cross- dataset study for text-based 3d human motion retrieval.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision A cross- dataset study for text-based 3d human motion retrieval

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:18.350693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:18:09.761688Z digest=sha256:ab79d2057870ed81aeff08e757423bf1b1208f05cbf0146ad9626b16f78633a4

Observation 3f7840b3-31ac-44ba-a1e3-a313f6891a9d · outbound

This paper cites Is space-time attention all you need for video understanding? InIcml, page 4, 2021.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Is space-time attention all you need for video understanding? InIcml, page 4, 2021

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:09.806437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:09.806437Z digest=sha256:de7d444051c6c5c9cfc8e3056817f51351a840a30686348f78e88a1e9984c95e

Observation 110b8fd9-c88c-4d7b-a86f-4d4abafa473c · outbound

This paper cites Sinkhorn distances: Lightspeed computation of optimal transport.Advances in neural information pro- cessing systems, 26, 2013.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Sinkhorn distances: Lightspeed computation of optimal transport.Advances in neural information pro- cessing systems, 26, 2013

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:18.086343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:18:09.916222Z digest=sha256:232b2a88dc204f077112365e98b6c9aa167710902224450da7882fd00f067907

Observation e1df4d46-a53f-4f95-8e51-83ea6a6c04b6 · outbound

This paper cites Segmo: Segment-aligned text to 3d human motion generation.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Segmo: Segment-aligned text to 3d human motion generation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:17.845673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:18:10.011811Z digest=sha256:73d6288c72e7b4bbf0652a6295e895b4bd9fd7d8cdc80aeba1880362b3e34f92

Observation 2d8e775e-2337-449d-bacd-e94edbb17081 · outbound

This paper cites Masked autoencoders as spatiotemporal learners.Advances in neural information processing systems, 35:35946–35958,.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Masked autoencoders as spatiotemporal learners.Advances in neural information processing systems, 35:35946–35958,

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:10.101270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:10.101270Z digest=sha256:f66c9d1a6470dc1dfb40b3b46f5db7378fb9c4dbfc5def7be27b5681ff7421c5

Observation 9396af1c-2063-4a0b-af53-c00773747281 · outbound

This paper cites Generating diverse and natural 3d human motions from text.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Generating diverse and natural 3d human motions from text

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:17.689155Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:18:10.184506Z digest=sha256:631234d52537279179806c39b0c4bc66a13ae8922a82dc52ebf1c57734f86a2e

Observation a814d231-cf1b-4a36-a09b-b2ee7c0b31b5 · outbound

This paper cites Momask: Generative masked model- ing of 3d human motions.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Momask: Generative masked model- ing of 3d human motions

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:17.448771Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:18:10.252169Z digest=sha256:7978074e86ba0113d2fdeded40b743899fd8bb72bb25013a93b31ae86a20a9f8

Observation 92145b61-a539-4f5f-bc8d-c4f3fe4662e8 · outbound

This paper cites Snapmogen: Human motion generation from expressive texts.arXiv preprint arXiv:2507.09122, 2025.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Snapmogen: Human motion generation from expressive texts.arXiv preprint arXiv:2507.09122, 2025

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:10.316066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:10.316066Z digest=sha256:565d6e5efabd5421245bd8872b7ac55c500d0215fa0ff2a72e265db2817c195e

Observation dd972f82-f165-447b-93fd-bb7817a4e9fc · outbound

This paper cites Amd: Autoregressive motion diffusion.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Amd: Autoregressive motion diffusion

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:17.223219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:18:10.394699Z digest=sha256:0e13eeae009a1befc66b5a833c9a9513f9ce7a37f2007dc5f90f3feffa882467

Observation da7653cb-97cf-45b6-9b9b-a0f7481e01ed · outbound

This paper cites Como: Controllable motion generation through language guided pose code edit- ing.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Como: Controllable motion generation through language guided pose code edit- ing

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:17.019348Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:18:10.457611Z digest=sha256:535a12dbfd380658613c06d121a0ce04be0945b37f96ca64bd3fd865051a0e37

Observation 4fbfc23a-42ea-4035-9321-2186955508ae · outbound

This paper cites Motiongpt: Human motion as a foreign lan- guage.Advances in Neural Information Processing Systems, 36:20067–20079, 2023.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Motiongpt: Human motion as a foreign lan- guage.Advances in Neural Information Processing Systems, 36:20067–20079, 2023

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:16.826747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:18:10.511210Z digest=sha256:4985b87561a24a64b75cf679d2af37b724bb38dbcce214edecc92860c1e39639

Observation e9676649-5f43-4969-8a10-99ff725ef267 · outbound

This paper cites Unimotion: Unifying 3d human motion synthesis and understanding.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Unimotion: Unifying 3d human motion synthesis and understanding

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:16.656878Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:18:10.609319Z digest=sha256:c97db25ea950cccafabcf62f8de97657a26e507adf9d1fafbd0512199b7071fa

Observation b6e88d32-1981-45cc-928b-82dd243fa9dd · outbound

This paper cites Frankenmotion: Part-level hu- man motion generation and composition.arXiv preprint arXiv:2601.10909, 2026.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Frankenmotion: Part-level hu- man motion generation and composition.arXiv preprint arXiv:2601.10909, 2026

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:10.686875Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:10.686875Z digest=sha256:392daf817fabba4e5f3601484e2b2e96d715f2371b594883c3932594e4eae26f

Observation 0baaaaff-ba9b-4d34-9537-72e342d8d4a2 · outbound

This paper cites Motion-x: A large-scale 3d expressive whole-body human motion dataset.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Motion-x: A large-scale 3d expressive whole-body human motion dataset

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:16.419404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:18:10.779221Z digest=sha256:370b44056c4f3c39499d19f9a0e01278ffe63e5d436fb66018b09e2d5e3fd50b

Observation cabcb29f-76bb-468c-98c7-20dadfc5b32b · outbound

This paper cites Multi-granularity Correspondence Learning from Long-term Noisy Videos.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Multi-granularity Correspondence Learning from Long-term Noisy Videos

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:10.857069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:10.857069Z digest=sha256:8a872a39277277989f0be1252b192eda6066625c5a7a491e893d692ce2f1ca44

Observation fc20a501-8fd9-45ba-8940-7ee9aff243e5 · outbound

This paper cites Rethinking diffusion for text-driven human motion generation: Redundant representations, evaluation, and masked autoregression.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Rethinking diffusion for text-driven human motion generation: Redundant representations, evaluation, and masked autoregression

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:16.238937Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:18:10.885666Z digest=sha256:159967b500b9a524767a4459c5aee6c34b7607cde8e6fb6bbb444b6cada86727

Observation 424881ac-e4b6-4324-a00e-6e48c5cc6a0d · outbound

This paper cites Proposal-free temporal action detection via global segmen- tation mask learning.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Proposal-free temporal action detection via global segmen- tation mask learning

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:16.000195Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:18:10.925323Z digest=sha256:67b431927304d08cd364d4e4a5cbc9582b295fef01d400d74bd6118177407742

Observation ecee7a8d-abd9-4d16-b5e1-d65c7a3ea633 · outbound

This paper cites Tmr: Text-to-motion retrieval using contrastive 3d human motion synthesis.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Tmr: Text-to-motion retrieval using contrastive 3d human motion synthesis

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:15.705738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:18:10.986515Z digest=sha256:a222a4f51ab441d2e2c28c5353243264e349352c8bcebc6522969f033d8539d9

Observation 7fbdb2c5-6e48-471e-97f2-46202e197982 · outbound

This paper cites Now Foun- dations and Trends, 2019.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Now Foun- dations and Trends, 2019

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:15.475588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:18:11.034615Z digest=sha256:f0e3bdf61802e66f67a2fb0b64cfc5b286ed118d07ce54cb7faea8acb220a890

Observation 9a584c9d-c0dd-48e8-a5ec-be2e508a0d10 · outbound

This paper cites Mmm: Generative masked motion model.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Mmm: Generative masked motion model

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:15.162114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:18:11.074653Z digest=sha256:9da7d4addf70c8785760e29f35919790788279004e20d82610348718786bbe9f

Observation b977c382-ccef-4a8b-8949-22bc2cc3fab5 · outbound

This paper cites The kit motion-language dataset.Big data, 4(4):236–252,.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision The kit motion-language dataset.Big data, 4(4):236–252,

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:11.165073Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:11.165073Z digest=sha256:3fc91efe55f6e49ff30fe5e42093eb3e1501c296efe9c19757fc3b1d73da85e8

Observation 5d8d8fc0-93e7-4e91-b1d5-9a10e7a9f51a · outbound

This paper cites Babel: Bodies, action and behavior with english la- bels.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Babel: Bodies, action and behavior with english la- bels

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:14.986967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:18:11.246845Z digest=sha256:cd4df5d8e627b05413a3691e1c3af4b4f8796a3956c0c6cd9edf3a02f18a530d

Observation 89c6b694-6741-4344-9dcb-dafa476b1c7c · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Learning transferable visual models from natural language supervi- sion

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:11.279566Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:11.279566Z digest=sha256:07c32f67e16d62c6abfcf5f09bb93ea713905623f723bf79eec4d3ea896d962c

Observation b11f5059-ada5-4929-8cb5-585a11dee210 · outbound

This paper cites an unresolved cited work.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Unresolved cited work

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:11.365326Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:11.365326Z digest=sha256:f826f464653089bd6b34b1f32148f96c25cef3f27e15765c8a1bdc9cb9cfe7db

Observation 232ce928-4530-4ff9-bb09-b7e511c4dfed · outbound

This paper cites Ot-clip: Un- derstanding and generalizing clip via optimal transport.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Ot-clip: Un- derstanding and generalizing clip via optimal transport

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:14.742352Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:18:11.436412Z digest=sha256:28a5503588db8141b9744abd097ecd0f8f8545487bec25002ec3caefb0551d2a

Observation 6b0e7d07-a54a-420e-9339-8cf55db0588d · outbound

This paper cites OpenAI GPT-5 System Card.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision OpenAI GPT-5 System Card

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:11.477503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:11.477503Z digest=sha256:c1d57a86f47ac273be4b045040a3f7485fcea7f287cdbfb6a28f3659a96ac253

Observation ee65441b-8813-4b88-8777-29ae06394ff9 · outbound

This paper cites Optimal transport on discrete domains.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Optimal transport on discrete domains

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:14.580551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:18:11.529984Z digest=sha256:88b9cb8be68d322119012fbd11939f797e7ae4b2aa8cb733f5216d13816b0b91

Observation c6481434-0f9e-4afa-9cc2-cdf53114e0fb · outbound

This paper cites CoMA: Compositional Human Motion Generation with Multi-modal Agents.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision CoMA: Compositional Human Motion Generation with Multi-modal Agents

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:11.588885Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:11.588885Z digest=sha256:0e6a6dc8f783c44bfb87ecc52ad9bdedb482e5ad0512e94459e94161b2dc4c4b

Observation 3ca679e8-3bff-47bc-a24a-7f0a4d873767 · outbound

This paper cites Motionclip: Exposing human motion generation to clip space.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Motionclip: Exposing human motion generation to clip space

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:14.464800Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:18:11.647315Z digest=sha256:60d381f13c1e3b3b738a5ad1fe706327a431f7da7d3c791c09b50e1a3cff52e5

Observation fcaeb8bd-9c66-4074-ae86-5bc2543f89f5 · outbound

This paper cites Human Motion Diffusion Model.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Human Motion Diffusion Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:11.728862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:11.728862Z digest=sha256:bae2f2e0995fc94cba0342cf26b8199b79e01a789bee51b082290da8875ae915

Observation 3f702020-fa09-4a3d-a137-91d7442beead · outbound

This paper cites Scaling Large Motion Models with Million-Level Human Motions.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Scaling Large Motion Models with Million-Level Human Motions

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:11.792396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:11.792396Z digest=sha256:32f040f0baffa09691629c0eded4e3adb08ff3b5a204b9cd5c774200d799cd0b

Observation f40afea6-d756-4ab6-bf7a-4babacb90ddc · outbound

This paper cites Mg-motionllm: A unified framework for motion comprehension and gener- ation across multiple granularities.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Mg-motionllm: A unified framework for motion comprehension and gener- ation across multiple granularities

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:14.206099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:18:11.853616Z digest=sha256:0f47cb947ae149b8af690de6ea9835d370ba7acbc8e85cb48bf88aaad0f1e85d

Observation 65b39cf5-8667-4a57-afbd-b6db884a8c0c · outbound

This paper cites Dense motion captioning.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Dense motion captioning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:14.014695Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:18:11.903996Z digest=sha256:f911be1583fcaa8613c66971d16ef1df5027d94573f9a8dfe4cd748960d90370

Observation f8392932-71c2-4c4c-acea-841098fe600c · outbound

This paper cites FILIP: Fine-grained Interactive Language-Image Pre-Training.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision FILIP: Fine-grained Interactive Language-Image Pre-Training

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:11.969697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:11.969697Z digest=sha256:2907f01691989e5a69c145f35025b3b41775d45db60f2145da8c61048aa8c7af

Observation c779391c-f27a-4d81-80ab-4fbcf11f13e8 · outbound

This paper cites Generating human motion from textual descrip- tions with discrete representations.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Generating human motion from textual descrip- tions with discrete representations

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:13.891286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:18:12.034560Z digest=sha256:77c3a9fcefaeaf35ab05ec0eda7a6ab6e988f7426967623d939e825a12a683e3

Observation f3756519-1f9f-4f69-b1dc-2a1634116f49 · outbound

This paper cites Re- modiffuse: Retrieval-augmented motion diffusion model.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Re- modiffuse: Retrieval-augmented motion diffusion model

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-06T00:18:12.118757Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T00:18:12.118757Z digest=sha256:ea52e10b4561ae5b08ae3a7fc78c39f8a813b64989f825f18f45c3b82b27fac8

Observation b402e152-3973-455d-932c-1835a5056a43 · outbound

This paper cites Finemogen: Fine-grained spatio- temporal motion generation and editing.Advances in Neural Information Processing Systems, 36:13981–13992, 2023.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Finemogen: Fine-grained spatio- temporal motion generation and editing.Advances in Neural Information Processing Systems, 36:13981–13992, 2023

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:13.726895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:18:12.173471Z digest=sha256:0cd06d2709465db1809924aaa6e8e457dd16876903c7180cd6e48c7c4c37e9cc

Observation 8c455f42-98d8-4d38-b96b-1abb68d58607 · outbound

This paper cites Pre- training clip against data poisoning with optimal transport- based matching and alignment.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Pre- training clip against data poisoning with optimal transport- based matching and alignment

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:13.640874Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:18:12.264675Z digest=sha256:fa9edf3b1651f547e95b97bc073e6178370b10560c00aeab4501e0e299fcb4c0

Observation 0973253c-3618-4620-86b4-6fcc9b32012d · outbound

This paper cites DartControl: A diffusion-based autoregressive motion model for real-time text-driven motion control.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision DartControl: A diffusion-based autoregressive motion model for real-time text-driven motion control

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:13.488133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:18:12.328454Z digest=sha256:daf77dc0e1ade985edafa0c638c141666beb81ef77377c8e41fde19af6195531

Observation 79f73070-d566-43e8-923e-878b1e21a7e2 · outbound

This paper cites an unresolved cited work.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:18:13.339249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:18:12.375527Z digest=sha256:fca69777347239e1ad93a27f0e9435ec7f07608c5530af114d8b77014d722933

Observation a6f6d641-ed17-44c8-8ab9-3f0f10634fdb · outbound

This paper cites an unresolved cited work.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision Unresolved cited work

Reference 43

Resolution
unresolved
raw_fallback, observed 2026-08-06T00:18:13.215087Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:18:12.439527Z digest=sha256:f69d09788e09c1ac1168599aad331292bfb798ffbe04e331956bb5d9433606e6

Observation 1845dedb-1e0b-4bd6-b0ef-58b78ec6c43d · outbound

This paper cites The annotation process involves man- ual segment-level boundary identification: for each action segmentj, human annotators specify the temporal interval [tstart, tend]in seconds.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision The annotation process involves man- ual segment-level boundary identification: for each action segmentj, human annotators specify the temporal interval [tstart, tend]in seconds

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:13.060514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:18:12.510219Z digest=sha256:788430550dcab4e424f69de463ddc6a58d65d074a7a860b3f27519492ed5df64

Observation db1aac17-a015-43c4-a123-5f27eda79b80 · outbound

This paper cites specialized human motion analyst.

FineMoLA: Towards Fine-Grained Motion-Language Alignment from Clip-Level Supervision specialized human motion analyst

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T00:18:12.924197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-06T00:18:12.550250Z digest=sha256:d3853c5abe9787bfda7d2e2d2718fe8a572d1568f31783ac3e8ff30bdaa8e688

Pith citing papers

No inbound Pith citation observations are available.