Pith. sign in

Paper Citation Record · LEDGER

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech

As of 13 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 1 inbound Pith citation observation for arXiv:2412.11409.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.11409 v3

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T15:01:33.912139Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:34:07.907336Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-11T04:34:07.970750Z

Reference resolution

38 of 38 outbound references displayed

  • verified exact0
  • verified fuzzy3
  • unresolved35
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c705524e-f197-4a1d-939c-0c3fa5483dc6 · outbound

This paper cites , " * write output.state after.block = add.period write newline.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech , " * write output.state after.block = add.period write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-11T15:01:33.739038Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:01:33.739038Z digest=sha256:e276ae89d73bbbec81bd22236d8f98324aa7e0f7daa5125a86ecf27cb655e2be

Observation 9e631196-c032-4965-bc13-04bbe0c34f52 · outbound

This paper cites write newline.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-11T15:01:33.744102Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:01:33.744102Z digest=sha256:23ac7b24448e598f38951aef46cd3484fd0a58a3113948da2675f5937648c297

Observation 5317d04d-bc0d-425d-83a1-b2e1a780c7e3 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 3

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.462181Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.749585Z digest=sha256:f0717173559455ef3c3c27c3bd64309bfabb62f37c7dcc53579d88593b0a536f

Observation ffb4a67a-d664-4158-8575-be662b556a44 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 4

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.446298Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.754047Z digest=sha256:b71b62af2192464cb9535751c01ab7d791745a9c8ec2aa599cf15cb7390f04d0

Observation acb71e47-d6c8-4f24-bbca-d13450814309 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.431945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.759161Z digest=sha256:ff82a0f4d74e9699243f7c1343d637054b78fcf0deaceee9533f0fb0ffc62697

Observation a79ebfa7-0d48-4c24-a165-d8486079fca1 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.416672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.763741Z digest=sha256:ab292e56b7e50509260da9e8c0d8d7cebd2b01bc02ed379f60789aa0c876729b

Observation 67654fb8-28d2-4741-824e-0d5c02c4aead · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.399707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.767991Z digest=sha256:93af7567f1f18d7b65923acb73cdb22bf249613a00b72e635893a1cf699e2e6e

Observation daace53f-7531-49e7-a731-8ac77f4de9ac · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.384054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.772569Z digest=sha256:69ded8c4824981dd56e8017e940cd8a3d07adc62c0300c5660140295c1a52060

Observation 6afcc842-ea5b-4c2d-8d9f-943ec20cf967 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-11T15:01:33.776974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:01:33.776974Z digest=sha256:6fc6553cce4f67c9532b9819ac0c733b4518a8e99c9c3379faf2832d366b24f8

Observation 1b34d21b-ea50-411c-9aee-745b07d2dab8 · outbound

This paper cites Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Scene-LLM: Extending Language Model for 3D Visual Understanding and Reasoning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-11T15:01:33.780829Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:01:33.780829Z digest=sha256:01d0e8bbcc3e7183b7eba9f76d40f0a012eadaf042efcc7c9ea7c726f73cbaa5

Observation e9b42dad-0124-4f16-8954-0d31983c5599 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.358773Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.786306Z digest=sha256:8bfe402b6c1bed5d8cca36c7ea9cda29b6491c5eaf8a974525ca1010d3baa75c

Observation 9ff317ad-9929-4402-a93f-f0be088005fe · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-11T15:01:33.791200Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:01:33.791200Z digest=sha256:4cc0e7b6a3f130ae35b7e2cc5566b157c6cc19625a45d88113870d134c4587a3

Observation 5ff458eb-6d03-4358-9d4f-977c5d726824 · outbound

This paper cites Multi-Source Spatial Knowledge Understanding for Immersive Visual Text-to-Speech.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Multi-Source Spatial Knowledge Understanding for Immersive Visual Text-to-Speech

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T15:01:33.795919Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:01:33.795919Z digest=sha256:bc81d9e06918f90db87fe7b6ed54d21fcb8a8229b4b539d9be68d57afbacb2dd

Observation 3947395a-d4de-42a0-bf82-50dbe3d93928 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.334532Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.800794Z digest=sha256:dde2948979521fd9ac3f73edcec342f568ed5016fc32445ba686fa16537f88c0

Observation 2001f691-e00a-460b-b8fe-fb14d7720ec7 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.320257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.805342Z digest=sha256:50757930fcc4fd0219f080203270499a2f5b6502d254187036ea1522311c23c1

Observation 1be592c6-f85d-4895-89d2-5d38ccae6684 · outbound

This paper cites I Want to Figure Things Out.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech I Want to Figure Things Out

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:01:34.306395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.810025Z digest=sha256:bfbd99ac5afd66626261357aab397c81d01bf9d87824fc1caf9cd00b3506a980

Observation 9574b814-ab1e-4dec-b20f-7b5d18629c11 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.292349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.814503Z digest=sha256:23f4b9923246bc4e70bf1299f1ce4360598a7e1afa4e1c54b5cf189d4c0d4720

Observation dd355004-bb8f-4cbb-8192-a963e173c590 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.279446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.819181Z digest=sha256:2d6f642a2a60c93017ee87b29d9a182fd8e91fb459ff0bd113ab6df2895a711a

Observation 40ae4745-caec-445a-a88f-6ae1dd587d2c · outbound

This paper cites BigVGAN: A Universal Neural Vocoder with Large-Scale Training.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech BigVGAN: A Universal Neural Vocoder with Large-Scale Training

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-11T15:01:33.823743Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:01:33.823743Z digest=sha256:b2a7eaca41ce7a62e8677041cfaf622863404f14ead581ab7a74a2ca8be52cc8

Observation c0978bca-366d-49ac-801b-55b80b065418 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.266073Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.829016Z digest=sha256:458493458a976ff79a2cf15b9644ad12db3d8e773bbe439c65898aea982ccef8

Observation 70e70d2c-3577-409f-8edb-c6d83fec4a91 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.252060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.833955Z digest=sha256:de9ede53f1e5577a77eda3b7da699b39d3f09feefdc182daa0aedffb199d9b17

Observation 69526ea0-acd6-4c1a-b90b-a65febcd5a91 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.237380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.838309Z digest=sha256:4526b18df4d3079054e42575fc6bc18b511c9a99f967d2890d5518f08ea2946a

Observation 7a359cd9-a89b-4151-b592-2b6075e3e73b · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.221868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.842440Z digest=sha256:3e15827bd72265c04bf0a93927c2caa2a89359e81a47910ff804587d747c12bd

Observation 95b7c3d1-10b7-4e3d-9774-a5ee7c04d747 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.206779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.846787Z digest=sha256:c45692fa321ee6d3aa46955da42654197f88045ce7cb9391def4f6dbfbecfcc2

Observation 55ec6a08-cbb9-4afb-bae2-853d5c06997c · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.191700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.851111Z digest=sha256:45b52e6861cb6bcdd78d72402801da81027937dab61835ff5e50788d9b1ebc23

Observation 0b52ec4d-6e79-4e64-b88e-82f8156726f0 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.176506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.855330Z digest=sha256:f322eb809058310e11f0933e8aab1a7fa5574afbaa610b897ae0b10249aa1fe8

Observation 4ed2f487-db09-40f8-856e-fd0699c39ed2 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.162005Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.859348Z digest=sha256:8631daced98ed668b780fe8e2cc8240b501f99ead7e2673d535eef754d3e9fe0

Observation b48ec891-ff5a-4114-a4ff-f89f1b5c907f · outbound

This paper cites W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; et al

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-11T15:01:33.864813Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:01:33.864813Z digest=sha256:324606fe1de9de3ed94c4dce667853ae1b1b2b9e9747b2762ad3d3508e5ed392

Observation 1e94e569-b79c-4f41-b772-23178a59a521 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.138379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.868936Z digest=sha256:aee8f2841473c378d4932725fc4eee6167121c8a7d85145a44a375a619fb1c3c

Observation 640690fe-27a4-4af8-98ac-b369067a71bd · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.125191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.872844Z digest=sha256:fdddd5cf636babe9c941676e372dc8a315d6fa74de19fc94a6c92ccbddf4822c

Observation 79c5ca32-baf5-4c51-a7dd-bb5b881743ab · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 31

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.112109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.878274Z digest=sha256:86ca8e6d8094de970f1edd0501cc6476166a49fb1a7ce32727f11699afd1e36e

Observation 0c14f8f4-2d1c-4f34-80e8-7f37a6a47d4e · outbound

This paper cites N.; Tran, S.; Yao, B.; Chilimbi, T.; and Shah, M.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech N.; Tran, S.; Yao, B.; Chilimbi, T.; and Shah, M

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:01:34.097133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.882704Z digest=sha256:1506b2fa410e3995130613e9117e5bf2b2e6a9057683238dd4a8629a2663080f

Observation 38206b27-7f94-472a-9891-5a092783cb75 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.081238Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.887301Z digest=sha256:47f11279efb2c349c783b8ddc48c0057be07019a04881bcc78d950d75bb1bfbd

Observation 1b240567-7cb5-4136-b312-0db0ea272faa · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T15:01:33.891600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T15:01:33.891600Z digest=sha256:d182be0ca5e2993b022063ee0bb564836a1b27274fe0859c8f64f3880d88ad1b

Observation 17aeb1c9-3e11-4734-853c-aad21ad18bd7 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.066593Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.897141Z digest=sha256:55c87911bcf4aff11e5c15dc12b628b707a6972ede6d11db6f4b873ff922f6c4

Observation 761eacb8-d295-4c79-bb62-0f9c58bee78e · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 36

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.052267Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.902340Z digest=sha256:01c6b51c1be1834520bf6a5decafc955b3817b7333d7a296e8f7a42f1081ee38

Observation bd7e2edf-23ba-465b-9cbc-7bd953f8cfcb · outbound

This paper cites C.; and YAN, S.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech C.; and YAN, S

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T15:01:34.036847Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.907180Z digest=sha256:706d94bee0d8058a187be35c1bdd06a648b134400e4214aeb7a8202b2d6fd167

Observation 01a873bd-32f7-49ba-9806-04828d57aec4 · outbound

This paper cites an unresolved cited work.

Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-11T15:01:34.020850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-11T15:01:33.912139Z digest=sha256:e88db964aaaa63b66b2176f48d7c58ed41d37d9a31b9ae072160e21475eae5e5

Pith citing papers

Observation 6ece83f0-cdae-48d8-b707-a9faaa591b34 · inbound

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction cites this paper.

Towards Expressive Video Dubbing with Multiscale Multimodal Context Interaction Multi-modal and Multi-scale Spatial Environment Understanding for Immersive Visual Text-to-Speech

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-11T04:34:07.977433Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-11T04:34:07.907336Z digest=sha256:b6b685c04b8209f128c8cff9e590d88fb37cdc3068d8b9917505674829dd2a34