Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T21:53:55.976780Z
Paper Citation Record · LEDGER
As of 15 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2412.04026.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-11T21:53:55.976780Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
62 of 62 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8593755e-5d6b-433e-baff-aa8db76bb587 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction DiffusionNER: Boundary diffusion for named entity recognition,
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 46c8358c-55d4-491d-87e7-4dafacc56f73 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Dual cache for long document neural coreference resolution,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 4a8f2460-1eaa-4815-850b-1764e609ff6b · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction An autoregressive text-to-graph framework for joint entity and relation extraction,
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 35ec8c6f-d5fe-4053-ad91-9bf35c1e0844 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Event extraction as question generation and answering,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation daa359ee-d090-4e94-80c5-e864f9b48fd6 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Rethinking boundaries: End-to-end recognition of discontinuous mentions with pointer networks,
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0b73483d-7035-4c59-aa5b-37e6ccfba00c · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A span-based model for joint overlapped and discontinuous named entity recognition,
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation fac3f90b-679e-4034-aa33-1d6960dc460e · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Unified named entity recognition as word-word relation classification,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 085f8686-478d-456b-9cb0-116e5f2da0b2 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Knowledge enhanced coreference resolution via gated attention,
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 7dd0163b-1bad-41fc-835e-eaa662995bae · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Double graph based reasoning for document-level relation extraction,
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 22b73191-308e-44e7-8b50-839af3353d3a · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Coreference resolution without span representations,
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f7f6b36d-0c67-497c-abdf-afe1d63d9955 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A sequence-to-sequence approach for document-level relation extraction,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 20c8838f-bcaa-4d4a-a2b0-2e48930fb651 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Visual attention model for name tagging in multimodal social media,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6d315406-ac50-46b1-aaea-0ba94cd620c1 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Adaptive co-attention network for named entity recognition in tweets,
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 127096dd-19e9-442e-b007-70ea11d91e36 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A large-scale chinese multimodal ner dataset with speech clues,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 21f13c5b-8483-4977-b277-ce850b99b0a7 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Who are you referring to? coreference resolution in image narrations,
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation d366b976-7698-4089-944c-07597330d6f4 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Mnre: A challenge multimodal dataset for neural relation extraction with visual evidence in social media posts,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ee792f39-f53b-48f1-8a22-1e680a79554c · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A hierarchical network for multimodal document-level relation extraction,
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 64464a7c-2ce7-4c67-b625-b49e06889e5a · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Grounded multimodal named entity recognition on social media,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3b718cd1-8fc2-412e-be63-439595997b9c · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Semi-supervised multimodal coreference resolution in image narrations,
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f55bcf04-1a0f-4fb6-869e-cfbbf9ab1d60 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Joint multimodal entity-relation extraction based on edge-enhanced graph alignment network and word- pair relation tagging,
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3bff049b-ab28-422b-b1af-03f0c210678b · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Multimodal relation extraction with efficient graph alignment,
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1700b069-affd-4949-a99c-3afae2238976 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Docred: A large-scale document-level relation extraction dataset,
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation ce4a9515-a6dd-414c-a1f1-7c77d4244199 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Improving multimodal named entity recognition via entity span detection with unified multimodal transformer,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation dcea31c6-3b17-4c7d-8673-8791ac23cdef · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A span-based multimodal variational autoencoder for semi-supervised multimodal named entity recognition,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a925a4fa-ce05-4b19-b1dd-99d624be8349 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Entity- level interaction via heterogeneous graph for multimodal named entity recognition,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 0b949fab-1e68-48aa-b1b8-73d0f3058adc · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Prompt- ing chatgpt in mner: Enhanced multimodal named entity recognition with auxiliary refined knowledge,
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6b13dd77-f9ea-4701-b6d7-a924b15827e6 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Gravl-bert: graphical visual-linguistic representations for multimodal coreference resolution,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 6a183391-336c-4d56-8247-c8f194b86873 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Good visual guidance make a better extractor: Hierarchical visual prefix for multimodal entity and relation extraction,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 289e4c81-903e-449d-9171-41c4d900c578 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Rethinking multimodal entity and relation extraction from a translation point of view,
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation aba6abb6-d9f3-43dd-b520-2b7b4884d727 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Information screening whilst exploiting! multimodal relation extraction with feature denoising and multimodal topic modeling,
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 59b3563e-06c6-444f-be96-14b66d40ae54 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Transvg: End-to-end visual grounding with transformers,
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 10d8a44e-499c-4f42-9bad-5b21ba5e3538 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Shifting more attention to visual backbone: Query-modulated refinement networks for end-to-end visual grounding,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation b6a867e3-aab2-4351-b594-4ae5589fcb64 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction You only look once: Unified, real-time object detection,
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e60e0871-744e-4ab5-a7c4-3cc522724eb6 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Ssd: Single shot multibox detector,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 70f282fc-3a62-43f4-821c-dbd7404185a1 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Ref-nms: Breaking proposal bottlenecks in two-stage referring expression grounding,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c63856b6-a7dc-4429-b84e-660cddeb2d03 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Missing modalities imputation via cascaded residual autoencoder,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 29585045-df22-4ae7-8611-bcf375eb6047 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Lrmm: Learning to recommend with missing modalities,
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3d5846a8-68ac-426e-9596-808564db30f7 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Dealing with missing modalities in the visual question answer-difference prediction task through knowledge distillation,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 93c3b3cb-0854-4ddc-807a-2ed9696d3f8a · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A unified self-distillation framework for multimodal sentiment analysis with uncertain missing modalities,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 398e2ac7-d9c4-4c6b-8fef-6f8e1c49bfbf · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Missing modality imagination network for emotion recognition with uncertain missing modalities,
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d58a6d9-b30a-42cb-a99e-9020de48ed42 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Mitigating inconsistencies in multimodal sentiment analysis under uncertain missing modalities,
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39c5209d-1341-4f87-a005-90bf62d45817 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Multimodal prompting with missing modalities for visual recognition,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1f90a499-dbcd-4942-8037-355dd9c53f14 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Multimodal prompt learning with missing modalities for sentiment analysis and emotion recognition,
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 29abe825-40e1-4009-826e-5e293f1e5201 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Longformer: The Long-Document Transformer
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2fd49c31-be75-4af4-84ec-3e9a8d5c17d2 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Visual Transformers: Token-based Image Representation and Processing for Computer Vision
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e36b2296-2cc4-4840-883f-2255bf452579 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A primer in bertology: What we know about how bert works,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 548d5047-ec8e-48f8-9686-4e80da3b10bc · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Bert: Pre-training of deep bidirectional transformers for language understanding,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 9696bcfe-f93b-4c3c-aff7-4a3c73dba5d9 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction An image is worth 16x16 words: Transformers for image recognition at scale,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 3609f568-0b54-49d9-94c9-774cda0cecc7 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Auto-encoding variational bayes,
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation c8b2cd50-66c0-47ab-a728-8d61aeeca6a8 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Exploring universal intrinsic task subspace for few-shot learning via prompt tuning,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 53c937e7-2558-431c-a72b-ec0c375c9853 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Attention is all you need,
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation f589483d-954d-4cfd-a5eb-b4ef49dc5c17 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Deep residual learning for image recognition,
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c679d900-ee18-442a-b05f-a707877b0d08 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Conditional random fields: Probabilistic models for segmenting and labeling sequence data,
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 1a07ebe0-1223-4048-bec1-804043acef90 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction VisualBERT: A Simple and Performant Baseline for Vision and Language
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 327576fc-5332-447d-a2c5-cdbbc9b28a7f · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc33cd84-8a54-47fb-be2e-8c7e25261b2f · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Vivit: A video vision transformer,
Reference 57
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f4c4426e-de7d-41ac-966a-d60a7f21d47e · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Videomae: Masked autoen- coders are data-efficient learners for self-supervised video pre-training,
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0ea64b8-955c-4ba8-972f-e48df9251217 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Video-llama: An instruction-tuned audio- visual language model for video understanding,
Reference 59
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dd061f07-0ff1-49b2-981b-b3c6017764f1 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Video-chatgpt: Towards detailed video understanding via large vision and language models,
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation 060b76a4-32cf-41e9-a904-9eddfe4745bb · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Parallel data helps neural entity coreference resolution,
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation a6c55131-0d59-431e-ba0e-6846f88eba82 · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Toe: A grid-tagging discontinuous ner model enhanced by embedding tag/word relations and more fine-grained tags,
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.
Observation acb1842e-afd2-49d1-9f2f-3db8a1d6502e · outbound
M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A fast and accurate one-stage approach to visual grounding,
Reference 63
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
No inbound Pith citation observations are available.