Pith. sign in

Paper Citation Record · LEDGER

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction

As of 15 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2412.04026.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04026 v2

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:53:55.976780Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy49
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8593755e-5d6b-433e-baff-aa8db76bb587 · outbound

This paper cites DiffusionNER: Boundary diffusion for named entity recognition,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction DiffusionNER: Boundary diffusion for named entity recognition,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:57.098637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.645493Z digest=sha256:df45da0e9d6af8b7f40ad889f30dff2ccabdeacd214257b7cb647684df891735

Observation 46c8358c-55d4-491d-87e7-4dafacc56f73 · outbound

This paper cites Dual cache for long document neural coreference resolution,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Dual cache for long document neural coreference resolution,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:57.082147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.651675Z digest=sha256:a53901d9f15a53c029f9f4c11f7590262e29646868646fe175ac6d608b47a2cf

Observation 4a8f2460-1eaa-4815-850b-1764e609ff6b · outbound

This paper cites An autoregressive text-to-graph framework for joint entity and relation extraction,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction An autoregressive text-to-graph framework for joint entity and relation extraction,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:57.064499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.656845Z digest=sha256:e57f79c63b3e074ec90fe979b6dda0ca501a87b013b1a06dec938a24483af561

Observation 35ec8c6f-d5fe-4053-ad91-9bf35c1e0844 · outbound

This paper cites Event extraction as question generation and answering,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Event extraction as question generation and answering,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:57.047700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.662619Z digest=sha256:adba4e3939fb65bd2afdb12807e2295486947df63953223da585d9970832b234

Observation daa359ee-d090-4e94-80c5-e864f9b48fd6 · outbound

This paper cites Rethinking boundaries: End-to-end recognition of discontinuous mentions with pointer networks,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Rethinking boundaries: End-to-end recognition of discontinuous mentions with pointer networks,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:57.029864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.667965Z digest=sha256:ad5592859577d3e2cefdf27d02c18a4aabe1fbecd72ec367d8007278b3e8caa4

Observation 0b73483d-7035-4c59-aa5b-37e6ccfba00c · outbound

This paper cites A span-based model for joint overlapped and discontinuous named entity recognition,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A span-based model for joint overlapped and discontinuous named entity recognition,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:57.010404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.674857Z digest=sha256:281e32474278a616835f149f16464cdd33e87d891c03840ae966600fddc6db72

Observation fac3f90b-679e-4034-aa33-1d6960dc460e · outbound

This paper cites Unified named entity recognition as word-word relation classification,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Unified named entity recognition as word-word relation classification,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.991325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.681483Z digest=sha256:71ff5d4f8469dc46e0051c9c452b896baef6105753f880265e2d6b90f036bafe

Observation 085f8686-478d-456b-9cb0-116e5f2da0b2 · outbound

This paper cites Knowledge enhanced coreference resolution via gated attention,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Knowledge enhanced coreference resolution via gated attention,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.972872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.686416Z digest=sha256:9732ddec4eddfb174767460aa783b8312c908dd29adc6481f6cc5bfd0c2f41ef

Observation 7dd0163b-1bad-41fc-835e-eaa662995bae · outbound

This paper cites Double graph based reasoning for document-level relation extraction,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Double graph based reasoning for document-level relation extraction,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.953381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.691524Z digest=sha256:5364ca86f39221bea8d84535702c9da6fbfaf4dd7e1672d2a5359680c2bd9655

Observation 22b73191-308e-44e7-8b50-839af3353d3a · outbound

This paper cites Coreference resolution without span representations,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Coreference resolution without span representations,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.933021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.696807Z digest=sha256:d55bb15f60070a84d73a73a4649eb4863b4d8e7b2a01c7611f24a24e6fc12f25

Observation f7f6b36d-0c67-497c-abdf-afe1d63d9955 · outbound

This paper cites A sequence-to-sequence approach for document-level relation extraction,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A sequence-to-sequence approach for document-level relation extraction,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.915822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.701825Z digest=sha256:f713a7d975792e3bcfa9bf5f4619a7c0fe7f96ce68350f69069f486d106cd01c

Observation 20c8838f-bcaa-4d4a-a2b0-2e48930fb651 · outbound

This paper cites Visual attention model for name tagging in multimodal social media,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Visual attention model for name tagging in multimodal social media,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.899191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.707133Z digest=sha256:1d752bdac50f6237a81267ae8fa6a4016cce9fe4c8f2c4ed3eca25253d0f8380

Observation 6d315406-ac50-46b1-aaea-0ba94cd620c1 · outbound

This paper cites Adaptive co-attention network for named entity recognition in tweets,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Adaptive co-attention network for named entity recognition in tweets,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.882580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.712237Z digest=sha256:705505984d5dd7794943afcd61c438451db7f7d840873e74fff106829bdaa80d

Observation 127096dd-19e9-442e-b007-70ea11d91e36 · outbound

This paper cites A large-scale chinese multimodal ner dataset with speech clues,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A large-scale chinese multimodal ner dataset with speech clues,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.866161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.717006Z digest=sha256:a77dc72c02e7b9a3928794a8add0067f62610e5c5fcf791f9d941247e0ee330d

Observation 21f13c5b-8483-4977-b277-ce850b99b0a7 · outbound

This paper cites Who are you referring to? coreference resolution in image narrations,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Who are you referring to? coreference resolution in image narrations,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.831366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.726731Z digest=sha256:e107b36963d8c5d18a3bf2d15fb690388f931896efab1cfcde7a3ed9b8e6512b

Observation d366b976-7698-4089-944c-07597330d6f4 · outbound

This paper cites Mnre: A challenge multimodal dataset for neural relation extraction with visual evidence in social media posts,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Mnre: A challenge multimodal dataset for neural relation extraction with visual evidence in social media posts,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.811261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.731553Z digest=sha256:20284a9c0a624813d11748ffb0981d51cfe249d146406fd06b7478dd8a2aa246

Observation ee792f39-f53b-48f1-8a22-1e680a79554c · outbound

This paper cites A hierarchical network for multimodal document-level relation extraction,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A hierarchical network for multimodal document-level relation extraction,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.794259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.736909Z digest=sha256:b0cb2913e4ff2416276314a7f71a9c4a6c2114f47cf8ee137eaf1d5f7f420209

Observation 64464a7c-2ce7-4c67-b625-b49e06889e5a · outbound

This paper cites Grounded multimodal named entity recognition on social media,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Grounded multimodal named entity recognition on social media,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.849101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.741883Z digest=sha256:b6207d4434b23f400631181ab8bde3274556509101ef3faf69e2fb556c168114

Observation 3b718cd1-8fc2-412e-be63-439595997b9c · outbound

This paper cites Semi-supervised multimodal coreference resolution in image narrations,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Semi-supervised multimodal coreference resolution in image narrations,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.775857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.748153Z digest=sha256:1bf66014d80003850cd3404bc471e4cc935bf2d3f7b360a9f50d22f0c4084366

Observation f55bcf04-1a0f-4fb6-869e-cfbbf9ab1d60 · outbound

This paper cites Joint multimodal entity-relation extraction based on edge-enhanced graph alignment network and word- pair relation tagging,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Joint multimodal entity-relation extraction based on edge-enhanced graph alignment network and word- pair relation tagging,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.759051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.753323Z digest=sha256:3fb6f5b83734043bb56222c6f8abf911af5c57110911b53153cfd415fd78cbee

Observation 3bff049b-ab28-422b-b1af-03f0c210678b · outbound

This paper cites Multimodal relation extraction with efficient graph alignment,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Multimodal relation extraction with efficient graph alignment,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.740778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.759044Z digest=sha256:4ada428030c0b4cf5bb2cd89d3f5ee45c0eb18960cdad49174e0b1db5c86a26f

Observation 1700b069-affd-4949-a99c-3afae2238976 · outbound

This paper cites Docred: A large-scale document-level relation extraction dataset,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Docred: A large-scale document-level relation extraction dataset,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.724088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.764038Z digest=sha256:dd48d99352e87df5dbd385d5314ff1feaad4b8c4b697b75e9cbc737992c34fa2

Observation ce4a9515-a6dd-414c-a1f1-7c77d4244199 · outbound

This paper cites Improving multimodal named entity recognition via entity span detection with unified multimodal transformer,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Improving multimodal named entity recognition via entity span detection with unified multimodal transformer,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.706299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.768999Z digest=sha256:080c49910d0fb88e1b92ad6d5e009753c3f08d99d8aefeb1e2293e0352dd431d

Observation dcea31c6-3b17-4c7d-8673-8791ac23cdef · outbound

This paper cites A span-based multimodal variational autoencoder for semi-supervised multimodal named entity recognition,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A span-based multimodal variational autoencoder for semi-supervised multimodal named entity recognition,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.688592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.774151Z digest=sha256:c0842d84adc59e12b2197a3ce1bbad85a6e75fe6bae2fef9cabc6f644a97f626

Observation a925a4fa-ce05-4b19-b1dd-99d624be8349 · outbound

This paper cites Entity- level interaction via heterogeneous graph for multimodal named entity recognition,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Entity- level interaction via heterogeneous graph for multimodal named entity recognition,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.670543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.779488Z digest=sha256:a86fd610cc622828563171d6c7a4926920c501e0f805314d580b9110ee65ed0a

Observation 0b949fab-1e68-48aa-b1b8-73d0f3058adc · outbound

This paper cites Prompt- ing chatgpt in mner: Enhanced multimodal named entity recognition with auxiliary refined knowledge,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Prompt- ing chatgpt in mner: Enhanced multimodal named entity recognition with auxiliary refined knowledge,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.650947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.785111Z digest=sha256:ac61b5eee96281520216368fea18cd83ef6d656d96b232a740f7ee7dcde50e71

Observation 6b13dd77-f9ea-4701-b6d7-a924b15827e6 · outbound

This paper cites Gravl-bert: graphical visual-linguistic representations for multimodal coreference resolution,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Gravl-bert: graphical visual-linguistic representations for multimodal coreference resolution,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.632958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.790468Z digest=sha256:628c10f100b2f4de3470d8e9d068619a78893fa2a8368734c2bcf36d82f64d16

Observation 6a183391-336c-4d56-8247-c8f194b86873 · outbound

This paper cites Good visual guidance make a better extractor: Hierarchical visual prefix for multimodal entity and relation extraction,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Good visual guidance make a better extractor: Hierarchical visual prefix for multimodal entity and relation extraction,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.615185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.795316Z digest=sha256:8d8019335b0ea40e2a8b4dbdd87a2abdfa1e37a2a2a8c925d2e3ca74b949ad78

Observation 289e4c81-903e-449d-9171-41c4d900c578 · outbound

This paper cites Rethinking multimodal entity and relation extraction from a translation point of view,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Rethinking multimodal entity and relation extraction from a translation point of view,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.597643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.802304Z digest=sha256:31ba8277130cdb7ec4f155afee9caba42f05e43ad8b3abdaa8c8677e2d2bb19d

Observation aba6abb6-d9f3-43dd-b520-2b7b4884d727 · outbound

This paper cites Information screening whilst exploiting! multimodal relation extraction with feature denoising and multimodal topic modeling,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Information screening whilst exploiting! multimodal relation extraction with feature denoising and multimodal topic modeling,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.579968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.807645Z digest=sha256:ac3edcdfaaf2d0c20b66709fcabb354a998ee8dea2e6d7e85a16ecee09057c38

Observation 59b3563e-06c6-444f-be96-14b66d40ae54 · outbound

This paper cites Transvg: End-to-end visual grounding with transformers,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Transvg: End-to-end visual grounding with transformers,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.812633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.812633Z digest=sha256:aaaa8ec600f4db6db1b9faced3f4462e16a7b2d19eadc3d67f1aa3d01478f7e0

Observation 10d8a44e-499c-4f42-9bad-5b21ba5e3538 · outbound

This paper cites Shifting more attention to visual backbone: Query-modulated refinement networks for end-to-end visual grounding,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Shifting more attention to visual backbone: Query-modulated refinement networks for end-to-end visual grounding,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.549750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.818110Z digest=sha256:ffd7e0e24ae750512e75a007e2ea219d1ca348b904c117802854816bf6c7114d

Observation b6a867e3-aab2-4351-b594-4ae5589fcb64 · outbound

This paper cites You only look once: Unified, real-time object detection,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction You only look once: Unified, real-time object detection,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.823763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.823763Z digest=sha256:75935048c023fae6c3c8b401caec0946c96dfe12f8534a8bc9baee4d66f81bd6

Observation e60e0871-744e-4ab5-a7c4-3cc522724eb6 · outbound

This paper cites Ssd: Single shot multibox detector,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Ssd: Single shot multibox detector,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.519786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.829059Z digest=sha256:a71df50110c3a5292c4e99c33e385a3750232ee8bc5221fd6bb4607fda3c1908

Observation 70f282fc-3a62-43f4-821c-dbd7404185a1 · outbound

This paper cites Ref-nms: Breaking proposal bottlenecks in two-stage referring expression grounding,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Ref-nms: Breaking proposal bottlenecks in two-stage referring expression grounding,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.500810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.834799Z digest=sha256:c4f3bd112d0b26f9a531c78f6ab8c21522cded352492afea0764f6d121955127

Observation c63856b6-a7dc-4429-b84e-660cddeb2d03 · outbound

This paper cites Missing modalities imputation via cascaded residual autoencoder,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Missing modalities imputation via cascaded residual autoencoder,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.481316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.839867Z digest=sha256:c347afbd06271ee6ec6e2a57579a8e16bdb854e7f57a79d19e42e25379664978

Observation 29585045-df22-4ae7-8611-bcf375eb6047 · outbound

This paper cites Lrmm: Learning to recommend with missing modalities,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Lrmm: Learning to recommend with missing modalities,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.463516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.844766Z digest=sha256:d808838b03ad832e65c4c4119ccaf47f26eddffbd89c900b11d9bec2910582d5

Observation 3d5846a8-68ac-426e-9596-808564db30f7 · outbound

This paper cites Dealing with missing modalities in the visual question answer-difference prediction task through knowledge distillation,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Dealing with missing modalities in the visual question answer-difference prediction task through knowledge distillation,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.445404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.850013Z digest=sha256:0696ee742eb2c472f1e5bb9d819c0c3ae2c000b512249c996a9ad4930fe8e148

Observation 93c3b3cb-0854-4ddc-807a-2ed9696d3f8a · outbound

This paper cites A unified self-distillation framework for multimodal sentiment analysis with uncertain missing modalities,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A unified self-distillation framework for multimodal sentiment analysis with uncertain missing modalities,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.427616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.855160Z digest=sha256:87088ef046aaacbcf1eea45525e8f76fa68253abbcd09d046eb200a35406d144

Observation 398e2ac7-d9c4-4c6b-8fef-6f8e1c49bfbf · outbound

This paper cites Missing modality imagination network for emotion recognition with uncertain missing modalities,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Missing modality imagination network for emotion recognition with uncertain missing modalities,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.860148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.860148Z digest=sha256:6f1a1fdb75a4610ec62f85e1a4c1507bdde5131540efd4bae27a917b54806600

Observation 3d58a6d9-b30a-42cb-a99e-9020de48ed42 · outbound

This paper cites Mitigating inconsistencies in multimodal sentiment analysis under uncertain missing modalities,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Mitigating inconsistencies in multimodal sentiment analysis under uncertain missing modalities,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.865429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.865429Z digest=sha256:07d419bcc7f7c44b3e655b8ee7549474833d04922319a4f616539a4e8e6b41c4

Observation 39c5209d-1341-4f87-a005-90bf62d45817 · outbound

This paper cites Multimodal prompting with missing modalities for visual recognition,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Multimodal prompting with missing modalities for visual recognition,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.386234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.870407Z digest=sha256:0653e854cddcdfcfc34a401133459a73073673bde069918e2534e6fe6c8ed787

Observation 1f90a499-dbcd-4942-8037-355dd9c53f14 · outbound

This paper cites Multimodal prompt learning with missing modalities for sentiment analysis and emotion recognition,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Multimodal prompt learning with missing modalities for sentiment analysis and emotion recognition,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.368424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.875900Z digest=sha256:59fb6853eec81b2b7132b7757301955a61951511fc39e0c936ed157d89c68144

Observation 29abe825-40e1-4009-826e-5e293f1e5201 · outbound

This paper cites Longformer: The Long-Document Transformer.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Longformer: The Long-Document Transformer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.881703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.881703Z digest=sha256:3b73090de103636de8c8979e967b421b023dec7d535b2ab1ceb4b77a3380d471

Observation 2fd49c31-be75-4af4-84ec-3e9a8d5c17d2 · outbound

This paper cites Visual Transformers: Token-based Image Representation and Processing for Computer Vision.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.887844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.887844Z digest=sha256:bc4dc1d67a0482b361e85c1b884ea4778330cd6ee0c9904f74c0403ba7b15f7b

Observation e36b2296-2cc4-4840-883f-2255bf452579 · outbound

This paper cites A primer in bertology: What we know about how bert works,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A primer in bertology: What we know about how bert works,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.349562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.894087Z digest=sha256:8e6bc65d5326ee8983bb450cb2c7d58b52fb70fb79e0875cf4620c119bd6a087

Observation 548d5047-ec8e-48f8-9686-4e80da3b10bc · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Bert: Pre-training of deep bidirectional transformers for language understanding,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.328333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.899056Z digest=sha256:137b16983713add3450cfa9f9eb2fa87799d3b76bda3bbcbe27e6f92ce35912d

Observation 9696bcfe-f93b-4c3c-aff7-4a3c73dba5d9 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.309219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.904139Z digest=sha256:ba18362461760c5432d44f94e334c5276070fe1a7d730aaa375f066801abd5b7

Observation 3609f568-0b54-49d9-94c9-774cda0cecc7 · outbound

This paper cites Auto-encoding variational bayes,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Auto-encoding variational bayes,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.289980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.910780Z digest=sha256:ef85bc9d75ff1bd9dc783fdcb0ccbc2fba606d6093296feab5c0aa0a2c223537

Observation c8b2cd50-66c0-47ab-a728-8d61aeeca6a8 · outbound

This paper cites Exploring universal intrinsic task subspace for few-shot learning via prompt tuning,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Exploring universal intrinsic task subspace for few-shot learning via prompt tuning,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.267587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.915496Z digest=sha256:0850946e86658e920a81515e29a9c834baee8657b9c9e3b8391a3e249f4f9ca8

Observation 53c937e7-2558-431c-a72b-ec0c375c9853 · outbound

This paper cites Attention is all you need,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Attention is all you need,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.247428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.920736Z digest=sha256:40cbd74c059d8866e41eaa012bc5d5152dfc4a5470cbe098bc6f2c44821a2c16

Observation f589483d-954d-4cfd-a5eb-b4ef49dc5c17 · outbound

This paper cites Deep residual learning for image recognition,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Deep residual learning for image recognition,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.925634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.925634Z digest=sha256:c673f5e518129eeb6f9d6d91c5c5fa39e0554517959b404363f93066ad49455a

Observation c679d900-ee18-442a-b05f-a707877b0d08 · outbound

This paper cites Conditional random fields: Probabilistic models for segmenting and labeling sequence data,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Conditional random fields: Probabilistic models for segmenting and labeling sequence data,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.215489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.931019Z digest=sha256:574a943065d6c655ac0d61e718bcfd8b044f336197bf00e16f4faeb8e06e29c4

Observation 1a07ebe0-1223-4048-bec1-804043acef90 · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.935726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.935726Z digest=sha256:d29c7215d72aa2fa420b244d98d5345e805f7cc2ddba12d5f54ef72afd497b7d

Observation 327576fc-5332-447d-a2c5-cdbbc9b28a7f · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.940778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.940778Z digest=sha256:678b50f7082eb16f9e77e23e2c6396bed26e7960c9b94dd89b57710aba590086

Observation dc33cd84-8a54-47fb-be2e-8c7e25261b2f · outbound

This paper cites Vivit: A video vision transformer,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Vivit: A video vision transformer,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.945512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.945512Z digest=sha256:cbb4cfdd6785ed424842d2ff324585bc62a91167297b6b58c0cb5436469164bb

Observation f4c4426e-de7d-41ac-966a-d60a7f21d47e · outbound

This paper cites Videomae: Masked autoen- coders are data-efficient learners for self-supervised video pre-training,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Videomae: Masked autoen- coders are data-efficient learners for self-supervised video pre-training,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.950362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.950362Z digest=sha256:a6588f901738e12bee4322fa1e61a5adb3aaa46dc300b9427cc5bab53737b70a

Observation e0ea64b8-955c-4ba8-972f-e48df9251217 · outbound

This paper cites Video-llama: An instruction-tuned audio- visual language model for video understanding,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Video-llama: An instruction-tuned audio- visual language model for video understanding,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.955213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.955213Z digest=sha256:08a287ac4afe4f799d917ce755fee22c5fc805e5b32d8bb784c5ecaed3b13c93

Observation dd061f07-0ff1-49b2-981b-b3c6017764f1 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Video-chatgpt: Towards detailed video understanding via large vision and language models,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.132655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.960826Z digest=sha256:c8611ace1b403ee9f9e95b02202e01c2964fe5a3c066ef913bc1870be3e5b9dc

Observation 060b76a4-32cf-41e9-a904-9eddfe4745bb · outbound

This paper cites Parallel data helps neural entity coreference resolution,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Parallel data helps neural entity coreference resolution,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.115435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.965982Z digest=sha256:071e0acf55a23e807eb2972bc06af34a475c8ac0006e2a1b46b08367b4df5d1c

Observation a6c55131-0d59-431e-ba0e-6846f88eba82 · outbound

This paper cites Toe: A grid-tagging discontinuous ner model enhanced by embedding tag/word relations and more fine-grained tags,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Toe: A grid-tagging discontinuous ner model enhanced by embedding tag/word relations and more fine-grained tags,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.095439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-08-11T21:53:55.971683Z digest=sha256:08b99599461ff05309d0f190260434b716811d713b2c3b3a3793408884776c88

Observation acb1842e-afd2-49d1-9f2f-3db8a1d6502e · outbound

This paper cites A fast and accurate one-stage approach to visual grounding,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A fast and accurate one-stage approach to visual grounding,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.976780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.976780Z digest=sha256:438e03d40d29d4144b4a498be7ae1c8f878dadd0d2fba1fd1598125ddff019e1

Pith citing papers

No inbound Pith citation observations are available.