Pith. sign in

Paper Citation Record · LEDGER

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction

As of 16 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 0 inbound Pith citation observations for arXiv:2412.04026.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2412.04026 v2

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-11T21:53:55.976780Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact0
  • verified fuzzy49
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8593755e-5d6b-433e-baff-aa8db76bb587 · outbound

This paper cites DiffusionNER: Boundary diffusion for named entity recognition,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction DiffusionNER: Boundary diffusion for named entity recognition,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:57.098637Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.645493Z digest=sha256:415a45d40aafab9b0fb823d4ca7cd3d57161d36eaeb4664f85de417e87dfbad7

Observation 46c8358c-55d4-491d-87e7-4dafacc56f73 · outbound

This paper cites Dual cache for long document neural coreference resolution,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Dual cache for long document neural coreference resolution,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:57.082147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.651675Z digest=sha256:ee15188a2c26c5034a7a1fef15c392778d225410a07cce0d086c55438081dda0

Observation 4a8f2460-1eaa-4815-850b-1764e609ff6b · outbound

This paper cites An autoregressive text-to-graph framework for joint entity and relation extraction,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction An autoregressive text-to-graph framework for joint entity and relation extraction,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:57.064499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.656845Z digest=sha256:6920518738cc16d0af38da177477ef10cf9989d78356ac020732c2f21a7980a8

Observation 35ec8c6f-d5fe-4053-ad91-9bf35c1e0844 · outbound

This paper cites Event extraction as question generation and answering,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Event extraction as question generation and answering,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:57.047700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.662619Z digest=sha256:d126587198d3c9fdb6a27930af28aad999ea9e6d8e34b65977aae7261ec1add0

Observation daa359ee-d090-4e94-80c5-e864f9b48fd6 · outbound

This paper cites Rethinking boundaries: End-to-end recognition of discontinuous mentions with pointer networks,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Rethinking boundaries: End-to-end recognition of discontinuous mentions with pointer networks,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:57.029864Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.667965Z digest=sha256:ecce59131442e35b99b217a7cbeca55d2b575c1b64b6173c8fb346cbe575a68f

Observation 0b73483d-7035-4c59-aa5b-37e6ccfba00c · outbound

This paper cites A span-based model for joint overlapped and discontinuous named entity recognition,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A span-based model for joint overlapped and discontinuous named entity recognition,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:57.010404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.674857Z digest=sha256:d7c2ab6cdb5d2862a0bd6e4b990b4f2935be6c98dd52b748acff0e55469938b1

Observation fac3f90b-679e-4034-aa33-1d6960dc460e · outbound

This paper cites Unified named entity recognition as word-word relation classification,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Unified named entity recognition as word-word relation classification,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.991325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.681483Z digest=sha256:041e94ed10637b8e57021781d4fd13c7183d39337bdda9e6203d86b07edd9f72

Observation 085f8686-478d-456b-9cb0-116e5f2da0b2 · outbound

This paper cites Knowledge enhanced coreference resolution via gated attention,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Knowledge enhanced coreference resolution via gated attention,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.972872Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.686416Z digest=sha256:8940b925bff43e9399f7d7100b91e7ea0912babb5d4d14411d94e563b6570c43

Observation 7dd0163b-1bad-41fc-835e-eaa662995bae · outbound

This paper cites Double graph based reasoning for document-level relation extraction,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Double graph based reasoning for document-level relation extraction,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.953381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.691524Z digest=sha256:4ea5d35f92e7f25b601fac1db89494a7be0de43eea22473dbefb5c013ab58aa2

Observation 22b73191-308e-44e7-8b50-839af3353d3a · outbound

This paper cites Coreference resolution without span representations,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Coreference resolution without span representations,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.933021Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.696807Z digest=sha256:5903fbd11b584c2836fe565fa027c68a7469925cbf6712746ae1aed1a9685887

Observation f7f6b36d-0c67-497c-abdf-afe1d63d9955 · outbound

This paper cites A sequence-to-sequence approach for document-level relation extraction,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A sequence-to-sequence approach for document-level relation extraction,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.915822Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.701825Z digest=sha256:8f9cd79cca41d4f7b156c7bd60968132055d417c6f35d53d66ef3c0c42e6c463

Observation 20c8838f-bcaa-4d4a-a2b0-2e48930fb651 · outbound

This paper cites Visual attention model for name tagging in multimodal social media,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Visual attention model for name tagging in multimodal social media,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.899191Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.707133Z digest=sha256:f2fa6cd1e66a5fc65c9b3ca265532989ee27268aabd7c8b8959446592388e58e

Observation 6d315406-ac50-46b1-aaea-0ba94cd620c1 · outbound

This paper cites Adaptive co-attention network for named entity recognition in tweets,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Adaptive co-attention network for named entity recognition in tweets,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.882580Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.712237Z digest=sha256:74c4356760afc14cbbba053940213745333a890ca6c4c189f0fd655e42800418

Observation 127096dd-19e9-442e-b007-70ea11d91e36 · outbound

This paper cites A large-scale chinese multimodal ner dataset with speech clues,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A large-scale chinese multimodal ner dataset with speech clues,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.866161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.717006Z digest=sha256:fdf213cf3e3d6d0972892488cea45545603616244635775dc8c5f59b4804c3ff

Observation 21f13c5b-8483-4977-b277-ce850b99b0a7 · outbound

This paper cites Who are you referring to? coreference resolution in image narrations,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Who are you referring to? coreference resolution in image narrations,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.831366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.726731Z digest=sha256:d41d199ad022514107733f4ac98dd25be23a3579b163266ad9a8e14d4e5476a9

Observation d366b976-7698-4089-944c-07597330d6f4 · outbound

This paper cites Mnre: A challenge multimodal dataset for neural relation extraction with visual evidence in social media posts,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Mnre: A challenge multimodal dataset for neural relation extraction with visual evidence in social media posts,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.811261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.731553Z digest=sha256:f514d81a11ee1c18646a0d084827b5a62720113553fb8fc9939e38defd87a172

Observation ee792f39-f53b-48f1-8a22-1e680a79554c · outbound

This paper cites A hierarchical network for multimodal document-level relation extraction,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A hierarchical network for multimodal document-level relation extraction,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.794259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.736909Z digest=sha256:4d8a56a3d76c51cbbdb340dd67321b90ef0fb071536715ae7c7c4c39da945bd9

Observation 64464a7c-2ce7-4c67-b625-b49e06889e5a · outbound

This paper cites Grounded multimodal named entity recognition on social media,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Grounded multimodal named entity recognition on social media,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.849101Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.741883Z digest=sha256:b4b2d78f7353d3adafc9a5bae9fc786082e1658e54b468825e6fcf7c27907279

Observation 3b718cd1-8fc2-412e-be63-439595997b9c · outbound

This paper cites Semi-supervised multimodal coreference resolution in image narrations,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Semi-supervised multimodal coreference resolution in image narrations,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.775857Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.748153Z digest=sha256:1f633bf69113f78dc6b6d3c44d870fa66a500a5250c4a02751013dbbbcb9d89d

Observation f55bcf04-1a0f-4fb6-869e-cfbbf9ab1d60 · outbound

This paper cites Joint multimodal entity-relation extraction based on edge-enhanced graph alignment network and word- pair relation tagging,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Joint multimodal entity-relation extraction based on edge-enhanced graph alignment network and word- pair relation tagging,

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.759051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.753323Z digest=sha256:cbf751e267041ebecfa3c64e298f7bdde93ee756ad928fba06c5a2e47c37091a

Observation 3bff049b-ab28-422b-b1af-03f0c210678b · outbound

This paper cites Multimodal relation extraction with efficient graph alignment,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Multimodal relation extraction with efficient graph alignment,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.740778Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.759044Z digest=sha256:33c55e1313bd257531d0cb8d52e23641fd8cddc131ff5fee93c856229a42edbd

Observation 1700b069-affd-4949-a99c-3afae2238976 · outbound

This paper cites Docred: A large-scale document-level relation extraction dataset,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Docred: A large-scale document-level relation extraction dataset,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.724088Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.764038Z digest=sha256:fc0fb56f8888da76fc9a46c643e0a6482d879127c7339b69ab8c99b6f8e028b6

Observation ce4a9515-a6dd-414c-a1f1-7c77d4244199 · outbound

This paper cites Improving multimodal named entity recognition via entity span detection with unified multimodal transformer,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Improving multimodal named entity recognition via entity span detection with unified multimodal transformer,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.706299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.768999Z digest=sha256:8abd1d4d9558f1322f5cc66b6d70071b4eaad72d26099fadcfae70df00c110e5

Observation dcea31c6-3b17-4c7d-8673-8791ac23cdef · outbound

This paper cites A span-based multimodal variational autoencoder for semi-supervised multimodal named entity recognition,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A span-based multimodal variational autoencoder for semi-supervised multimodal named entity recognition,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.688592Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.774151Z digest=sha256:ac1e50687ca24ac00ea3ef15d9220d5c697393cd07d7431d2fd1689d4bf19247

Observation a925a4fa-ce05-4b19-b1dd-99d624be8349 · outbound

This paper cites Entity- level interaction via heterogeneous graph for multimodal named entity recognition,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Entity- level interaction via heterogeneous graph for multimodal named entity recognition,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.670543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.779488Z digest=sha256:0b46a63b7de79425e3d48d2e7edbb044bc78b69cd96ea195e31170c2db1e463e

Observation 0b949fab-1e68-48aa-b1b8-73d0f3058adc · outbound

This paper cites Prompt- ing chatgpt in mner: Enhanced multimodal named entity recognition with auxiliary refined knowledge,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Prompt- ing chatgpt in mner: Enhanced multimodal named entity recognition with auxiliary refined knowledge,

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.650947Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.785111Z digest=sha256:b496bed771a1a481fe1a6e1a896830b32a7413a181b4e14f05d456dbbe76f6c1

Observation 6b13dd77-f9ea-4701-b6d7-a924b15827e6 · outbound

This paper cites Gravl-bert: graphical visual-linguistic representations for multimodal coreference resolution,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Gravl-bert: graphical visual-linguistic representations for multimodal coreference resolution,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.632958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.790468Z digest=sha256:90751d28320aaed3a935ca40f62abf01fb6af59b56f8dc8fa45134195f0f7185

Observation 6a183391-336c-4d56-8247-c8f194b86873 · outbound

This paper cites Good visual guidance make a better extractor: Hierarchical visual prefix for multimodal entity and relation extraction,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Good visual guidance make a better extractor: Hierarchical visual prefix for multimodal entity and relation extraction,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.615185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.795316Z digest=sha256:1eb85acb1d74c0d8bf433d571771429d91807cc8d19d43dfd4dbfba03aae6663

Observation 289e4c81-903e-449d-9171-41c4d900c578 · outbound

This paper cites Rethinking multimodal entity and relation extraction from a translation point of view,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Rethinking multimodal entity and relation extraction from a translation point of view,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.597643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.802304Z digest=sha256:51436c487da54c0a92232757572b44f1748f33024a4e63d56638e0d0df34360a

Observation aba6abb6-d9f3-43dd-b520-2b7b4884d727 · outbound

This paper cites Information screening whilst exploiting! multimodal relation extraction with feature denoising and multimodal topic modeling,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Information screening whilst exploiting! multimodal relation extraction with feature denoising and multimodal topic modeling,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.579968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.807645Z digest=sha256:02755ea19b7fde81e2c3fab7183994bd9f036a9c6f7789a9e5de7a53854f6eee

Observation 59b3563e-06c6-444f-be96-14b66d40ae54 · outbound

This paper cites Transvg: End-to-end visual grounding with transformers,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Transvg: End-to-end visual grounding with transformers,

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.812633Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.812633Z digest=sha256:5ae519b5f4e2e033082733f54065b1e5de4e8f4490bc5fdfd605c69ebcd443eb

Observation 10d8a44e-499c-4f42-9bad-5b21ba5e3538 · outbound

This paper cites Shifting more attention to visual backbone: Query-modulated refinement networks for end-to-end visual grounding,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Shifting more attention to visual backbone: Query-modulated refinement networks for end-to-end visual grounding,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.549750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.818110Z digest=sha256:ba3cf4601462b94e5956234d19c54e2d5716e44c2dd0256fb9716ea78d7cac27

Observation b6a867e3-aab2-4351-b594-4ae5589fcb64 · outbound

This paper cites You only look once: Unified, real-time object detection,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction You only look once: Unified, real-time object detection,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.823763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.823763Z digest=sha256:262c826413a6463771c9a3f9e90a28f0a2eb36387c0808d4d885c08929d7398f

Observation e60e0871-744e-4ab5-a7c4-3cc522724eb6 · outbound

This paper cites Ssd: Single shot multibox detector,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Ssd: Single shot multibox detector,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.519786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.829059Z digest=sha256:fdb1ef7e079840c1256135a7b3ad9cf1a1079dc40bd84f5a89820cea2e7eb171

Observation 70f282fc-3a62-43f4-821c-dbd7404185a1 · outbound

This paper cites Ref-nms: Breaking proposal bottlenecks in two-stage referring expression grounding,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Ref-nms: Breaking proposal bottlenecks in two-stage referring expression grounding,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.500810Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.834799Z digest=sha256:a27d6298d2bebe73f0118badc0fcc380f43bfba969bd806ac73c5dd9b3fbfe2a

Observation c63856b6-a7dc-4429-b84e-660cddeb2d03 · outbound

This paper cites Missing modalities imputation via cascaded residual autoencoder,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Missing modalities imputation via cascaded residual autoencoder,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.481316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.839867Z digest=sha256:b7b9e9187f17e7facd00bc6cb0359f93a106baafe5814ac9374144a2704feec5

Observation 29585045-df22-4ae7-8611-bcf375eb6047 · outbound

This paper cites Lrmm: Learning to recommend with missing modalities,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Lrmm: Learning to recommend with missing modalities,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.463516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.844766Z digest=sha256:d75297209768a2f888e271f6e4b445e4b4bd7874ffe81481c5d7fd42a83a6388

Observation 3d5846a8-68ac-426e-9596-808564db30f7 · outbound

This paper cites Dealing with missing modalities in the visual question answer-difference prediction task through knowledge distillation,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Dealing with missing modalities in the visual question answer-difference prediction task through knowledge distillation,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.445404Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.850013Z digest=sha256:a31fd23bc6d092afca0687cbf7c0f413c8cc64a1c6669a043fb798fbeb4e5872

Observation 93c3b3cb-0854-4ddc-807a-2ed9696d3f8a · outbound

This paper cites A unified self-distillation framework for multimodal sentiment analysis with uncertain missing modalities,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A unified self-distillation framework for multimodal sentiment analysis with uncertain missing modalities,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.427616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.855160Z digest=sha256:1fd23c9ebba97c373504570d2370df1516640cdef60586befd9a32d216d30c94

Observation 398e2ac7-d9c4-4c6b-8fef-6f8e1c49bfbf · outbound

This paper cites Missing modality imagination network for emotion recognition with uncertain missing modalities,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Missing modality imagination network for emotion recognition with uncertain missing modalities,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.860148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.860148Z digest=sha256:b0022a5330bc5f4fe5c022566db4c199854cbc1d3de570d75304cde247c07fc3

Observation 3d58a6d9-b30a-42cb-a99e-9020de48ed42 · outbound

This paper cites Mitigating inconsistencies in multimodal sentiment analysis under uncertain missing modalities,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Mitigating inconsistencies in multimodal sentiment analysis under uncertain missing modalities,

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.865429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.865429Z digest=sha256:8b668b5ebf7ddb5ce15147cc51e36088169ede51915526a303936fb259f63903

Observation 39c5209d-1341-4f87-a005-90bf62d45817 · outbound

This paper cites Multimodal prompting with missing modalities for visual recognition,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Multimodal prompting with missing modalities for visual recognition,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.386234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.870407Z digest=sha256:7ffd42a990f950c19bb595427c54bfda35dc1163da44847b9827cfdecc49e881

Observation 1f90a499-dbcd-4942-8037-355dd9c53f14 · outbound

This paper cites Multimodal prompt learning with missing modalities for sentiment analysis and emotion recognition,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Multimodal prompt learning with missing modalities for sentiment analysis and emotion recognition,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.368424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.875900Z digest=sha256:fd526a2a17ba1dbb3082c500c5a98211bee704601d75a14739e0851a74cb962b

Observation 29abe825-40e1-4009-826e-5e293f1e5201 · outbound

This paper cites Longformer: The Long-Document Transformer.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Longformer: The Long-Document Transformer

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.881703Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.881703Z digest=sha256:7b11f20155f09723f4f95b73fccc02e858ec033365238575613a8e42fa9856f9

Observation 2fd49c31-be75-4af4-84ec-3e9a8d5c17d2 · outbound

This paper cites Visual Transformers: Token-based Image Representation and Processing for Computer Vision.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Visual Transformers: Token-based Image Representation and Processing for Computer Vision

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.887844Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.887844Z digest=sha256:9257140cf950e2619211dda3c5cd1e2f9c4f8a4db988283ffd5a240bf10816f1

Observation e36b2296-2cc4-4840-883f-2255bf452579 · outbound

This paper cites A primer in bertology: What we know about how bert works,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A primer in bertology: What we know about how bert works,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.349562Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.894087Z digest=sha256:dc6154ac2378bc821b79fef424b9a482e203100446e0880a421767ec66a83071

Observation 548d5047-ec8e-48f8-9686-4e80da3b10bc · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Bert: Pre-training of deep bidirectional transformers for language understanding,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.328333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.899056Z digest=sha256:b9492e96a79f0acdccaa676dda6b5318ffac30fbddec3cae6c108961e07a1736

Observation 9696bcfe-f93b-4c3c-aff7-4a3c73dba5d9 · outbound

This paper cites An image is worth 16x16 words: Transformers for image recognition at scale,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction An image is worth 16x16 words: Transformers for image recognition at scale,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.309219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.904139Z digest=sha256:c0c29b1d715e99c1dcf594f1c6079c4d66a0f146a85391494b785cb7661f1026

Observation 3609f568-0b54-49d9-94c9-774cda0cecc7 · outbound

This paper cites Auto-encoding variational bayes,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Auto-encoding variational bayes,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.289980Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.910780Z digest=sha256:67544b74e6146703a01052b806be3d8c1d4136249d753b272adbb8cb0ef79918

Observation c8b2cd50-66c0-47ab-a728-8d61aeeca6a8 · outbound

This paper cites Exploring universal intrinsic task subspace for few-shot learning via prompt tuning,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Exploring universal intrinsic task subspace for few-shot learning via prompt tuning,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.267587Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.915496Z digest=sha256:fd4a8fa588deb19cd5eb3429a16cb1d6c31955e0583c4d611f9763710f685bb1

Observation 53c937e7-2558-431c-a72b-ec0c375c9853 · outbound

This paper cites Attention is all you need,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Attention is all you need,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.247428Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.920736Z digest=sha256:0a3520cd72d7f2dcfa82f551b69077f766af9204ae97f73bd680352b66f0ddd5

Observation f589483d-954d-4cfd-a5eb-b4ef49dc5c17 · outbound

This paper cites Deep residual learning for image recognition,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Deep residual learning for image recognition,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.925634Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.925634Z digest=sha256:93c8c962ac1896361d58f923b184e868ec5d4ce4e52d7746fe7183f1de9226e9

Observation c679d900-ee18-442a-b05f-a707877b0d08 · outbound

This paper cites Conditional random fields: Probabilistic models for segmenting and labeling sequence data,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Conditional random fields: Probabilistic models for segmenting and labeling sequence data,

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.215489Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.931019Z digest=sha256:7e25b98660ebc429558ff39113a41d9dc84d9499c2c0b6f0a585c27ed3e2c865

Observation 1a07ebe0-1223-4048-bec1-804043acef90 · outbound

This paper cites VisualBERT: A Simple and Performant Baseline for Vision and Language.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction VisualBERT: A Simple and Performant Baseline for Vision and Language

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.935726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.935726Z digest=sha256:a2119957cc4ac103f7eecf94877e88882335632fd008a88f19a7478ab39d9bad

Observation 327576fc-5332-447d-a2c5-cdbbc9b28a7f · outbound

This paper cites Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Vilbert: Pretraining task-agnostic visiolinguistic representations for vision-and-language tasks,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.940778Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.940778Z digest=sha256:47b21364a1a469e79802a06a6f49770915c7c5aacf2c74959f35690be50e08ac

Observation dc33cd84-8a54-47fb-be2e-8c7e25261b2f · outbound

This paper cites Vivit: A video vision transformer,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Vivit: A video vision transformer,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.945512Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.945512Z digest=sha256:6e57ce39bfd41cad954edf4257ab0cc8036a5132aa37250cb6977f4ec0f65d53

Observation f4c4426e-de7d-41ac-966a-d60a7f21d47e · outbound

This paper cites Videomae: Masked autoen- coders are data-efficient learners for self-supervised video pre-training,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Videomae: Masked autoen- coders are data-efficient learners for self-supervised video pre-training,

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.950362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.950362Z digest=sha256:8ba4d680722a15986f85889d479e0748925f326fa5ef199f4aa073ab58acd7ff

Observation e0ea64b8-955c-4ba8-972f-e48df9251217 · outbound

This paper cites Video-llama: An instruction-tuned audio- visual language model for video understanding,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Video-llama: An instruction-tuned audio- visual language model for video understanding,

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.955213Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.955213Z digest=sha256:dbdd47994f0716d6c0130c810a9d5295551889381312ff04a7c15d0207101679

Observation dd061f07-0ff1-49b2-981b-b3c6017764f1 · outbound

This paper cites Video-chatgpt: Towards detailed video understanding via large vision and language models,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Video-chatgpt: Towards detailed video understanding via large vision and language models,

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.132655Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.960826Z digest=sha256:6493e8e89240ea745b62c2444dd8720247e712ae8a9c3281fce2ed158dbb0e0c

Observation 060b76a4-32cf-41e9-a904-9eddfe4745bb · outbound

This paper cites Parallel data helps neural entity coreference resolution,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Parallel data helps neural entity coreference resolution,

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.115435Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.965982Z digest=sha256:71a4352a265e41b71c311a23a428633973c38f4334c1489d7042ce1ec5c42236

Observation a6c55131-0d59-431e-ba0e-6846f88eba82 · outbound

This paper cites Toe: A grid-tagging discontinuous ner model enhanced by embedding tag/word relations and more fine-grained tags,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction Toe: A grid-tagging discontinuous ner model enhanced by embedding tag/word relations and more fine-grained tags,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-11T21:53:56.095439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-11T21:53:55.971683Z digest=sha256:280b3b0b904240433a2f5df7331fb81836d5ce83593b2e6534f2815e85ff0714

Observation acb1842e-afd2-49d1-9f2f-3db8a1d6502e · outbound

This paper cites A fast and accurate one-stage approach to visual grounding,.

M$^{3}$D: A Multimodal, Multilingual and Multitask Dataset for Grounded Document-level Information Extraction A fast and accurate one-stage approach to visual grounding,

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-11T21:53:55.976780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T21:53:55.976780Z digest=sha256:0b6f62977a7db303b2956c1f243fe65233330a651f02257eabb386927cc177dd

Pith citing papers

No inbound Pith citation observations are available.