Pith. sign in

Paper Citation Record · LEDGER

MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding

As of 10 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 7 inbound Pith citation observations for arXiv:2104.12763.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2104.12763 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 7 of 7 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 7 of 7 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:22:37.398957Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-24T11:09:22.426941Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation a2219cda-1b5c-4460-976a-a4b5fa618fa1 · inbound

DetailCLIP: Injecting Image Details into CLIP's Feature Space cites this paper.

DetailCLIP: Injecting Image Details into CLIP's Feature Space MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-24T11:09:22.430430Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-24T11:08:20.298043Z digest=sha256:7793990a320b85af88de0cf9f2d55cc8779e2b858330de68a9e22fa14724504f

Observation e5177977-f620-44b3-af16-da781b15a257 · inbound

Grounding-Aware Token Pruning: Recovering from Drastic Performance Drops in Visual Grounding Caused by Pruning cites this paper.

Grounding-Aware Token Pruning: Recovering from Drastic Performance Drops in Visual Grounding Caused by Pruning MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T22:22:37.398957Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:22:37.398957Z digest=sha256:5adc8dc9d1e56ac2ee7937cdf1a2d542ae34a6cd59f0baf875dbd442a4640811

Observation d0c7ffae-0651-4c6d-88ca-aa9489c7923b · inbound

Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset cites this paper.

Smelly, dense, and spreaded: The Object Detection for Olfactory References (ODOR) dataset MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-06T18:25:20.637769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:25:20.637769Z digest=sha256:7b6239df6b7332c51563871bad86164235f4e5ae04220e41131fb1fd1052e853

Observation dd15ed5d-d967-417b-8ed0-c9edfe0d9bd8 · inbound

STORM: End-to-End Referring Multi-Object Tracking in Videos cites this paper.

STORM: End-to-End Referring Multi-Object Tracking in Videos MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:56:00.501349Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T16:25:31.777907Z digest=sha256:462bbf9352577927616dbe7b932880fc9d71709e8fbabb524b93e1fa7f03d75b

Observation ec4bf94c-242f-46d6-8637-a030cbb128e8 · inbound

AutoVQA-G: Self-Improving Agentic Framework for Automated Visual Question Answering and Grounding Annotation cites this paper.

AutoVQA-G: Self-Improving Agentic Framework for Automated Visual Question Answering and Grounding Annotation MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-10T06:36:36.455044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-10T06:35:44.016054Z digest=sha256:3bdd31c631ddd240e543b76508cb34c65c35614e5d72c298efeeb796d5db6acb

Observation e4e15ac1-93fa-4776-b7e4-8f2b71b184a1 · inbound

TIGER-FG: Text-Guided Implicit Fine-Grained Grounding for E-commerce Retrieval cites this paper.

TIGER-FG: Text-Guided Implicit Fine-Grained Grounding for E-commerce Retrieval MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding

Reference 8

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T23:57:53.190292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-19T23:54:32.646978Z digest=sha256:d69ff0df3dbf950b5c4bcd814e5f4121e441c11af5992b667681a15597dd0058

Observation 59a5792e-d8c7-485b-a0e8-c3f2bce2d99d · inbound

Vision Harnessing Agent for Open Ad-hoc Segmentation cites this paper.

Vision Harnessing Agent for Open Ad-hoc Segmentation MDETR -- Modulated Detection for End-to-End Multi-Modal Understanding

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-20T05:53:04.442806Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-20T05:52:40.429412Z digest=sha256:66a6cb23c7aeae48f83bbaf944ff7854a97a59e02cc9edea9eec538ab8217c53