Pith. sign in

Paper Citation Record · LEDGER

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning

As of 22 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 0 inbound Pith citation observations for arXiv:2506.16082.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.16082 v1

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T23:49:30.743524Z

measured 68 of 68 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

68 of 68 outbound references displayed

  • verified exact1
  • verified fuzzy45
  • unresolved22
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fa5e75c7-e623-44cf-a38d-39cffd799f59 · outbound

This paper cites Swinbert: End-to-end transformers with sparse attention for video captioning,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Swinbert: End-to-end transformers with sparse attention for video captioning,

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.281463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.533987Z digest=sha256:d34bceea76436fbc837f34b8455c2a563c1bf81ae92b96da90c6b5baee33c1f3

Observation 1d89fd1c-5d43-41f3-b082-3be8c3e7a0bd · outbound

This paper cites Univl: A unified video and language pre-training model for multimodal understanding and generation,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Univl: A unified video and language pre-training model for multimodal understanding and generation,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.274458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.538020Z digest=sha256:483e28832139a0f0c3e56ab0aa543294c5b68e0cd62c125c8165ca7eecaf2e63

Observation 44fa5252-8e30-4483-9f19-d2bcfec2807d · outbound

This paper cites End-to-end generative pretraining for multimodal video captioning,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning End-to-end generative pretraining for multimodal video captioning,

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.267518Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.541611Z digest=sha256:fbbb3b57ce59aab139f07cdac9f7be69334eb8ad41fbd01cfc83e38e2ca171a4

Observation e4831189-774a-4eac-8d97-886958a8f233 · outbound

This paper cites Memory- attended recurrent network for video captioning,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Memory- attended recurrent network for video captioning,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.260793Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.545285Z digest=sha256:89b3239da75e3733491dc11a44dd0328c4cbbf1bb07c595b01fdae8eb1f61cee

Observation 113ddb04-1b16-413b-99b6-452427e883e1 · outbound

This paper cites Sports video captioning via attentive motion representation and group relationship modeling,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Sports video captioning via attentive motion representation and group relationship modeling,

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.253789Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.548825Z digest=sha256:6f4adaa94c7f50b82ac5818b789446d2449ddd0171c133a6d8a6e7b363f27edd

Observation d43601c3-4508-4875-bfd2-808817b94359 · outbound

This paper cites Reconstruction network for video captioning,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Reconstruction network for video captioning,

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.246803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.551991Z digest=sha256:35cfeffbb8596aa1765471625a90830ec3589efcf050139099249622172f57f3

Observation 10703529-50cb-490d-9161-a037c9997068 · outbound

This paper cites Concept-aware video captioning: Describing videos with effective prior information,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Concept-aware video captioning: Describing videos with effective prior information,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.238970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.555325Z digest=sha256:9a71e25c95a6ba2575201148b016dc6f797708daea5163054e5f66b882ff707e

Observation ae159af9-be8a-438c-b1f6-4b7a4d07989c · outbound

This paper cites Dense- captioning events in videos,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Dense- captioning events in videos,

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.227919Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.559277Z digest=sha256:16000ec9acdf263f3799744db77b3e4fe6432ac972b5f5d8e1d9a6a53bf999b9

Observation cfeb5a98-db2d-42d5-9589-5adad7828359 · outbound

This paper cites Jointly localizing and describing events for dense video captioning,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Jointly localizing and describing events for dense video captioning,

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.219010Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.562266Z digest=sha256:cf82917a31ec4d4d1ed6ed154f176e80f4b4686c30d80dbf06d8f244dd08320f

Observation 1822c6d2-42aa-4df6-b921-e077a0235f16 · outbound

This paper cites End-to- end dense video captioning with parallel decoding,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning End-to- end dense video captioning with parallel decoding,

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.209998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.565004Z digest=sha256:f6fca9517a684fbaea0d0a6da7052c741b6670a86becb33ec6a1d32549f13c10

Observation 11f72725-4a3f-4343-a00f-0edd60a4472f · outbound

This paper cites Vid2seq: Large-scale pretraining of a visual language model for dense video captioning,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Vid2seq: Large-scale pretraining of a visual language model for dense video captioning,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.200147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.568648Z digest=sha256:6da9acd7e04c3bae62ed183f19b6252f41f2a1ace50f4d33e6c8e515bfe28b96

Observation 542d302d-477c-4189-895d-7c8ed67132f5 · outbound

This paper cites Cap4video: What can auxiliary captions do for text-video retrieval?.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Cap4video: What can auxiliary captions do for text-video retrieval?

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.191147Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.571989Z digest=sha256:4a62f5776147774974e538769abfe5c43f800d950f3f67eb82e8544feb53a9c3

Observation d891ad30-fab8-40a4-8db5-83a0afcbc4d0 · outbound

This paper cites Multi-event video-text retrieval,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Multi-event video-text retrieval,

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.182933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.574902Z digest=sha256:0608b671061c1aec2a6355b00de51862c84fc6d33b8ba7f46c9bc3ef47e5684f

Observation f2c9c0eb-2240-4a09-8218-19fdfc163885 · outbound

This paper cites Exploiting unlabeled videos for video-text retrieval via pseudo-supervised learning,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Exploiting unlabeled videos for video-text retrieval via pseudo-supervised learning,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.174629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.577717Z digest=sha256:b75b5ea17b35360ec59316facf542e1319bc2fdbbfb1fa402992d28b6cd474cd

Observation 4e662bfc-e59e-4372-9249-c8be4ba23f18 · outbound

This paper cites Video recap: Recursive captioning of hour-long videos,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Video recap: Recursive captioning of hour-long videos,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.165632Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.580462Z digest=sha256:71a2684330797d519ac983249e2a65ca6822bd0567ef30e1a5376d47fc9f484a

Observation 5699f10b-f5c1-46ad-b6f2-758d831f54c2 · outbound

This paper cites Vidchapters- 7m: Video chapters at scale,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Vidchapters- 7m: Video chapters at scale,

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.156243Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.583741Z digest=sha256:d25a5d61cdf6eb155c5491b3069786ae76361978676280fb7f7b47e8a1de7218

Observation 0bc1ec0c-a38e-4e8d-b366-121891cef46e · outbound

This paper cites Hierarchical representation network with auxiliary tasks for video captioning and video question answering,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Hierarchical representation network with auxiliary tasks for video captioning and video question answering,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.147504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.586459Z digest=sha256:ca14f17b9fbc0cae8c063eb9051dfd1566cdafd792c2bc7a5eef63004065a849

Observation ad283bb7-0429-447c-b561-e24c5efdff8a · outbound

This paper cites Multi-modal dense video captioning,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Multi-modal dense video captioning,

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.138529Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.589292Z digest=sha256:229d5f645ad98d3cc22443d65eb8d6dce4b2a426fd083f5b2095ce9a48536ee4

Observation ab34a21b-64be-463e-a64b-1a9239f2c474 · outbound

This paper cites A Better Use of Audio-Visual Cues: Dense Video Captioning with Bi-modal Transformer.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning A Better Use of Audio-Visual Cues: Dense Video Captioning with Bi-modal Transformer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:30.592133Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:30.592133Z digest=sha256:04555b435343256733445097dc14091cac1c1c5797194da58cf1d99c89210f17

Observation 1b4428d7-f918-40f0-af3a-9538660fe4e1 · outbound

This paper cites Hierarchical context encoding for events caption- ing in videos,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Hierarchical context encoding for events caption- ing in videos,

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.130059Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.595745Z digest=sha256:08f7aff4dace04565991195ae98adf0bbb60897402b3b9d14f50e28692e76e0d

Observation 3aed2096-8ddd-4045-a9a1-361f48bb4193 · outbound

This paper cites End-to-end object detection with transformers,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning End-to-end object detection with transformers,

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:30.598874Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:30.598874Z digest=sha256:9b55d2b679eee05a63d66bc62859902361eddfadba829d86936dab6e764657aa

Observation 794f98ca-fcfe-46b4-94b1-bf96aafb570d · outbound

This paper cites Do you remember? dense video captioning with cross-modal memory retrieval,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Do you remember? dense video captioning with cross-modal memory retrieval,

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.116790Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.601360Z digest=sha256:26d598e4b3cf831aa5769ffa28805225951a8c0f45a40d39766de3e5f13ec03e

Observation f70b8490-cd78-443e-bf6e-3592beb5a3a6 · outbound

This paper cites Parallel pathway dense video captioning with deformable transformer,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Parallel pathway dense video captioning with deformable transformer,

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.108621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.604312Z digest=sha256:96a3db3892f70cd3ca78ec133fd11a6fcd3d6966a4f21c017b2b1a1dee54b678

Observation 2a23eb09-4550-4e63-901e-be803f5120dc · outbound

This paper cites Dibs: Enhancing dense video captioning with unlabeled videos via pseudo boundary enrichment and online refinement,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Dibs: Enhancing dense video captioning with unlabeled videos via pseudo boundary enrichment and online refinement,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.100504Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.607160Z digest=sha256:2510faacba5c744773434bfcab80f108d68b4fabbefd695d68d8681943397a5b

Observation 0f88833e-93cf-45b4-90b9-7002b9add345 · outbound

This paper cites Towards automatic learning of procedures from web instructional videos,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Towards automatic learning of procedures from web instructional videos,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.092046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.609759Z digest=sha256:490dafc2526c7377dcc762ad7c13f97dc36f7c33277343d780e3b5ce834da3a6

Observation ca198214-842b-4c2c-9d69-00f1bbf8c895 · outbound

This paper cites Bidirectional attentive fusion with context gating for dense video captioning,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Bidirectional attentive fusion with context gating for dense video captioning,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.081982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.612617Z digest=sha256:79804768117f48d6f46387295edbc120c1537508610db6552b3fbed6b81c5d81

Observation 67355e09-8a00-4481-bef1-1ce32a7536e0 · outbound

This paper cites Sketch, ground, and refine: Top-down dense video captioning,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Sketch, ground, and refine: Top-down dense video captioning,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.063355Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.619718Z digest=sha256:852779074e3e03485c3a711ac0e80e7e35a7969f2dad9ff27c7edb76fd3e7384

Observation 5ed51ade-4dbb-4174-b9d4-1445d517539f · outbound

This paper cites iPerceive: Applying Common-Sense Reasoning to Multi-Modal Dense Video Captioning and Video Question Answering.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning iPerceive: Applying Common-Sense Reasoning to Multi-Modal Dense Video Captioning and Video Question Answering

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-08-06T23:49:30.817671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.622323Z digest=sha256:916f9b2c640eaeff37117386aafdfafde1078cdf6db4a90396908798919142a3

Observation 790c75d0-3028-4619-9d21-93b99435fccb · outbound

This paper cites Towards bridging event captioner and sentence localizer for weakly supervised dense event captioning,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Towards bridging event captioner and sentence localizer for weakly supervised dense event captioning,

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.054536Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.625464Z digest=sha256:aea4662bedd1a35beb9be0ad4b850f05955057febc3ed3e4e1ce73e0c33fb78f

Observation ec21ce32-6afd-4ff0-b1c6-82e02b299266 · outbound

This paper cites Streamlined dense video captioning,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Streamlined dense video captioning,

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.045077Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.629079Z digest=sha256:7c6a1e16de811c38bbc0c5ef5185eb649d1b6220304197ccd3e0aa2a32d4e229

Observation 31f8ffc7-98bc-439b-9cf4-32cdbbc7a81d · outbound

This paper cites Watch, listen and tell: Multi- modal weakly supervised dense event captioning,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Watch, listen and tell: Multi- modal weakly supervised dense event captioning,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.036326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.632967Z digest=sha256:e4cc69f4d5498cee688abd38b036ff3ded166ee2a08caaecb8873bf35ad12969

Observation ef6ff805-c600-4479-8e8e-ce78bfbff472 · outbound

This paper cites Dense procedure captioning in narrated instructional videos,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Dense procedure captioning in narrated instructional videos,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.028261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.636182Z digest=sha256:de00a4dc23a86c54e891a5f7fae05839832bb9f10b9a1e35bf51f7e236533c0f

Observation a99fc2d2-df09-4f3c-8b8a-7e48293598ba · outbound

This paper cites Retrieval- augmented generation for knowledge-intensive nlp tasks,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Retrieval- augmented generation for knowledge-intensive nlp tasks,

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:30.639410Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:30.639410Z digest=sha256:4d942c836b20da9f131a0b95b605d54a5544bc65aca3822c05fe2774f8ccf261

Observation 01aac300-6314-4204-b877-fb2c2bf0ad17 · outbound

This paper cites Streaming dense video captioning,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Streaming dense video captioning,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.014998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.642357Z digest=sha256:ef066675626244055ccb5a315e87c82904a953a78305be687965b15cb9a085b8

Observation f24b0380-e52d-42a3-9f68-cff204ddeec8 · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:30.645161Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:30.645161Z digest=sha256:e6564a3e8958a5d1b9e289b675d56f20be983d3413e4a58d03af5517977ea024

Observation b8abfb3a-ee01-41e9-aca8-7db98dba97f2 · outbound

This paper cites Learning texture transformer network for image super-resolution,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Learning texture transformer network for image super-resolution,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.006881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.647963Z digest=sha256:43fd9d27d00a6e420d255683dcef8c4c370f27a59301f270b7406fc740f6ccce

Observation f755325a-cf61-4d1f-8217-8c1c20993a7f · outbound

This paper cites Siamese-detr for generic multi-object tracking,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Siamese-detr for generic multi-object tracking,

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:30.998262Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.650996Z digest=sha256:0995cb2f2385a628fe9f1bad65d82cb0324b34e53722641e7b6d694c2c9705a3

Observation 655c5caf-df19-4c6f-82d8-1c05d1e651d4 · outbound

This paper cites Depth anything: Unleashing the power of large-scale unlabeled data,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Depth anything: Unleashing the power of large-scale unlabeled data,

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:30.653780Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:30.653780Z digest=sha256:fd08e1b5c800ce9aac41aec45742eac66568e9df8af4f8f09c89f9e7a7ffad88

Observation 06e03206-93a2-498f-9b16-467cad8297a0 · outbound

This paper cites Spectralgpt: Spectral remote sensing foun- dation model,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Spectralgpt: Spectral remote sensing foun- dation model,

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:30.656528Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:30.656528Z digest=sha256:ece15c0b9ab6adb7f55db3fddb899bd4f8128b458bd3e81856d23777e1c735a1

Observation 2f31e8ee-f887-4a7a-96eb-47f4b62d8179 · outbound

This paper cites Segment anything in medical images,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Segment anything in medical images,

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:30.659718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:30.659718Z digest=sha256:9bb0244b2c4e3ef1189eefacf2593b07882893cf8c72867d3d158caecaa5b8d5

Observation e83cc329-7a9f-4831-a30e-7d562858aaaa · outbound

This paper cites Point to set similarity based deep feature learning for person re-identification,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Point to set similarity based deep feature learning for person re-identification,

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:30.976259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.663035Z digest=sha256:f3ce32aca2a6bcc016ddfc52e24b4690565ccece54618eeb0adde6f552926ef4

Observation 2a888963-af32-49ba-a627-5475c9bec640 · outbound

This paper cites Visual-linguistic feature align- ment with semantic and kinematic guidance for referring multi-object tracking,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Visual-linguistic feature align- ment with semantic and kinematic guidance for referring multi-object tracking,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:30.967527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.666073Z digest=sha256:7d67fd153571053b5a347a09d55ef10e44bac57fc82d90a60acd3031b40a8d8b

Observation ca1d6cf8-70dc-426e-840d-56823da30168 · outbound

This paper cites Explainability enhanced object detection transformer with feature disentanglement,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Explainability enhanced object detection transformer with feature disentanglement,

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:30.957909Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.669729Z digest=sha256:1b9391aca39a1708d826cf4981fe94706fd2126cf61befe73589964923fae4ba

Observation 0992a42d-65ad-4aba-a10c-c5cd0f1dfdc8 · outbound

This paper cites Deformable DETR: Deformable Transformers for End-to-End Object Detection.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Deformable DETR: Deformable Transformers for End-to-End Object Detection

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:30.672800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:30.672800Z digest=sha256:8ebfdd63d5bec196d3e01ea9d832df0fd7fde69388c3d5e403420d916762e45b

Observation 9095b9ae-19f1-417b-93fe-41701a75a88f · outbound

This paper cites Fast temporal activity proposals for efficient detection of human actions in untrimmed videos,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Fast temporal activity proposals for efficient detection of human actions in untrimmed videos,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:30.949177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.675862Z digest=sha256:3d4a9d908ac1c01a0119229394c240b55824c83af0c84871aef75cb2d27277b5

Observation 99bd5ad5-0eee-4749-8580-f048a6992829 · outbound

This paper cites Turn tap: Temporal unit regression network for temporal action proposals,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Turn tap: Temporal unit regression network for temporal action proposals,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:30.940183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.678711Z digest=sha256:b28c6f930740d1f4a785e37092c7b756616a6db40a0f9e6bd846e29327d2904f

Observation 2b2f46a3-1035-4d4e-a0ba-a3755a11b930 · outbound

This paper cites Daps: Deep action proposals for action understanding,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Daps: Deep action proposals for action understanding,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:30.931901Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.681671Z digest=sha256:05d5a51529ec8a96b6ee034607a875eba47789ca4927ab34abad92629885f259

Observation 3964fccf-844e-4393-8b9e-5d074e022e84 · outbound

This paper cites Bmn: Boundary-matching network for temporal action proposal generation,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Bmn: Boundary-matching network for temporal action proposal generation,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:30.923491Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.684841Z digest=sha256:fd09ecc64375d3c06e900c3f7f28065fb1ea36280c8be853c186664f5e3fcfa2

Observation 6a62d0aa-f909-4ba9-adb5-d9e489ea7341 · outbound

This paper cites Bsn: Boundary sensitive network for temporal action proposal generation,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Bsn: Boundary sensitive network for temporal action proposal generation,

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:30.913477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.687518Z digest=sha256:f376b153ed628f9e79f519cef5cbe9166748c95432b0a1ccc0870790787c1ecc

Observation fff87b8a-cc00-4133-a777-d22d4d091316 · outbound

This paper cites Temporal action detection with structured segment networks,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Temporal action detection with structured segment networks,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:30.904862Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.690224Z digest=sha256:e69fe66084c46cd5d16607ee94489a0f74d3a211e1b4d7d8d7436620c0212a5a

Observation e99f9f09-0088-4c18-ab25-7fc4378b8618 · outbound

This paper cites Bert: Pre-training of deep bidirectional transformers for language understanding,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Bert: Pre-training of deep bidirectional transformers for language understanding,

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:30.693135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:30.693135Z digest=sha256:661c8d914c72d370f1d6ceb52697df459f084a946387e686af7fa018e65381a8

Observation 26838943-bc32-48b3-a35b-11a34589037b · outbound

This paper cites Learning transferable visual models from natural language supervision,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Learning transferable visual models from natural language supervision,

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:30.695794Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:30.695794Z digest=sha256:723b40fb619ac93b1fbe19c5c7f43a21852e0b8660fab679fa91fb2fec3fe3da

Observation d2378daf-4e9e-4f59-8b20-ba5aae863ffe · outbound

This paper cites Attention is all you need,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Attention is all you need,

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:30.698590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:30.698590Z digest=sha256:0a357f73d83068e5151a635176b3b7a8f7b89e4443902dd0ceb5d2efde695f10

Observation dad2060a-0920-499d-8779-5c23d59dc7e8 · outbound

This paper cites The hungarian method for the assignment problem,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning The hungarian method for the assignment problem,

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:30.702059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:30.702059Z digest=sha256:adc14367c6285f2024b42732b72c6c4ca3e0f676d8732bf5b04de574a889b418

Observation 4a2d00ab-e957-48be-85e3-7dd8d9569679 · outbound

This paper cites Generalized intersection over union: A metric and a loss for bounding box regression,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Generalized intersection over union: A metric and a loss for bounding box regression,

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:30.705186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:30.705186Z digest=sha256:06bc58c4f86c42e56a394a41f82d8bf000a12a66b7c0ac408348b2d4941545de

Observation 28fb2958-0a4d-4bef-b95f-b53ab4c74fca · outbound

This paper cites Focal loss for dense object detection,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Focal loss for dense object detection,

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:30.707762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:30.707762Z digest=sha256:d58e4191290802e45fd5531b009cf7f614da462f961ff783dd3eeae8c8d4defc

Observation 02b0983d-d264-4de7-a0a4-78a13078d2d4 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:30.710577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:30.710577Z digest=sha256:c30551fc80f677c8316768918a154121af0d0cf74be1f1428a59fe5100a3c5fd

Observation d0abc37a-0e0c-421d-b03e-a92d76fd7b81 · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning VideoChat: Chat-Centric Video Understanding

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:30.713603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:30.713603Z digest=sha256:3b2b35fa183401a139fe71a510833ee63e876af82a668b4d26df93d1cbb12c33

Observation cf0a8937-3623-44e7-970f-abd7ac383a08 · outbound

This paper cites Timechat: A time-sensitive multimodal large language model for long video understanding,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Timechat: A time-sensitive multimodal large language model for long video understanding,

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:30.716783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:30.716783Z digest=sha256:5ba3b14d499e19326cf9e773940dad1840f24c04be422262fe28d872672cbd98

Observation b96baa8c-f98f-4b68-aac6-179d53c9f81e · outbound

This paper cites End-to-end Dense Video Captioning as Sequence Generation.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning End-to-end Dense Video Captioning as Sequence Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:30.719432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:30.719432Z digest=sha256:4d08f090b7de7b495528a028710ffcf8fad7707a1afc51a3ef0246aec9903cb9

Observation add7b680-b287-4f8a-8d02-c4296cecaa0a · outbound

This paper cites Event-centric hier- archical representation for dense video captioning,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Event-centric hier- archical representation for dense video captioning,

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:30.865177Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.722485Z digest=sha256:81789af3414c2eecb81774ad7be14c2829685a01a4724816a8207d6821fb0edf

Observation b5151edd-cabd-4bd3-8084-49f77d5a5779 · outbound

This paper cites End-to-end dense video captioning with masked transformer,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning End-to-end dense video captioning with masked transformer,

Reference 63

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:31.073170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.725394Z digest=sha256:025af7b1e1bbcc516c16f3a4599d30b9a1d8e96ce2270a73cdd40e8aa3b14f30

Observation 739b80c0-5372-437d-932c-593f7a4026f5 · outbound

This paper cites Dense-Captioning Events in Videos: SYSU Submission to ActivityNet Challenge 2020.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Dense-Captioning Events in Videos: SYSU Submission to ActivityNet Challenge 2020

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:30.728030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:30.728030Z digest=sha256:b7a44a17b107c23e4ba922cf234331472caba17a1fd4b9e801338d23edc3bb9d

Observation 93559e2d-a349-4524-9cea-c70b6032fe40 · outbound

This paper cites Cider: Consensus- based image description evaluation,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Cider: Consensus- based image description evaluation,

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:30.731021Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:30.731021Z digest=sha256:fcf81940508b1285ef40aeb63b512801d7784222a556c761adc248887c575a49

Observation 9ef41eb8-1756-4316-bd3d-1e91b39e8633 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Bleu: a method for automatic evaluation of machine translation,

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:30.734331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:30.734331Z digest=sha256:3b8d8135e1556ca34067442989e156d02be1dd296b5feaf0d6b7e9a4b7983a0c

Observation 6ddaf0f5-115d-49e5-8bf2-e35f1bde38af · outbound

This paper cites Meteor: An automatic metric for mt evalua- tion with improved correlation with human judgments,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Meteor: An automatic metric for mt evalua- tion with improved correlation with human judgments,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:30.737908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:30.737908Z digest=sha256:c45db7c65543c5f99b9792aaa788c35593e0856605e2594dffde1994c131b64f

Observation 2d775db5-68d8-4dba-8e81-96629aa209ef · outbound

This paper cites Soda: Story oriented dense video captioning evaluation framework,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Soda: Story oriented dense video captioning evaluation framework,

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:30.842462Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.740876Z digest=sha256:f499ac4f63f1c749ce89349dad688198e0a0f60c7908267a7cb35e15c26975cc

Observation 560e4972-7cc4-490a-b7fa-6460964bbfeb · outbound

This paper cites Move forward and tell: A progressive generator of video descriptions,.

PR-DETR: Injecting Position and Relation Prior for Dense Video Captioning Move forward and tell: A progressive generator of video descriptions,

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T23:49:30.834153Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-06T23:49:30.743524Z digest=sha256:d27360486b73556a03031ed7126ba729ea7633b171b32e5de44a0208a9cfa5d0

Pith citing papers

No inbound Pith citation observations are available.