Pith. sign in

Paper Citation Record · LEDGER

MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos

As of 20 August 2026, this Paper Citation Record lists 13 of 13 outbound references and 0 inbound Pith citation observations for arXiv:2506.12623.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.12623 v1

Coverage vector

measured 13 of 13 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T00:48:30.581939Z

measured 13 of 13 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

13 of 13 outbound references displayed

  • verified exact0
  • verified fuzzy5
  • unresolved6
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 34ea7804-2841-406c-a928-502d210f5a08 · outbound

This paper cites BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension.

MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos BART: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:29.881132Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:29.881132Z digest=sha256:44a915cc3f13120588b7f6a71686798e53b8073d9355f2eccaea9c7914c63c38

Observation bf0cc22b-2fe0-4413-a07d-e2ec1215f176 · outbound

This paper cites InProceedings of the 2020 Con- ference on Empirical Methods in Natural Language Processing (EMNLP), pages 4707–4716, Online.

MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos InProceedings of the 2020 Con- ference on Empirical Methods in Natural Language Processing (EMNLP), pages 4707–4716, Online

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:31.404579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:48:29.990078Z digest=sha256:49a1c7be1a44c228c543d699b31df5cc745c902f111567296b9e195428fdc71f

Observation 795be711-6f3f-4c84-a376-8019c03f01d3 · outbound

This paper cites MMSum: A Dataset for Multimodal Summarization and Thumbnail Generation of Videos.

MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos MMSum: A Dataset for Multimodal Summarization and Thumbnail Generation of Videos

Reference 10

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T00:48:30.791494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:48:30.247861Z digest=sha256:6b42abe5c9ec56fe5081291853cc2dadb0b271b573350ddaa820d37b6456c196

Observation 2de4cfa6-e2a6-4e58-af28-867351195999 · outbound

This paper cites How2: A Large-scale Dataset for Multimodal Language Understanding.

MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos How2: A Large-scale Dataset for Multimodal Language Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.364282Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.364282Z digest=sha256:7159a1184e553dce1334047d02443a8c8d35e2cc2449ca01ce88ad19c466c2e8

Observation 5f5296aa-02e4-4b02-892e-3f33da3e84d6 · outbound

This paper cites Very Deep Convolutional Networks for Large-Scale Image Recognition.

MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos Very Deep Convolutional Networks for Large-Scale Image Recognition

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.467237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.467237Z digest=sha256:444f72a6544d27112b939ffc1f09a6336874539d7f3600b3397127f7c5d0931f

Observation 70c46022-23c8-4b93-8fac-770eac918e80 · outbound

This paper cites InProceed- ings of the IEEE conference on computer vision and pattern recognition, pages 1059–1067.

MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos InProceed- ings of the IEEE conference on computer vision and pattern recognition, pages 1059–1067

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:31.211184Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:48:30.581939Z digest=sha256:381271a58906ceec1cddf9923620a73e5eefe9cff1fbfde462ba35bd6830035d

Observation 7607221a-21b7-4272-9467-33ef23071a16 · outbound

This paper cites InComputer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part VII 13, pages 505–520.

MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos InComputer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part VII 13, pages 505–520

Reference 2014

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:31.904581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:48:29.588655Z digest=sha256:ff157ff9a89b0a6026e379c4c75e98c42cea9328991fea7a90f9bdb52c9279df

Observation 4a6f7f72-e090-462a-9b56-96a4427c9d89 · outbound

This paper cites Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond.

MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond

Reference 2016

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:30.130052Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:30.130052Z digest=sha256:e6c9d8a554e14cbf4f8457596110fe16ac42f526fafe521faf580f453728e0ed

Observation 264c2662-35b9-4326-84c6-0d60e4ab4ab3 · outbound

This paper cites Deep Communicating Agents for Abstractive Summarization.

MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos Deep Communicating Agents for Abstractive Summarization

Reference 2018

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:29.217734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:29.217734Z digest=sha256:ba9400992b88f1ee069c056e6e0f5dc9a8f328e971a76f790f4bfb8e04d7b280

Observation 63b842e4-0c12-47d4-98b7-1c8586ae3834 · outbound

This paper cites InPro- ceedings of the 2019 CHI Conference on Human Factors in Computing Systems, pages 1–10.

MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos InPro- ceedings of the 2019 CHI Conference on Human Factors in Computing Systems, pages 1–10

Reference 2019

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:31.740163Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:48:29.687369Z digest=sha256:8f36c2dc5b020a1b5b9d51259f0ea491e8c6b987aac0ad93d271ef9c256cc30e

Observation 8063f4d8-c1ec-4ac4-89b5-c4169252c7e1 · outbound

This paper cites Multi-modal Summarization for Video-containing Documents.

MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos Multi-modal Summarization for Video-containing Documents

Reference 2020

Resolution
metadata mismatch
local_arxiv, observed 2026-08-07T00:48:31.016644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:48:29.447866Z digest=sha256:36cf08b63ebfb4d757ae4610468cdbe9959cf03a8121b48455be64c19eea0268

Observation 76ee0f58-4157-4dd3-b9f3-e554262b00db · outbound

This paper cites bert2BERT: Towards Reusable Pretrained Language Models.

MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos bert2BERT: Towards Reusable Pretrained Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-07T00:48:29.332019Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T00:48:29.332019Z digest=sha256:ff130af93dd9f2c297771a8a2742aaedaa85fd961362d4ee6f04debaa1f4341d

Observation a5eb272a-6f4d-4fad-a5ed-e63d7755510e · outbound

This paper cites InFindings of the Association for Computa- tional Linguistics: EACL 2023, pages 880–894.

MS4UI: A Dataset for Multi-modal Summarization of User Interface Instructional Videos InFindings of the Association for Computa- tional Linguistics: EACL 2023, pages 880–894

Reference 2023

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T00:48:31.573166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-07T00:48:29.789251Z digest=sha256:f59a3906f7be18fd3e9a623e793cfe1d6c64e104a54fb6f4cff829bf494cfd86

Pith citing papers

No inbound Pith citation observations are available.