Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T04:30:58.855512Z
Paper Citation Record · LEDGER
As of 13 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2508.03410.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-06T04:30:58.855512Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T04:31:30.702800Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-06T04:31:30.785644Z
36 of 36 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 853abf69-ed9e-4ba1-9e1c-aac611e0a97d · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Self- supervised object-centric learning for videos
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation afa7eafe-1c6b-495e-aa81-aad8e3b9446f · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Invariant slot attention: object discovery with slot- centric reference frames
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 788d4df3-a31c-4f4f-98b0-77e2e490960f · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations MONet: Unsupervised Scene Decomposition and Representation
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e0b0dcf3-e677-4c6c-bdf3-e74073ec048f · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Emerg- ing properties in self-supervised vision transformers
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation a835b49b-c214-476c-becb-a1db42252975 · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Sobolev training for neural networks
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 7ea62f96-40c5-4b5c-a6ae-7891b212fdbe · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations CTRL-O: Language-Controllable Object-Centric Visual Representation Learning
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f930be02-3da2-4fcc-92d6-29c250c9037a · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations SA Vi++: Towards end-to-end object-centric learning from real-world videos
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13164b92-a76f-4537-9434-634e112842e6 · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Adap- tive slot attention: Object discovery with dynamic slot num- ber
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d236eda9-cf8d-4093-ab34-4d7c9f54fc61 · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations MoVi: A large multi-purpose human motion and video dataset.PLoS One, 16(6):e0253157, 2021
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 9bebab07-a7b3-4fab-b6e9-c7711aca7fd3 · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Tagger: Deep un- supervised perceptual grouping
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 3dbd392c-dc5c-4cfe-87f1-d1009247cf78 · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Neural expectation maximization
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5d1fec2e-4287-4024-bcb6-510d26a8e419 · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Multi-object representation learning with iterative variational inference
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f8ef2d4b-b7b5-40ef-922d-7bb24ddd1f00 · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations MiniLLM: Knowledge Distillation of Large Language Mod- els
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 5782148d-ac6e-4093-9c78-edd1cd0cfef9 · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Masked autoencoders are scal- able vision learners
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 00520d13-c187-4762-9a1a-6a8358dab990 · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Distilling the Knowledge in a Neural Network
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bc62d036-bb5f-4dc8-bd36-387ca04eca9e · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Multi-level feature distillation of joint teachers trained on distinct image datasets
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d71c245f-57bd-4c87-a2cc-7f44af9ecddb · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Improving object- centric learning with query optimization
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 00622ac3-50ac-4062-885f-51679a8509d7 · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations SPOT: Self-Training with Patch-Order Permutation for Object-Centric Learning with Autoregressive Transformers
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4b2b0fd7-dc79-405d-a25a-498f0a23ad8a · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations DIOD: Self-Distillation Meets Ob- ject Discovery
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 62c6b32d-722d-4918-a9fb-e5f478849866 · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Elsayed, Aravindh Mahen- dran, Austin Stone, Sara Sabour, Georg Heigold, Rico Jon- schkowski, Alexey Dosovitskiy, and Klaus Greff
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8c313c9c-04d6-4ee1-82ae-ddb00b39d19e · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Object-centric cross- modal feature distillation for event-based object detection
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 21f4a567-791d-4453-ab7a-96d1ec00b5b7 · outbound
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d0f3c9ce-1e28-420f-822d-f92e0ddfed7a · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Microsoft COCO: Common Objects in Context
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 492b0791-372a-44d9-a45a-58deda931865 · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Object- centric learning with slot attention
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation d13e6d06-ee1d-4dfd-866b-a4b6a6ad3d0e · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Temporally consistent object-centric learning by contrasting slots
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 01100cb7-ef8f-4908-ad72-21d653f1ccaf · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations DINOv2: Learning robust visual features without supervi- sion.Transactions on Machine Learning Research, 2024
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 998a1e18-719b-4f53-b3f6-dbb637f54a42 · outbound
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f3dca698-d562-411f-996e-b2d9e1717786 · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Torr, and Song Bai
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation f2d0397d-f3bc-4e86-8f64-26d5d111339e · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations FitNets: Hints for Thin Deep Nets
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0adba32-53f7-49ed-83e2-d08b174b2882 · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Bridging the gap to real-world object-centric learning
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 623dff23-c2ad-4960-95f6-3d55c66f6bb5 · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Simple unsu- pervised object-centric learning for complex and naturalis- tic videos
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 4109caf7-3d09-4470-9f27-fc1938d7b9df · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Self-supervised video object segmentation by motion grouping
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 60a88745-0076-4b62-9070-000ae8c311de · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations The 3rd large-scale video object segmentation challenge - video in- stance segmentation track, 2021
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 92a05be0-e7c7-42ad-ac55-0297dd531122 · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Object-centric learning for real-world videos by predict- ing temporal feature similarities
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation 8855759e-aea4-430e-ab96-00dc24dfd321 · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations The SLOTMATCHstudent based on DINOv2 is compared with an equivalent architecture without distillation (no KD), as well as its corresponding teacher model
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation fe00adde-0429-40ae-8600-88453e16f351 · outbound
VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations In preliminary experiments, we found the standard deviation for FG-ARI, and mBO across seeds to be within±0.06 and±0.29 on YTVIS, indicating stable convergence behavior
Reference 2048
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.
Observation bb072915-1027-4b07-bcb8-ca1f9d69bb33 · inbound
Cornelis Easton:The Milky Way as a spiral galaxy VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.