Pith. sign in

Paper Citation Record · LEDGER

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations

As of 13 August 2026, this Paper Citation Record lists 36 of 36 outbound references and 1 inbound Pith citation observation for arXiv:2508.03410.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.03410 v1

Coverage vector

measured 36 of 36 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:30:58.855512Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T04:31:30.702800Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T04:31:30.785644Z

Reference resolution

36 of 36 outbound references displayed

  • verified exact0
  • verified fuzzy29
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 853abf69-ed9e-4ba1-9e1c-aac611e0a97d · outbound

This paper cites Self- supervised object-centric learning for videos.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Self- supervised object-centric learning for videos

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:05.285548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:56.315227Z digest=sha256:8a15baebb382e5fc1021a05b8114b14279a4178dab03baf70d3bbea4ed1805d6

Observation afa7eafe-1c6b-495e-aa81-aad8e3b9446f · outbound

This paper cites Invariant slot attention: object discovery with slot- centric reference frames.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Invariant slot attention: object discovery with slot- centric reference frames

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:05.128437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:56.385986Z digest=sha256:b7b09c9035f0bb728f51dc7902e442d8866f8bb6913218510f127527e4dbd148

Observation 788d4df3-a31c-4f4f-98b0-77e2e490960f · outbound

This paper cites MONet: Unsupervised Scene Decomposition and Representation.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations MONet: Unsupervised Scene Decomposition and Representation

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:56.451793Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:56.451793Z digest=sha256:f1a57fc5a764746216db8be6bdd58ad2ceca9d7b9681121d27455b533ad0d1b7

Observation e0b0dcf3-e677-4c6c-bdf3-e74073ec048f · outbound

This paper cites Emerg- ing properties in self-supervised vision transformers.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Emerg- ing properties in self-supervised vision transformers

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:04.939414Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:56.539206Z digest=sha256:c05c2851bcd259009d8ffdeff866aa220c1e4640667309148dd35b4515d9d7aa

Observation a835b49b-c214-476c-becb-a1db42252975 · outbound

This paper cites Sobolev training for neural networks.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Sobolev training for neural networks

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:04.752038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:56.608318Z digest=sha256:4a9a71b52cabcda62ae1bdb476f651e4dd86699c6a9b26d397eed8c0c7ccd2b1

Observation 7ea62f96-40c5-4b5c-a6ae-7891b212fdbe · outbound

This paper cites CTRL-O: Language-Controllable Object-Centric Visual Representation Learning.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations CTRL-O: Language-Controllable Object-Centric Visual Representation Learning

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:56.664516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:56.664516Z digest=sha256:e61c27ec53b5af17e253bb5005789ae6202a381211f816ada98c871c8cb2ae7f

Observation f930be02-3da2-4fcc-92d6-29c250c9037a · outbound

This paper cites SA Vi++: Towards end-to-end object-centric learning from real-world videos.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations SA Vi++: Towards end-to-end object-centric learning from real-world videos

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:56.725716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:56.725716Z digest=sha256:1b4436360162057dccd5d1457795aa5270e4f67655cc45215ca066ee75fa4d0f

Observation 13164b92-a76f-4537-9434-634e112842e6 · outbound

This paper cites Adap- tive slot attention: Object discovery with dynamic slot num- ber.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Adap- tive slot attention: Object discovery with dynamic slot num- ber

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:56.774375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:56.774375Z digest=sha256:7970bd9c36b8c6fa46872bf219af468f0ad39d72566b034b2d9ed30939237710

Observation d236eda9-cf8d-4093-ab34-4d7c9f54fc61 · outbound

This paper cites MoVi: A large multi-purpose human motion and video dataset.PLoS One, 16(6):e0253157, 2021.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations MoVi: A large multi-purpose human motion and video dataset.PLoS One, 16(6):e0253157, 2021

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:04.582904Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:56.839087Z digest=sha256:241125c1dd247ce54253016674a879631a9449d8e96369b86612cd06b9ec53d7

Observation 9bebab07-a7b3-4fab-b6e9-c7711aca7fd3 · outbound

This paper cites Tagger: Deep un- supervised perceptual grouping.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Tagger: Deep un- supervised perceptual grouping

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:04.421492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:56.896320Z digest=sha256:4cbabb2a3e8cc637a15e5ab0a0b755f9204902005a943e1c37277d79f82386a0

Observation 3dbd392c-dc5c-4cfe-87f1-d1009247cf78 · outbound

This paper cites Neural expectation maximization.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Neural expectation maximization

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:04.211712Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:56.954368Z digest=sha256:ab1c3198f819c2981f7509fcbde35ac69d4d5e397f778f011ca88fd2ec1c6eb4

Observation 5d1fec2e-4287-4024-bcb6-510d26a8e419 · outbound

This paper cites Multi-object representation learning with iterative variational inference.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Multi-object representation learning with iterative variational inference

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:04.031103Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:57.010076Z digest=sha256:6a0cff988b5dbe6dfa9c4972c60fed81d3232adf36ef75d8a45d67f0c969c0f1

Observation f8ef2d4b-b7b5-40ef-922d-7bb24ddd1f00 · outbound

This paper cites MiniLLM: Knowledge Distillation of Large Language Mod- els.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations MiniLLM: Knowledge Distillation of Large Language Mod- els

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:03.799045Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:57.073725Z digest=sha256:fc0efd7f33ac38624ca22b1f94af79f649666a0d60f449f7b98e8c710bd23e73

Observation 5782148d-ac6e-4093-9c78-edd1cd0cfef9 · outbound

This paper cites Masked autoencoders are scal- able vision learners.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Masked autoencoders are scal- able vision learners

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:03.587311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:57.122200Z digest=sha256:0d7f188f0e0a833c2b748ea1eef3f2f1945bdfae2474643865174ebded2dee97

Observation 00520d13-c187-4762-9a1a-6a8358dab990 · outbound

This paper cites Distilling the Knowledge in a Neural Network.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Distilling the Knowledge in a Neural Network

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:57.206596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:57.206596Z digest=sha256:6132164312a0468be5f4160ba955f54f244dca7af6a45b62b2cc947ada646f81

Observation bc62d036-bb5f-4dc8-bd36-387ca04eca9e · outbound

This paper cites Multi-level feature distillation of joint teachers trained on distinct image datasets.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Multi-level feature distillation of joint teachers trained on distinct image datasets

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:03.365189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:57.274758Z digest=sha256:848a7f0d62cb70fffd216aff9c11ef2f04bf1036de4cc4edacba09310c4de6eb

Observation d71c245f-57bd-4c87-a2cc-7f44af9ecddb · outbound

This paper cites Improving object- centric learning with query optimization.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Improving object- centric learning with query optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:57.347323Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:57.347323Z digest=sha256:01b2500c9f43b078e3ae6fd83bf66b8830a045bf3927c5d22b1f20aacba3fa52

Observation 00622ac3-50ac-4062-885f-51679a8509d7 · outbound

This paper cites SPOT: Self-Training with Patch-Order Permutation for Object-Centric Learning with Autoregressive Transformers.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations SPOT: Self-Training with Patch-Order Permutation for Object-Centric Learning with Autoregressive Transformers

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:03.157503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:57.389490Z digest=sha256:b511f3cda3dff653e2ef1287633d54ad23cac43efaf03a9714de06e0d700807a

Observation 4b2b0fd7-dc79-405d-a25a-498f0a23ad8a · outbound

This paper cites DIOD: Self-Distillation Meets Ob- ject Discovery.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations DIOD: Self-Distillation Meets Ob- ject Discovery

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:02.866586Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:57.480626Z digest=sha256:7e15bc856ba19b8aebeb19f687fee90848fd6ca9c2bcbff65ffb2fe4aca3ee4c

Observation 62c6b32d-722d-4918-a9fb-e5f478849866 · outbound

This paper cites Elsayed, Aravindh Mahen- dran, Austin Stone, Sara Sabour, Georg Heigold, Rico Jon- schkowski, Alexey Dosovitskiy, and Klaus Greff.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Elsayed, Aravindh Mahen- dran, Austin Stone, Sara Sabour, Georg Heigold, Rico Jon- schkowski, Alexey Dosovitskiy, and Klaus Greff

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:02.637383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:57.547467Z digest=sha256:da919fccf684fb583753f3a98d0aa6514d91ae9a4a909bbced439974e1dc34c3

Observation 8c313c9c-04d6-4ee1-82ae-ddb00b39d19e · outbound

This paper cites Object-centric cross- modal feature distillation for event-based object detection.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Object-centric cross- modal feature distillation for event-based object detection

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:02.437533Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:57.618336Z digest=sha256:0a1747532965e53592eec2554f24a3fb6a410b3abc12ee6b29343f300b4d4ac6

Observation 21f4a567-791d-4453-ab7a-96d1ec00b5b7 · outbound

This paper cites Hashimoto.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Hashimoto

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:02.138735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:57.690861Z digest=sha256:72364072ff2567030f0d25b2cf294e91eff1856b472d3926b874375f87ccfbdc

Observation d0f3c9ce-1e28-420f-822d-f92e0ddfed7a · outbound

This paper cites Microsoft COCO: Common Objects in Context.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Microsoft COCO: Common Objects in Context

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:01.961664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:57.778202Z digest=sha256:67c3a42cf0759b2b1604e153180dde2057b25cbc9d6b4e94b39f4373fe56b76d

Observation 492b0791-372a-44d9-a45a-58deda931865 · outbound

This paper cites Object- centric learning with slot attention.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Object- centric learning with slot attention

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:01.761165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:57.874877Z digest=sha256:86b8ed7ef3ae77762690f2ca06fc925b4dae679aad6b72cfe464b6ddb5417a75

Observation d13e6d06-ee1d-4dfd-866b-a4b6a6ad3d0e · outbound

This paper cites Temporally consistent object-centric learning by contrasting slots.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Temporally consistent object-centric learning by contrasting slots

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:01.576513Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:57.969711Z digest=sha256:bcc04a2fa203c7634edfffb988a072296151491937c32147ba944d14d02084c3

Observation 01100cb7-ef8f-4908-ad72-21d653f1ccaf · outbound

This paper cites DINOv2: Learning robust visual features without supervi- sion.Transactions on Machine Learning Research, 2024.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations DINOv2: Learning robust visual features without supervi- sion.Transactions on Machine Learning Research, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:01.311964Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:58.071804Z digest=sha256:7255c587202d2220bae3cf36f68a911fff7a427f73cf87c65b95f293b7047788

Observation 998a1e18-719b-4f53-b3f6-dbb637f54a42 · outbound

This paper cites Perazzi, J.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Perazzi, J

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:01.016609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:58.150727Z digest=sha256:e7fb89fdee24f7b77905dd48dc3960374cec53a02d3af5fbb18b2739ac78b659

Observation f3dca698-d562-411f-996e-b2d9e1717786 · outbound

This paper cites Torr, and Song Bai.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Torr, and Song Bai

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:00.692495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:58.223978Z digest=sha256:17ca45723257626cfed20314e3d36975c4a5ab721579d76b4f10559723de1aa5

Observation f2d0397d-f3bc-4e86-8f64-26d5d111339e · outbound

This paper cites FitNets: Hints for Thin Deep Nets.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations FitNets: Hints for Thin Deep Nets

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T04:30:58.248095Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T04:30:58.248095Z digest=sha256:e81850c54f0dbaabb056e63745444807c365a645cffdf9d86bf042a0e15a8671

Observation c0adba32-53f7-49ed-83e2-d08b174b2882 · outbound

This paper cites Bridging the gap to real-world object-centric learning.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Bridging the gap to real-world object-centric learning

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:00.372198Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:58.306675Z digest=sha256:a5e354b6e53cddeeb7608074be687030159040e6712ae9d4b7a37dcd761d0d25

Observation 623dff23-c2ad-4960-95f6-3d55c66f6bb5 · outbound

This paper cites Simple unsu- pervised object-centric learning for complex and naturalis- tic videos.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Simple unsu- pervised object-centric learning for complex and naturalis- tic videos

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:31:00.032058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:58.384400Z digest=sha256:8ceb54b116a9611a67b4c5828d5feaefacaf57a94ca6155fae5403c1dec0cbbc

Observation 4109caf7-3d09-4470-9f27-fc1938d7b9df · outbound

This paper cites Self-supervised video object segmentation by motion grouping.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Self-supervised video object segmentation by motion grouping

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:59.789129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:58.454779Z digest=sha256:a81ddaf006930d287cf4412f764371a46bf9c1fbeed9dbfa038ab44035611d02

Observation 60a88745-0076-4b62-9070-000ae8c311de · outbound

This paper cites The 3rd large-scale video object segmentation challenge - video in- stance segmentation track, 2021.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations The 3rd large-scale video object segmentation challenge - video in- stance segmentation track, 2021

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:59.606311Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:58.585604Z digest=sha256:214d6d150108130fb30e26c3fe69b66b23bcd4c1c2ae51c17d7fe7f251bcf4d6

Observation 92a05be0-e7c7-42ad-ac55-0297dd531122 · outbound

This paper cites Object-centric learning for real-world videos by predict- ing temporal feature similarities.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations Object-centric learning for real-world videos by predict- ing temporal feature similarities

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:59.442514Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:58.699996Z digest=sha256:a4fa53538505e355a054bf74ff393017ce2fec365089c84c9735cf2e4ca1ce08

Observation 8855759e-aea4-430e-ab96-00dc24dfd321 · outbound

This paper cites The SLOTMATCHstudent based on DINOv2 is compared with an equivalent architecture without distillation (no KD), as well as its corresponding teacher model.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations The SLOTMATCHstudent based on DINOv2 is compared with an equivalent architecture without distillation (no KD), as well as its corresponding teacher model

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:59.223002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:58.777597Z digest=sha256:c89fbe02cd535903374d287a2695025d634285c61b5efbd05016acb52cf2302c

Observation fe00adde-0429-40ae-8600-88453e16f351 · outbound

This paper cites In preliminary experiments, we found the standard deviation for FG-ARI, and mBO across seeds to be within±0.06 and±0.29 on YTVIS, indicating stable convergence behavior.

VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations In preliminary experiments, we found the standard deviation for FG-ARI, and mBO across seeds to be within±0.06 and±0.29 on YTVIS, indicating stable convergence behavior

Reference 2048

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T04:30:59.064396Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:30:58.855512Z digest=sha256:d6b1ca3c8f6077d06ce0804f5b87c1b4cfd1868ec1fd59b00cd03871570cb44a

Pith citing papers

Observation bb072915-1027-4b07-bcb8-ca1f9d69bb33 · inbound

Cornelis Easton:The Milky Way as a spiral galaxy cites this paper.

Cornelis Easton:The Milky Way as a spiral galaxy VisAug: Facilitating Speech-Rich Web Video Navigation and Engagement with Auto-Generated Visual Augmentations

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-06T04:31:30.867218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T04:31:30.702800Z digest=sha256:dbc161706ed9d08aa8b2fdcfe6b3329741cd5bf83a431c38b859b44b2314296c