Pith. sign in

Paper Citation Record · LEDGER

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos

As of 14 August 2026, this Paper Citation Record lists 100 of 132 outbound references and 2 inbound Pith citation observations for arXiv:2411.08753.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.08753 v4

Coverage vector

measured 100 of 132 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T21:27:16.884770Z

measured 102 of 102 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-11T04:45:38.441181Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T06:00:58.843605Z

Reference resolution

100 of 132 outbound references displayed

  • verified exact1
  • verified fuzzy27
  • unresolved72
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bb025d11-7abc-45fa-941e-be55fa65712e · outbound

This paper cites Deep Learning using Rectified Linear Units (ReLU).

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Deep Learning using Rectified Linear Units (ReLU)

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.239636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.239636Z digest=sha256:a31f05bee12695be05f7552fa6e4066f1c93c5401f8ddfb43d15fd4d2a94a1c4

Observation bbbd5838-3012-4a00-bed1-68e5f1a12eb2 · outbound

This paper cites McCrae, Kenton Murray, Maria Nadejde, Satoshi Nakamura, Matteo Negri, Ha Nguyen, Jan Niehues, Xing Niu, Atul Kr.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos McCrae, Kenton Murray, Maria Nadejde, Satoshi Nakamura, Matteo Negri, Ha Nguyen, Jan Niehues, Xing Niu, Atul Kr

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.246965Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.246965Z digest=sha256:a39c5afc1e07c15827de0c69145b86dc5de0a88115748db11f347c0495ab7562

Observation 5f7e63c4-d784-42c4-9cd6-80c062bfff39 · outbound

This paper cites A dataset for develop- ing and benchmarking active vision.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos A dataset for develop- ing and benchmarking active vision

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.253155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.253155Z digest=sha256:fdd1343e19b46d7aec72bfa48f011b59a42ad0af54c412d6f28ba1f44637089f

Observation 5f35b769-c0b0-401e-b12b-d158305d7e1d · outbound

This paper cites Automatic editing of footage from multi- ple social cameras.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Automatic editing of footage from multi- ple social cameras

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.258665Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.258665Z digest=sha256:615322f347b6bd3f873984210fff1c0c0008eefc47ce4944d27bda8b20e5c706

Observation 2b809dd6-454e-4425-a264-bae399839d18 · outbound

This paper cites an unresolved cited work.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Unresolved cited work

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.266710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.266710Z digest=sha256:3dc23bc880a2498f671581a08b2839addb9c5d16c988225efe4fa5254c481d23

Observation cc8180e4-325e-4e9e-8930-f4a1fede591c · outbound

This paper cites WeaQA: Weak Supervision via Captions for Visual Question Answering.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos WeaQA: Weak Supervision via Captions for Visual Question Answering

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.273850Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.273850Z digest=sha256:afe28cc2b13a5f508ef3a2160df9acf1aa63896b278534a23b87de79ce61de02

Observation a3137ed3-695d-457c-b997-5a27ba7b661f · outbound

This paper cites METEOR: An auto- matic metric for MT evaluation with improved correlation with human judgments.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos METEOR: An auto- matic metric for MT evaluation with improved correlation with human judgments

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.281097Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.281097Z digest=sha256:6000a918a6c14467317640594fa2889004cccd77230d23e353d18a107b5a497f

Observation 1fd79043-4e34-4a2c-892b-dd4ab3137967 · outbound

This paper cites Is space-time attention all you need for video understanding? In ICML, page 4, 2021.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Is space-time attention all you need for video understanding? In ICML, page 4, 2021

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.287219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.287219Z digest=sha256:56d469bb7206e67dceccea720c348fa71aee62ec55fb7b0c0c077d78797c73d5

Observation 71bc575b-f19d-4b46-8a5c-939063e2b0b4 · outbound

This paper cites High- lightme: Detecting highlights from human-centric videos.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos High- lightme: Detecting highlights from human-centric videos

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.296787Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.296787Z digest=sha256:d2000ed227222af3a2db8224de5ce3949c2b5d03eac33e9fa4fe76af960e37c2

Observation 1970d812-711b-4c45-a19f-73c8b38c4a28 · outbound

This paper cites Extreme rotation estimation using dense correlation volumes.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Extreme rotation estimation using dense correlation volumes

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.304503Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.304503Z digest=sha256:28d6ccbb39a1cdd707c24dde7b1bfed394bdeedf39350885239723074481ef5e

Observation 0a0c9b5d-fb6a-4dba-a770-a30c788473ae · outbound

This paper cites Davis, and Lei Zhang.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Davis, and Lei Zhang

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.312059Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.312059Z digest=sha256:9a14206b63e6d6959d1ddee81acaea9ad6f04cce01f44cedccb7aa58179d59f4

Observation e8215ca4-bb01-45c4-917f-b677ab0198c0 · outbound

This paper cites Enhanced interactive 360° viewing via automatic guidance.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Enhanced interactive 360° viewing via automatic guidance

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.318489Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.318489Z digest=sha256:918883fc68859deb28f35622a2b929ad40ef6c59700b14ea61e15b8ab89a6270

Observation 20ccd794-4833-45b6-95bf-6553e0024242 · outbound

This paper cites Learn- ing sports camera selection from internet videos.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Learn- ing sports camera selection from internet videos

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.324741Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.324741Z digest=sha256:60fc654da19e18bf123d50b7fdcae5cd8adb0a6785a367319640009c4c73f0ad

Observation 9f637a90-1a36-4165-9a78-1202d46acbd4 · outbound

This paper cites Wide- baseline relative camera pose estimation with directional learning.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Wide- baseline relative camera pose estimation with directional learning

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.331782Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.331782Z digest=sha256:82ab2a0425e9365ed350c25e1ef27de5a884f2a87def000ec3fd470ea2d43bb4

Observation 3899a749-9f6c-48a5-bf29-582af09d9ade · outbound

This paper cites Geometry-aware recurrent neural networks for active visual recognition.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Geometry-aware recurrent neural networks for active visual recognition

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.337905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.337905Z digest=sha256:eefcd4978337bf8f1fc92c866c780652f5bfb95704142c40016d26e264c333c1

Observation 128e8bf9-6538-4aff-8651-901cd8535971 · outbound

This paper cites Towards a richer 2d understanding of hands at scale.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Towards a richer 2d understanding of hands at scale

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.351292Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.351292Z digest=sha256:b5653e970d8523f9c4924bece1613f08c56701a185c7993aab7f710eef67a44b

Observation f512be94-a5f9-47ad-9c48-3fe5ebc5064b · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Gonzalez, Ion Stoica, and Eric P

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.357941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.357941Z digest=sha256:8e7d7a9c3dc604b9741d792eada1cd2ffd13c1564137e76d6b9ee98fa5dfa136

Observation 85501816-c5e0-4811-a324-9150f4030214 · outbound

This paper cites Self-view Grounding Given a Narrated 360{\deg} Video.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Self-view Grounding Given a Narrated 360{\deg} Video

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.365365Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.365365Z digest=sha256:48adbcf13f5df5fa495311699338318f9694c8ab874b79c23a29444ac2fff04d

Observation 0b7dded2-0479-4e31-860c-2989caae0cf8 · outbound

This paper cites Video co-summarization: Video summarization by visual co- occurrence.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Video co-summarization: Video summarization by visual co- occurrence

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.374896Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.374896Z digest=sha256:df60d66bd624c28ff19ed7383c0fb4bcd9d6936edc2801ba7505e585a699ef40

Observation 7ab7f5a7-5e64-401d-b3e0-1b9d56e9b9a2 · outbound

This paper cites elochoice.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos elochoice

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.381981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.381981Z digest=sha256:e7b38aeb25a120f6060733c149dab49c1df9c69120b9b9906953df23593c3344

Observation 15cdcf51-af90-48d1-9276-bccd1ddac3ad · outbound

This paper cites Scaling egocentric vision: The epic- kitchens dataset.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Scaling egocentric vision: The epic- kitchens dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.387951Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.387951Z digest=sha256:d96d1c4b5c2b41ebcabb7ddd11f83680419c56a6eb8ea8dd6a6bd99a426bf6a2

Observation 13f61efd-951b-41cd-a55b-661009310c20 · outbound

This paper cites FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.393588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.393588Z digest=sha256:76521dc41834356b57c9df025ac17c864dc42cc3552b75a42f4cc71353fee662

Observation 58020f30-05f1-4628-8f05-3ada0da5fc75 · outbound

This paper cites Flashattention: Fast and memory-efficient exact attention with io-awareness.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Flashattention: Fast and memory-efficient exact attention with io-awareness

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.399483Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.399483Z digest=sha256:8aafe11d2420b330e7908917cbde4788efa95f65368d8e92f8db97c2f392c46d

Observation 9cb4724f-b0dd-49f3-825f-fe2f18162725 · outbound

This paper cites Velastin.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Velastin

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.404733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.404733Z digest=sha256:dcef16b742aa3448ae3c7120dad4c6aa64becb483c99a2808e0f75483356a320

Observation d0be0b1f-e948-4cb3-b0af-4b86378f439e · outbound

This paper cites Virtex: Learning visual representations from textual annotations.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Virtex: Learning visual representations from textual annotations

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.410436Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.410436Z digest=sha256:a5625bc2a7f0c92b8d39e2ce2c3cd54b9e56a5f13f9b5916124d935f154e96c2

Observation acde62f6-cf94-455a-bc65-539a3d3f779d · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.416234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.416234Z digest=sha256:7f4a8ee3d478a08618ae55d12c65b1c15a42d2b1507c8743c93e4568de8d1db6

Observation 786245bb-34cc-4782-b6da-f8e960b5843f · outbound

This paper cites Dense and aligned captions (dac) promote compositional reasoning in vl models.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Dense and aligned captions (dac) promote compositional reasoning in vl models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.423479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.423479Z digest=sha256:dc157079e55aa76dd088cf720fbfb0b23956666aa665135e7ab29d4aef1bc12f

Observation face0574-4d45-4a9a-b547-52f30f5a2cb4 · outbound

This paper cites Multi-view active fine- grained visual recognition.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Multi-view active fine- grained visual recognition

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.429622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.429622Z digest=sha256:2ee4e6532790b6f94ff98b9653c6b1414c92f3cc456b6cfb39d5888287765fc4

Observation ae7829db-5a61-4529-b7b2-53d081487132 · outbound

This paper cites Multi- stream dynamic video summarization.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Multi- stream dynamic video summarization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.436093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.436093Z digest=sha256:e80772f61a60326ae65aa4783af0c4e2ccae41da981fbfc2d8cbada5cc2ae94d

Observation 6c41b363-bd5c-48fc-8940-827deb8740d4 · outbound

This paper cites Elson and Mark O.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Elson and Mark O

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.441602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.441602Z digest=sha256:c9515ab8b810a81583000afda97080457d02003f0e269dc735365e00ecb1307e

Observation 20a10721-ce03-4f24-8f5d-68767cb25750 · outbound

This paper cites Foote and D.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Foote and D

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.450299Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.450299Z digest=sha256:36d25ec33349665d98b67505f177592caeb86a8b68cd5348fdadbb9fd08a57a2

Observation 0904ea70-ee80-4a76-9aac-ec3227472dea · outbound

This paper cites Multi-view video summa- rization.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Multi-view video summa- rization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.457120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.457120Z digest=sha256:4d687809a6739622fc70c0bf31334de3dbe56aab4e3469a37da2b9ff5717a0bc

Observation aea03a6d-8bf6-41c2-823e-ae6aaa5262d3 · outbound

This paper cites Gleicher, Rachel M.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Gleicher, Rachel M

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.463492Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.463492Z digest=sha256:9f8c560c8922f41c35aed0478350061b904724559cfa7929b2a1e9fa8fc8d2e4

Observation 69cd9eef-ad25-4bde-be6c-0f56da3ca3fa · outbound

This paper cites PEAVS: Perceptual Evaluation of Audio-Visual Synchrony Grounded in Viewers' Opinion Scores.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos PEAVS: Perceptual Evaluation of Audio-Visual Synchrony Grounded in Viewers' Opinion Scores

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.469572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.469572Z digest=sha256:d8a7b7552a6fba63715d33f69c4f4bd40828fed00300ea4a35198e0cff184544

Observation b399e38c-c960-45a5-8172-81faef191d2f · outbound

This paper cites Diverse sequential subset selection for supervised video summarization.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Diverse sequential subset selection for supervised video summarization

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.479571Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.479571Z digest=sha256:02798f1b12d17e7060cb39977d3f852488ee86c6029d493aab72ecb4d35cf82c

Observation 1fb9a401-425c-4436-beaa-101ae1ca7ace · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Ego4d: Around the world in 3,000 hours of egocentric video

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.485303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.485303Z digest=sha256:9cae3e77b54212e4c3726d3fc44dbb5aaef7e376814af15ea335873e938af462

Observation c0b2dc9a-8d8c-43ed-9738-c519d6a0c022 · outbound

This paper cites Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Ego-Exo4D: Understanding Skilled Human Activity from First- and Third-Person Perspectives

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.490447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.490447Z digest=sha256:3a58ec3dbc6c7aab26dd6433d1e4d60ea7242ccbc9844d5c56ed1c548eaedd1d

Observation 36ada9b3-36b4-4a4b-bdaf-73478b66a597 · outbound

This paper cites Temporal Difference Variational Auto-Encoder.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Temporal Difference Variational Auto-Encoder

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.496223Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.496223Z digest=sha256:271d4d2614ff6aafa4d231564265f6e88d1154928f2c816e9854750ea5e0615f

Observation 82cf2f02-60ce-48a7-b2b7-45c7745d63b4 · outbound

This paper cites From Images to Textual Prompts: Zero-shot VQA with Frozen Large Language Models.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos From Images to Textual Prompts: Zero-shot VQA with Frozen Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.503459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.503459Z digest=sha256:7dddd78746bc962630b4ce314bcc3a92fd5aff48a0e242fde8c6e3c315e7f8ca

Observation bb7f5658-868b-4da4-8003-6fea52ca632f · outbound

This paper cites Using closed captions as supervision for video activity recognition.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Using closed captions as supervision for video activity recognition

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.510211Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.510211Z digest=sha256:73b869693398ebfc38aa558fe1f0ee65c6bfde6560d7d8a8a34cbc56ae9a1fa3

Observation ee3b7321-6456-4c50-a586-d310a09ced60 · outbound

This paper cites Creating summaries from user videos.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Creating summaries from user videos

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.520424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.520424Z digest=sha256:5a29d6b6963656ad897604f3c11db89a022a4177063c480275bb928cd7add32d

Observation adf155bf-20d6-4c5c-b6cc-30895c8ea6ba · outbound

This paper cites Video summarization by learning submodular mixtures of objec- tives.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Video summarization by learning submodular mixtures of objec- tives

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.533989Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.533989Z digest=sha256:905554213c273802d0a596e7d8ef225e51ff3084913d373f61c5f8ba679b93eb

Observation 51c0c39d-eb07-4531-b1e0-32883fb4dd9b · outbound

This paper cites Align and attend: Multimodal summarization with dual contrastive losses.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Align and attend: Multimodal summarization with dual contrastive losses

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.539969Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.539969Z digest=sha256:12bce3e8a40d3333bebb19e27c01f2be660e324aad2e67d4e55cb98e7eeb2e6b

Observation ccfe8115-fc24-4cb6-8957-db3605f10956 · outbound

This paper cites Cohen, and David H.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Cohen, and David H

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.545269Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.545269Z digest=sha256:ed6f87db288a7366f5cace0c9f4ab32fa0546cbdc212dcbb7e6da5b67401d55b

Observation 4771a9c2-2ac3-4484-be01-c8bc7857d19c · outbound

This paper cites Cohen, and David H.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Cohen, and David H

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.551100Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.551100Z digest=sha256:d73a2a5cd42cdd73e1005f212e24704cc5b4996d51e3e1a40cda9925297f8c3b

Observation 772d4bde-00fa-4203-a507-cb95825b6041 · outbound

This paper cites Vir- tual videography.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Vir- tual videography

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.556351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.556351Z digest=sha256:bd4f3f7168d26a81bd9e146087546060b8d1179a5fa1e4c66d9b4ec167941ed5

Observation 00f95f17-6312-4f34-91fb-bd507296c3bd · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos LoRA: Low-Rank Adaptation of Large Language Models

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.561315Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.561315Z digest=sha256:74d5e58cbc8cc7c8be38710e5fcc8a8b25de6643529734100f17c2e3c48422d3

Observation d38383b6-f858-4707-a017-1dede39c461d · outbound

This paper cites Deep 360 pilot: Learning a deep agent for piloting through 360deg sports videos.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Deep 360 pilot: Learning a deep agent for piloting through 360deg sports videos

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.566908Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.566908Z digest=sha256:c726bbb7d801e7371fc4e2c4f2be05f0033ff5225e53e54daaa77036b135a013

Observation eaaccbea-74bd-4aa4-a9a1-0ed54ce64a55 · outbound

This paper cites EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos EgoExoLearn: A Dataset for Bridging Asynchronous Ego- and Exo-centric View of Procedural Activities in Real World

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.572054Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.572054Z digest=sha256:90911cf8861c10a59ee58a06d5838bdf4f8cc8d8a36e650055f4ed92ec461037

Observation d3df167d-f2ce-44a1-97d9-fcefb64f1699 · outbound

This paper cites Batch normalization: accelerating deep network training by reducing internal co- variate shift.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Batch normalization: accelerating deep network training by reducing internal co- variate shift

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.577856Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.577856Z digest=sha256:bc53caa9fd7d408c1449ffa306513b90b3b2a5f4f8d163a445e72f682a2a5002

Observation c0397a64-a4a6-40dc-9e1d-f306b0a6255a · outbound

This paper cites Look-ahead be- fore you leap: end-to-end active recognition by forecasting the effect of motion.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Look-ahead be- fore you leap: end-to-end active recognition by forecasting the effect of motion

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.584139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.584139Z digest=sha256:b7b64dca892655d7aa1a98d6c3bd887722494039e62129aa731d3963cf3002e5

Observation 8ea4e95b-7769-44eb-b8c4-b79cb4233be3 · outbound

This paper cites Learning to look around: Intelligently exploring unseen environments for unknown tasks.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Learning to look around: Intelligently exploring unseen environments for unknown tasks

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.589659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.589659Z digest=sha256:8504278a8a104644ef65ecb79dd93e3a1ea3ece65e5aa7f95ead64f7e6265698

Observation 512bb144-e0c5-4bf9-a5f8-977456cfdf53 · outbound

This paper cites End-to-end policy learning for active visual categorization.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos End-to-end policy learning for active visual categorization

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.595297Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.595297Z digest=sha256:07be7d7a06e962da6ca13c99eecd676ec6adc84e76e5b9a8032b7a2d4c5935fa

Observation 2835bd81-9f36-41e4-a29d-cbf613bc3efc · outbound

This paper cites Time-Agnostic Prediction: Predicting Predictable Video Frames.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Time-Agnostic Prediction: Predicting Predictable Video Frames

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.600840Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.600840Z digest=sha256:f7795c6f808e9d36968829947014e5119d38ed9af9918d8bdc0016e6bec666d0

Observation 5c8021b6-ddd4-4df0-a69b-51c0e2d0c7ce · outbound

This paper cites Simglim: Simplifying glimpse based active visual reconstruction.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Simglim: Simplifying glimpse based active visual reconstruction

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.607025Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.607025Z digest=sha256:0051c590026b5e51959d174b9dc117a99d917c2e4368fecd5ef95d612f8f9953

Observation d3adfba1-3679-4606-90ce-c0a8591e3d1d · outbound

This paper cites Lemma: A multi-view dataset for le arning m ulti-agent m ulti-task a ctivities.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Lemma: A multi-view dataset for le arning m ulti-agent m ulti-task a ctivities

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.612958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.612958Z digest=sha256:05d95c861c26e462942c380b5b6e817f036e97d6915c6185e50548f29a6515f1

Observation 75304319-c3b5-4630-854f-269cd0168e43 · outbound

This paper cites RTMPose: Real-Time Multi-Person Pose Estimation based on MMPose.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos RTMPose: Real-Time Multi-Person Pose Estimation based on MMPose

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.619160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.619160Z digest=sha256:7a8ebb43dce8de99fedd2013e4fa9b6365c571f252025f51cfe1691c8eecc854

Observation cece5df7-2619-49f6-be83-94a62c0fa7f0 · outbound

This paper cites Large-scale video summarization using web-image priors.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Large-scale video summarization using web-image priors

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.624405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.624405Z digest=sha256:ff26355aec68f82dc619403028605d3558224dd77c136349f7d07d9aa568302d

Observation 4afe155d-955c-48ed-bbcc-329d56768cd9 · outbound

This paper cites an unresolved cited work.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Unresolved cited work

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.629375Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.629375Z digest=sha256:7baa6044692ee721a162c85a599f06c75198c5ee6813d4501527e52199fd563d

Observation cb661ab1-ee16-4e98-81ab-c87070f9aa7b · outbound

This paper cites Segment any- thing.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Segment any- thing

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.635697Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.635697Z digest=sha256:53fa642c89c3c5674f9c1277c019f60ee0835709a1ed1f6336946c61c0335d98

Observation 883c89c2-d078-4ec9-9dee-f367fe283eba · outbound

This paper cites Hyperbolic Learning with Synthetic Captions for Open-World Detection.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Hyperbolic Learning with Synthetic Captions for Open-World Detection

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-08-12T21:27:17.290381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.641426Z digest=sha256:c7b30b99c67d924e2651b4a7a8f62dffc0bdc29194ebb875c4d00452a7c8ccfd

Observation 41dfb3fc-a295-4cc4-895d-81fc3c3afa56 · outbound

This paper cites A memory network approach for story-based temporal summarization of 360° videos.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos A memory network approach for story-based temporal summarization of 360° videos

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.650230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.650230Z digest=sha256:48f259d93f4e412f97dcd4a5af2a60bfe5b60eddbfd82f9ca15baac781164764

Observation f6e45f0f-680a-4afb-9c6e-ea8b2a8a707a · outbound

This paper cites Predicting important objects for egocentric video summarization.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Predicting important objects for egocentric video summarization

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.657826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.657826Z digest=sha256:ec9a17b55d3e5d86699e9af201d30601bbe64d826079e04859db6af2e7266b9a

Observation ebfefeba-e6cc-4e98-ad04-5e03f150fd7a · outbound

This paper cites MVBench: A Comprehensive Multi-modal Video Understanding Benchmark.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos MVBench: A Comprehensive Multi-modal Video Understanding Benchmark

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.663653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.663653Z digest=sha256:d93b83cd62c4770ecbee17d48f997e971327060d59c2fd37ae1e270b9be4580a

Observation e15fd5d7-482f-47e8-9912-453c1c504e1c · outbound

This paper cites Grounded language-image pre-training.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Grounded language-image pre-training

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.669888Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.669888Z digest=sha256:7527b1c5b0f7f78082f4b01263d17ddb172342e1e227ab30443b95dada499062

Observation cb28cbc6-1078-4f06-883f-5b54aacfb4cb · outbound

This paper cites How local is the local diversity? reinforcing sequen- tial determinantal point processes with dynamic ground sets for supervised video summarization.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos How local is the local diversity? reinforcing sequen- tial determinantal point processes with dynamic ground sets for supervised video summarization

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.675226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.675226Z digest=sha256:f77b82dcaf7e46bf52f75f6c5eb1ccac8e5cd2d524aa34ad4834a1c496bd1879

Observation 756366c4-920a-4fa0-adff-8122bda308df · outbound

This paper cites Egocentric video-language pretraining.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Egocentric video-language pretraining

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.569353Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.680256Z digest=sha256:8411ebbbf8f34a6eafebe25b23ba6471209b2116e903cc67ce5d09d5d9243945

Observation cadd1a59-85f1-4e16-aa5a-d1d422355b18 · outbound

This paper cites Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Grounding DINO: Marrying DINO with Grounded Pre-Training for Open-Set Object Detection

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.685434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.685434Z digest=sha256:55f164b7af4e7e1422b90fa5f61a5b9fb3737dc2a47729db7cd2994b2fb6f22a

Observation d116433c-7261-411f-861d-40dd7605c0ca · outbound

This paper cites SGDR: Stochastic Gradient Descent with Warm Restarts.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos SGDR: Stochastic Gradient Descent with Warm Restarts

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.691686Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.691686Z digest=sha256:55bb336fe5023c0238194c0c5f717ebdcbdd3cfc633c2c92d1daf71c40c73010

Observation 30c3a2d4-49dd-4644-9a7e-f73abf1a3b68 · outbound

This paper cites Decoupled Weight Decay Regularization.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Decoupled Weight Decay Regularization

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.697244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.697244Z digest=sha256:6a8da8db9840e111a1e22b1902b46106109dd256a3c35fae9c879e84723e990b

Observation 5fefd853-04f0-49ae-bd8c-3265e7192f38 · outbound

This paper cites Story-driven summariza- tion for egocentric video.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Story-driven summariza- tion for egocentric video

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.549057Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.702606Z digest=sha256:7ddc721e89c30e7df4a21c5bc432fa4965cd79232843bfd4cad5a885392859ab

Observation a9ae5bc4-44f2-43e1-b13f-e77415b31e02 · outbound

This paper cites Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.709927Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.709927Z digest=sha256:af96034527d2d07d251af80c564fe0b03074f596b40f1ccec61982ee2757677e

Observation 04784fba-8090-4375-bf5d-e45497f5a31e · outbound

This paper cites Video summarization via multi- view representative selection.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Video summarization via multi- view representative selection

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.525516Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.715943Z digest=sha256:bd1c440659c03323fa0e2709ab42ae701439b8ee8bb8df48fbc7a33f9b28d227

Observation fe0807fc-36ba-4265-b1fc-4fa9b5e49f6a · outbound

This paper cites Howto100m: Learning a text-video embedding by watching hundred million narrated video clips.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Howto100m: Learning a text-video embedding by watching hundred million narrated video clips

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.496074Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.721421Z digest=sha256:96769451e03ff2c864f0e1640585c237ef8dec7eb86e8bf031944abba8999d40

Observation 8106cd4f-a3de-45d2-9b5d-29a59d8cef9b · outbound

This paper cites Srinivasan, Matthew Tancik, Jonathan T.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Srinivasan, Matthew Tancik, Jonathan T

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.726553Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.726553Z digest=sha256:cb2fcc3cabc33421c219844e5a388cf3bedf1bd89af0e37d83fd358fc2a51d9d

Observation 03ae2a1e-6682-4d9f-a1ec-1b14efc1a1e2 · outbound

This paper cites Automatized summarization of multi- player games.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Automatized summarization of multi- player games

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.454733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.731426Z digest=sha256:a4643a4b9587452262db1e6ab332e1d9eb451e01fc18e7cec51d4a1a74903acf

Observation 61da99fb-d31c-4453-92b7-c0c58c127885 · outbound

This paper cites Egoenv: Human- centric environment representations from egocentric video.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Egoenv: Human- centric environment representations from egocentric video

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.410732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.745576Z digest=sha256:589cc85a3f77e1eef5062abc6dde4d38e5dd31ec1a36125aa47578fe42bf7bce

Observation 67c3248e-d510-4503-9440-8e7f5cc3a3bb · outbound

This paper cites Tl; dw? summarizing instructional videos with task relevance and cross-modal saliency.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Tl; dw? summarizing instructional videos with task relevance and cross-modal saliency

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.385894Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.750940Z digest=sha256:41daa6a99b6fd73dba4995e5201cfddd4fc1cdee6e631cce7a77f992a0f9ba85

Observation 82e5f804-aefd-4974-b8dc-ef24f38a6c01 · outbound

This paper cites Adaptive skip intervals: Temporal abstraction for recurrent dynamical models.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Adaptive skip intervals: Temporal abstraction for recurrent dynamical models

Reference 79

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.357873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.756958Z digest=sha256:0f1c1d6b0af804784953ae7cd5e907ff2553fc44e897871c786d0724e7d08ae6

Observation 3e63b266-ec5b-4f6c-abe9-3c919f19deda · outbound

This paper cites Au- tomatic video summarization by graph modeling.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Au- tomatic video summarization by graph modeling

Reference 80

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.336833Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.763234Z digest=sha256:2c52f4812aed3c15f64a414105327325ff0576ec46c4b867732337ce5e174152

Observation ad9bab32-aae4-4415-8174-2a602a225541 · outbound

This paper cites Collabora- tive summarization of topic-related videos.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Collabora- tive summarization of topic-related videos

Reference 81

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.316696Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.771983Z digest=sha256:83580efb4bfbc7019f9707885e40ef5da39ac2ea014f4864e678555e47af06c7

Observation 1deecac3-7788-4b32-ab77-5d47a72a0932 · outbound

This paper cites Multi-view surveillance video summarization via joint embedding and sparse optimization.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Multi-view surveillance video summarization via joint embedding and sparse optimization

Reference 82

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.295299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.778053Z digest=sha256:01625b168fea1270a3119bcadebd3ae2bfdd7fbdd51006b469e612b4e6067725

Observation 8f1a830f-752c-4357-a063-ac2da8b6188e · outbound

This paper cites Roy-Chowdhury.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Roy-Chowdhury

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.274325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.783830Z digest=sha256:ebe2a1e47a45e2030844a1ee20fff4acee27dbda128cbf885d139104ffbc53b3

Observation c21a3630-06c2-4b52-af5b-4584cb1ac0f1 · outbound

This paper cites Bleu: a method for automatic evaluation of machine translation.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Bleu: a method for automatic evaluation of machine translation

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.257179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.789486Z digest=sha256:5adddbb364e37b0d34557eb5efee7fbf2f3cc12f181eba8096e169b76a5bf05b

Observation 5a1a688b-0e61-4405-8da4-40835fb472d8 · outbound

This paper cites Sumgraph: Video summarization via recursive graph modeling.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Sumgraph: Video summarization via recursive graph modeling

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.238943Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.795037Z digest=sha256:d4f8ea4b2f438009318d8970e360794bf50ee31ecfb2516e4f6ed4525dd339c7

Observation 7d7ac8df-72f4-4c62-9db1-cc869eb1d79d · outbound

This paper cites Egovlpv2: Egocentric video-language pre-training with fusion in the backbone.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Egovlpv2: Egocentric video-language pre-training with fusion in the backbone

Reference 86

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.215657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.799995Z digest=sha256:a85d91932d5ebff6c2153952c1482c453fffb70eeb772a6bb129d012f4034d16

Observation 378c0f34-b5b3-4334-b4f1-51e9efeac2e5 · outbound

This paper cites Vloc- net++: Deep multitask learning for semantic visual localiza- tion and odometry.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Vloc- net++: Deep multitask learning for semantic visual localiza- tion and odometry

Reference 87

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.193781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.805015Z digest=sha256:6e609ef92015521623c85a8cb7be4ec30aae0c5a1a690574d5e235068a9a5174

Observation a4a08e10-1781-4d09-a398-42aed2923331 · outbound

This paper cites Sidekick policy learning for active visual exploration.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Sidekick policy learning for active visual exploration

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.176412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.810155Z digest=sha256:4fb69b39b501fc68abcd4759533ba539083903edde563429822b5e4a10a7fb0f

Observation af605da0-9852-40f0-bac3-870705489a30 · outbound

This paper cites Emergence of exploratory look-around behaviors through active observation completion.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Emergence of exploratory look-around behaviors through active observation completion

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.160511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.814977Z digest=sha256:bb8ad3ec207b7d06c5383cec394762c6caf6b5f05c21d5e2f9134350e9752fde

Observation e496eb70-24af-4f73-94f6-c2f539eb4e2a · outbound

This paper cites Naq: Leveraging narrations as queries to super- vise episodic memory.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Naq: Leveraging narrations as queries to super- vise episodic memory

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.143654Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.820501Z digest=sha256:aaacf7e78c163083c690640c0180e291f9eab3c9c48c11a930af61f523eb836b

Observation d19f8cfa-e3cc-4253-ad30-7ea6ae3fa489 · outbound

This paper cites Video summarization by learning from unpaired data.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Video summarization by learning from unpaired data

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.126653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.829184Z digest=sha256:1517266f937cb3181208be2d1f025d5209e7b10b4fd4bc147d9a59da05904edb

Observation 0d4725b1-c25f-457f-9870-5f756309db84 · outbound

This paper cites Adaptive video highlight detection by learning from user history.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Adaptive video highlight detection by learning from user history

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.835818Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.835818Z digest=sha256:ee3737d8e45b7cb64dad6fb174f8c2d8c4fcaf93ce904707bf32fe765626f8e4

Observation 2fc92f75-2740-4f13-a654-84a748667606 · outbound

This paper cites Chowdhury.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Chowdhury

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.097889Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.841151Z digest=sha256:76723d97c6cdd3063832624483315ea7efb1eb9636cb652d2ca6dfdc39603165

Observation c5a9690f-59df-487e-b0b1-bb2c99d7e9e3 · outbound

This paper cites Attend and segment: Attention guided active semantic segmentation.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Attend and segment: Attention guided active semantic segmentation

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-12T21:27:16.846399Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T21:27:16.846399Z digest=sha256:2dd293c1af9921a1d334efe724ecad0867dfa5876d515bb4ff7cd315c20ae74c

Observation 121c8f1a-836b-4abb-a24b-3acd089351ec · outbound

This paper cites Glimpse- attend-and-explore: Self-attention for active visual explo- ration.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Glimpse- attend-and-explore: Self-attention for active visual explo- ration

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.070161Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.851964Z digest=sha256:ba28ded2243e17287c248ec32b5f14b97c6713af9f2060cad0e0eac31c77ca0a

Observation d6b9d89a-40b3-4049-8dfd-16e538200d9a · outbound

This paper cites Actor and observer: Joint modeling of first and third-person videos.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Actor and observer: Joint modeling of first and third-person videos

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.052780Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.858613Z digest=sha256:b49dca971be4e494a6c47534cc989eaa5f12fb8faa633d7f93a59997cd8ee629

Observation 8481d850-ddf3-403a-8c8c-56f49059cce9 · outbound

This paper cites Tvsum: Summarizing web videos using titles.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Tvsum: Summarizing web videos using titles

Reference 97

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.034387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.863934Z digest=sha256:d5f15e2021222ec68ee6e6a33f9046b4dbd1c901ee15059ab561d387e3d4678a

Observation 6266f243-3dcb-4190-a5e2-0dda93bef23a · outbound

This paper cites Making 360 ° video watchable in 2d: Learning videography for click free view- ing.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Making 360 ° video watchable in 2d: Learning videography for click free view- ing

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:18.015861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.872398Z digest=sha256:59d55270150a278b8bf8ddd690253fd98536228f3d6e48decb5b4bf1c786b882

Observation 67a9f5c6-e0a6-41cf-9c32-e40ba2a166ae · outbound

This paper cites Pano2vid: Automatic cinematography for watching 360 videos.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Pano2vid: Automatic cinematography for watching 360 videos

Reference 99

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:17.994590Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.878297Z digest=sha256:c5bc4c27e7ac23b350727228f552ba384fc060198bb231a2461757973f8f6734

Observation fbe32558-5b46-4a57-9a06-394c4ba8a462 · outbound

This paper cites Automatic con- cept discovery from parallel text and visual corpora.

Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos Automatic con- cept discovery from parallel text and visual corpora

Reference 100

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T21:27:17.975142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-12T21:27:16.884770Z digest=sha256:ee2ce0e363806d0b47dce2caf6d64de98413a4ae57c719367063ed6307272565

Pith citing papers

Observation 6cd707ae-c6d6-4b32-b3a1-eec071956336 · inbound

Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos cites this paper.

Switch-a-View: View Selection Learned from Unlabeled In-the-wild Videos Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-11T04:45:38.441181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T04:45:38.441181Z digest=sha256:3a3ba31ed329190802c3467bbd525315a8ac7e06403adfe6ceb8c0a1cbd17c47

Observation ebe94ece-1e15-49a0-bcc6-e24f2cfb1f29 · inbound

Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision cites this paper.

Bridging Perspectives: A Survey on Cross-view Collaborative Intelligence with Egocentric-Exocentric Vision Which Viewpoint Shows it Best? Language for Weakly Supervising View Selection in Multi-view Instructional Videos

Reference 187

Resolution
verified exact
local_arxiv, observed 2026-08-07T06:00:58.847019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T06:00:58.555825Z digest=sha256:d8df74fa30d2002c088b131867877b3a25961f0ef952eb1e4a5d78517588020b