Pith. sign in

Paper Citation Record · LEDGER

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks

As of 14 August 2026, this Paper Citation Record lists 100 of 127 outbound references and 3 inbound Pith citation observations for arXiv:2506.07016.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07016 v2

Coverage vector

measured 100 of 127 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:49:53.545730Z

measured 103 of 103 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T22:41:05.001058Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-06-28T19:22:34.453758Z

Reference resolution

100 of 127 outbound references displayed

  • verified exact2
  • verified fuzzy2
  • unresolved96
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation c9c349f6-6edd-46a0-bdcb-665277de79c7 · outbound

This paper cites Crab: A Unified Audio-Visual Scene Understanding Model with Explicit Cooperation.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Crab: A Unified Audio-Visual Scene Understanding Model with Explicit Cooperation

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:49:54.269852Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:49:53.235638Z digest=sha256:ab40e3538089cdcce2f0f336e3df2c88e74fa112aa789a54e4d18fb3e830e6d7

Observation ae26d71d-10cd-4f1b-987c-f9fc48a03f20 · outbound

This paper cites Meerkat: Audio-visual large language model for grounding in space and time.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Meerkat: Audio-visual large language model for grounding in space and time

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.239537Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.239537Z digest=sha256:fb0905a5dd91925fde51df8f177212fa7f41f0479349c2dcda47ea00ed61f291

Observation 6fbb984b-5ef9-49ec-abb0-41193774fd76 · outbound

This paper cites VITA: Towards Open-Source Interactive Omni Multimodal LLM.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.242431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.242431Z digest=sha256:2065eced6908b87879faa87e363044cc0f0e73a1ed12157ccdc6f62d54b56d29

Observation 9281e5fc-12a0-44b1-b467-6ee8a1b32bbc · outbound

This paper cites Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Unified-io 2: Scaling autoregressive multimodal models with vision language audio and action

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.245607Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.245607Z digest=sha256:4915b9f9845f9eb42c14ea0f0ded5b3304926cee5eccc0b4da34c52d0ec404a4

Observation cb2ada9f-3ebe-453f-974a-ce660e33b6ce · outbound

This paper cites CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks CREMA: Generalizable and Efficient Video-Language Reasoning via Multimodal Modular Fusion

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.248569Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.248569Z digest=sha256:1d22c9141deb971de2e6a5a4db7ca15e2b0905ebaf72881601da0f11b9b62e11

Observation 07fa958e-1721-4972-a904-1b52d70a42c5 · outbound

This paper cites Avicuna: Audio-visual llm with interleaver and context-boundary alignment for temporal referential dialogue.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Avicuna: Audio-visual llm with interleaver and context-boundary alignment for temporal referential dialogue

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.251727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.251727Z digest=sha256:ec5b592069dbc52d751aaa52da4ec60a2a4e1455aa3d10984e45b8fa5ecaf9fc

Observation 2aa1aa53-cd40-479c-b381-cab4192fa0bd · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks MLVU: Benchmarking Multi-task Long Video Understanding

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.267029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.267029Z digest=sha256:b049e549f17711ebe5e65242f6ce21d6c6bb57b2a228e15f95fc283e1449d993

Observation ffb5b741-b0a6-47a7-9850-a1b6c96facdf · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Sharegpt4video: Improving video understanding and generation with better captions

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.270357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.270357Z digest=sha256:29bb9fd370d6ef276f16d554721e21ae300bf1b42091ba4d2f84b82d9bfbac77

Observation a55a95d4-5a08-4bdb-a540-3c8398416327 · outbound

This paper cites CinePile: A Long Video Question Answering Dataset and Benchmark.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks CinePile: A Long Video Question Answering Dataset and Benchmark

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.273176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.273176Z digest=sha256:2e1b3692676124f37ac89b3e093ba5d03eff2bafea80f09138e84594dd3eb52f

Observation b68cb8c7-9fec-421a-85ad-614890e1b7d6 · outbound

This paper cites Moviechat: From dense token to sparse memory for long video understanding.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Moviechat: From dense token to sparse memory for long video understanding

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.279518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.279518Z digest=sha256:2727cc12fc4487bffd57d908e0cdb4611ba3e2eb162d32e9b03d7d89938ac6c6

Observation 0642fdb3-3293-452d-aace-d241ec4e6a73 · outbound

This paper cites Dense-localizing audio-visual events in untrimmed videos: A large-scale benchmark and baseline.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Dense-localizing audio-visual events in untrimmed videos: A large-scale benchmark and baseline

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.281960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.281960Z digest=sha256:b74027b7b81ee851e00d0dcceead1f799c36de55e70c42a76bda7425ded5e4e4

Observation 28cd0823-f30f-498c-9861-143eb3a8bc9f · outbound

This paper cites Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.284473Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.284473Z digest=sha256:179e62210d1ce14a1b1e2ddbf91c49e590867b4411340f42b60d0b95141fc680

Observation 2f204c4f-25e5-4659-bdd6-a4b9f95b604a · outbound

This paper cites Avqa: A dataset for audio-visual question answering on videos.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Avqa: A dataset for audio-visual question answering on videos

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.287137Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.287137Z digest=sha256:3281cc0f45c7b8ca37bd8ce56b76b83582a2a98e4313c186d13e56390004f642

Observation e8a4da05-4271-4769-9f0e-d9b74fa231d7 · outbound

This paper cites Cat: Enhancing multimodal large language model to answer questions in dynamic audio-visual scenarios.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Cat: Enhancing multimodal large language model to answer questions in dynamic audio-visual scenarios

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.290455Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.290455Z digest=sha256:6dee413374fbcdc206d32224112e43e861e3a5d05c5e74f1cff687ec6b4a1a5a

Observation 86e2d76f-14bd-411b-a289-923b64fd4cb2 · outbound

This paper cites Learning to answer questions in dynamic audio-visual scenarios.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Learning to answer questions in dynamic audio-visual scenarios

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.293237Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.293237Z digest=sha256:de49e854f57b449893eb9847ec425f364ae715cd0e324cbf0fc05765e3581c4e

Observation 2b52dfc9-6683-4c2b-abb7-c81cab78b27a · outbound

This paper cites Vggsound: A large-scale audio- visual dataset.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Vggsound: A large-scale audio- visual dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.296583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.296583Z digest=sha256:3a666f6876e02ccf448807b10c1f749eafc75dd36bc0330318857f837310b5b4

Observation 74a42b6f-062f-4581-8e7b-74b0436b721c · outbound

This paper cites SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks SAVEn-Vid: Synergistic Audio-Visual Integration for Enhanced Understanding in Long Video Context

Reference 22

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:49:54.215338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:49:53.299116Z digest=sha256:2ec956d1c5499362d0016809d7e8fb8e5fcb61f064afcb49da1b4406482c4d0d

Observation 9743a3e4-b48c-45de-a516-1a9cbe36c5de · outbound

This paper cites Imagebind: One embedding space to bind them all.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Imagebind: One embedding space to bind them all

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.302017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.302017Z digest=sha256:444cee764f766fb3a2f2b0430a8c67855ac3b605f6ac2b80c9f1c4c594d3f0e9

Observation fd4a3517-5586-4f10-a753-2a6c64bfd37b · outbound

This paper cites Gemini: Google’s multimodal ai model.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Gemini: Google’s multimodal ai model

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.305234Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.305234Z digest=sha256:fea615b5eae7576afefaa6a3611a02f48b26815428ff2d9435d6372d233a7291

Observation 14d0b676-6f3f-4568-95a6-a122675a9208 · outbound

This paper cites Qwen2.5 Technical Report.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Qwen2.5 Technical Report

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.307666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.307666Z digest=sha256:0522dfd5119d2669feb434017921df394e0abd4057e7e6804b4d819a90f9ab99

Observation d21a88a6-d50d-4291-8c0a-83ba536d916d · outbound

This paper cites GPT-4o System Card.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks GPT-4o System Card

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.310525Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.310525Z digest=sha256:09daa22a867a211cd13441c51e3f5145ee25ece0fc326fb4633ad1c77c374f05

Observation 43ef4468-98cc-4c9e-9be2-577564468c64 · outbound

This paper cites Towards General Text Embeddings with Multi-stage Contrastive Learning.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Towards General Text Embeddings with Multi-stage Contrastive Learning

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.313653Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.313653Z digest=sha256:53bf67860fa771545ea0eab02eab19627d3b4fc241ef56aaad0ae21d757b686b

Observation 0cc2a8d5-55cd-4c78-9013-06881199978f · outbound

This paper cites Variants of the hungarian method for assignment problems.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Variants of the hungarian method for assignment problems

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.316432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.316432Z digest=sha256:e40562b9d1ac66b4f25736ea741b8b50a6a8ec8bcbcee3e0c0acf417408bca5f

Observation c9527294-ac5e-4fa7-bc35-06407f19f055 · outbound

This paper cites Video-rag: Visually-aligned retrieval-augmented long video comprehension.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Video-rag: Visually-aligned retrieval-augmented long video comprehension

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.319165Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.319165Z digest=sha256:cab62226a1f94551c39befdcf9c55a3cc4225f95595933395d5dc4a52bb539e8

Observation 75b41b82-8f33-443b-a107-50688a16e005 · outbound

This paper cites video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks video-SALMONN: Speech-Enhanced Audio-Visual Large Language Models

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.321734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.321734Z digest=sha256:d3772b2805da9344b08dac6e21262bea77002a02e408d8b086cde7f6d132b5c5

Observation 861d624e-75d6-47de-be27-79ed826e7e49 · outbound

This paper cites High-fidelity audio compression with improved rvqgan.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks High-fidelity audio compression with improved rvqgan

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.324462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.324462Z digest=sha256:cb8d7a002117fd8c366ca2f356f83c1f3cdb1994ed8d2e9878db1426d18fd6b7

Observation 4faf8e1e-e314-4294-9137-d87b8caf369f · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.326881Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.326881Z digest=sha256:095f44e889d432554f9ec2afd754a400067a315b36277ea54bbaa5f828474226

Observation 521efbb0-9174-4f8f-9d04-0aca524bec44 · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.329657Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.329657Z digest=sha256:edea214d49b528ee99126182e878f56d96e1dd42cdd50f6d7521b2bd5ddb76e7

Observation 64c22047-d71b-4f29-9eed-692710e1357a · outbound

This paper cites Mvbench: A comprehensive multi-modal video understanding benchmark.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Mvbench: A comprehensive multi-modal video understanding benchmark

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.332855Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.332855Z digest=sha256:8c07e9e0a92bf217a3af2a85c800963882f2467a05d6cf2c4700d47b9b5d7b01

Observation 1e9f3a76-d7a7-4118-adb0-2a0b0df3377f · outbound

This paper cites Video Question Answering: Datasets, Algorithms and Challenges.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Video Question Answering: Datasets, Algorithms and Challenges

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.335527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.335527Z digest=sha256:c620f1f45f3dc5efd0da6882b7037984ecbcf3f5c39d4476bf58b6c43cd91f4a

Observation d005edee-7e59-453d-a80b-70df346af50f · outbound

This paper cites LVBench: An Extreme Long Video Understanding Benchmark.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks LVBench: An Extreme Long Video Understanding Benchmark

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.338244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.338244Z digest=sha256:2fb3d6c72f8acb404a6bb7a19ab0e39ed1b72745f98954258b7382f432432d36

Observation 44fd02b2-02c6-4094-a565-5d80afc81ff9 · outbound

This paper cites Movieqa: Understanding stories in movies through question-answering.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Movieqa: Understanding stories in movies through question-answering

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.341351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.341351Z digest=sha256:e28de855c559bcddd870a8df8e161588ea4d0b257e2d3a0940333f0a110eb248

Observation 1e25c88d-2c88-4a7a-9074-80b411e855b5 · outbound

This paper cites Are we asking the right questions in movieqa? In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, pages 0–0, 2019.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Are we asking the right questions in movieqa? In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, pages 0–0, 2019

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.344344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.344344Z digest=sha256:4fbd71e8f8b4e4136538aa40fd4170d5549ee2ec6339416bdc21f2a52fda1a07

Observation e13897b5-9982-4b69-b512-dcf1b32f27c5 · outbound

This paper cites Activitynet-qa: A dataset for understanding complex web videos via question answering.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Activitynet-qa: A dataset for understanding complex web videos via question answering

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.347140Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.347140Z digest=sha256:d0818ccbb401f578fa7156c67b5ef017bffb288c75bcfd0705438a9046905ec9

Observation 04042d7a-4529-42c6-add3-d490954fb018 · outbound

This paper cites How2: A Large-scale Dataset for Multimodal Language Understanding.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks How2: A Large-scale Dataset for Multimodal Language Understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.349615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.349615Z digest=sha256:57666f21ec80f5277c002d67c2363da1641033079a94a99b4750b5a868b32414

Observation dd516ceb-d6e4-4528-8a6e-b47a87b04001 · outbound

This paper cites Next-qa: Next phase of question-answering to explaining temporal actions.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Next-qa: Next phase of question-answering to explaining temporal actions

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.352727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.352727Z digest=sha256:b3be42698104908906609cb3635fc387db03a26bb46f25dec04886fe213d69d2

Observation 4de816fd-bc2a-4bcb-90f1-f11d582fca64 · outbound

This paper cites Perception test: A diagnostic benchmark for multimodal video models.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Perception test: A diagnostic benchmark for multimodal video models

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.354968Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.354968Z digest=sha256:8d3f60daa5c424f88f3695c1d6d831b8a0db74ee27216bcfd92345f6d227a1f6

Observation 28fd6f61-85cf-4000-9df3-5bdea4755f80 · outbound

This paper cites STAR: A Benchmark for Situated Reasoning in Real-World Videos.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks STAR: A Benchmark for Situated Reasoning in Real-World Videos

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.357596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.357596Z digest=sha256:eba5a2673f3cbba553e13eb4f3d3a62d33f5d71b326fefabd88e85c1f60b2048

Observation c14e6391-f72e-47a3-9cbe-5663d75b4092 · outbound

This paper cites Agqa: A benchmark for compositional spatio-temporal reasoning.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Agqa: A benchmark for compositional spatio-temporal reasoning

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.360494Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.360494Z digest=sha256:80726613a140a8326c8d5e7b4c3c58eaf853f696485dc253d8700f53b3019d08

Observation dfa4da0a-2546-47a2-9d07-f89a2ae3df00 · outbound

This paper cites Egoschema: A diagnostic benchmark for very long-form video language understanding.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Egoschema: A diagnostic benchmark for very long-form video language understanding

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.363689Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.363689Z digest=sha256:1dce797370d28f480c344e0e4e535fa04beea9b0d1dcfadfc05271c14f6fb308

Observation 4129b2ae-f882-4af7-9895-bbe0823ac0cb · outbound

This paper cites Longvideobench: A benchmark for long-context interleaved video-language understanding.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Longvideobench: A benchmark for long-context interleaved video-language understanding

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.366642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.366642Z digest=sha256:36aa143cc4b30f93c219feb52ec12eeb2636dae02b72d5fa3cd25fc267e81a12

Observation c8204057-19fa-4593-b216-cc2fbaa405a1 · outbound

This paper cites Just ask: Learning to answer questions from millions of narrated videos.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Just ask: Learning to answer questions from millions of narrated videos

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.369194Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.369194Z digest=sha256:fe486f7d93781d89f9627a832facc991a2572350d95853daeba12fdaac5ccfa8

Observation b3a54a4b-e9d3-4d54-a0b1-7cbe0b1e541f · outbound

This paper cites InstructionBench: An Instructional Video Understanding Benchmark.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks InstructionBench: An Instructional Video Understanding Benchmark

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.372532Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.372532Z digest=sha256:f6a5e7965ed74e71c0b5a13295431de1f8c54291ae05621ce3a735d35b196025

Observation cb94633a-c344-46d0-a195-ee1cbbcc352c · outbound

This paper cites HD-EPIC: A Highly-Detailed Egocentric Video Dataset.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks HD-EPIC: A Highly-Detailed Egocentric Video Dataset

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.375278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.375278Z digest=sha256:939feeb359b848dbeeff1af753ed8fbaafe5002d50b94c70ca3c723355eb812d

Observation 7fa9c4d4-f755-44ca-9822-11dfc0246416 · outbound

This paper cites Ego4d: Around the world in 3,000 hours of egocentric video.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Ego4d: Around the world in 3,000 hours of egocentric video

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.378573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.378573Z digest=sha256:a3047fcc73ab9796a622d6ebd76e53205b4cab3182043ba9faaed8ef90e72c09

Observation 812bdbe8-357f-4318-92ae-ef9fe1546578 · outbound

This paper cites Ego-exo4d: Understand- ing skilled human activity from first-and third-person perspectives.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Ego-exo4d: Understand- ing skilled human activity from first-and third-person perspectives

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.381710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.381710Z digest=sha256:3c68481c90c295664971d380165f420de8fda7ed9eaaac14b8b4f5da814aed3e

Observation 4a9a7a8e-fa8e-465d-a7c3-0011f7212f13 · outbound

This paper cites Scaling egocentric vision: The epic-kitchens dataset.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Scaling egocentric vision: The epic-kitchens dataset

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.384201Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.384201Z digest=sha256:cd860864275aade895ccd01aea2c32f8464a8c99d21e68d6796926a0e2862eee

Observation 5c323aad-84cb-4215-9c20-e0af2b7a6b7e · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.387583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.387583Z digest=sha256:0f23b109ac4743c65e8d3946c4205b6719de1d13a749828d89b844b464bcbcf2

Observation f5324068-7c28-4598-a0ac-5e7f23433fe8 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks LLaMA: Open and Efficient Foundation Language Models

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.390941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.390941Z digest=sha256:e9bbb0026edd057e9f3d5b14a43630f080f755970c52399c04b0a56f77566675

Observation 4714920f-c24b-4c1a-a7e2-0299bb14b562 · outbound

This paper cites Mistral 7B.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Mistral 7B

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.393475Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.393475Z digest=sha256:41da280db665326d96ac333733d522fc1aca27a20bdd32f89b05ed61e5a2359c

Observation 20fd0780-27b5-4a90-872a-f5ccf4089a5d · outbound

This paper cites VideoChat: Chat-Centric Video Understanding.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VideoChat: Chat-Centric Video Understanding

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.397026Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.397026Z digest=sha256:016d75ba66467cfa1ef2768d2920e8c51a501b97827eb094a19b713d768f4dda

Observation 78fa4c01-ccd7-4fb7-b01c-789b8ccbdd51 · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.399967Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.399967Z digest=sha256:190e603aaffaed17de0f347452f9dedec9d3b6a60cafd31a9814ebfadfdd21e3

Observation 3e0fbcc7-cc27-4663-a7cd-b84c9c0db2f4 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks LLaVA-OneVision: Easy Visual Task Transfer

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.403166Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.403166Z digest=sha256:bb8000384c9ba221fad1dfdcd9a373e1116782edc84d1349b77e1b5ad534712c

Observation 01eb85a9-d3e8-4f7e-9a2d-ffdb046686dc · outbound

This paper cites World Model on Million-Length Video And Language With Blockwise RingAttention.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks World Model on Million-Length Video And Language With Blockwise RingAttention

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.405853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.405853Z digest=sha256:7432f60db7b461b564432b0f52b8b201937d84a8ebb41c6edd8eda8e4175b8fc

Observation 372c0c6c-ddb6-46d8-8879-8a2903415c3f · outbound

This paper cites Long Context Transfer from Language to Vision.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Long Context Transfer from Language to Vision

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.408552Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.408552Z digest=sha256:b8935b1b76f047d03b655e2b73733ec2c7d1481ec95f494280bc8ce1cc055c0c

Observation ea34d88d-cd06-4fee-838a-1704492c1e8c · outbound

This paper cites LongVILA: Scaling Long-Context Visual Language Models for Long Videos.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks LongVILA: Scaling Long-Context Visual Language Models for Long Videos

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.411481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.411481Z digest=sha256:6c402d231b2d9d3247263039ab818f5a5b6f29cb3ec6c14564f66bdff852c6d6

Observation 785eeebb-0c74-458b-91e0-3df35fdf5fc7 · outbound

This paper cites Longllava: Scaling multi-modal llms to 1000 images efficiently via hybrid architecture.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Longllava: Scaling multi-modal llms to 1000 images efficiently via hybrid architecture

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.414273Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.414273Z digest=sha256:e775e8c0d2801a642304eadc663bda62323abbcf4e35c6fe680c4f616756dbe8

Observation dc8fecce-635a-4350-823e-a604ad8268a5 · outbound

This paper cites LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks LLaMA-VID: An Image is Worth 2 Tokens in Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.416913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.416913Z digest=sha256:a2bf6e28440d4042dfcf405571ecc26a251071a025dcf6e74ee10e2550dda7e3

Observation 5c7ac51c-e03f-4c20-be75-737f1f5e4b4e · outbound

This paper cites Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.419480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.419480Z digest=sha256:0f913cb379a07bc542c126b1fd209bc5974881ec6a9a5cac53a215c3dc65680c

Observation 1eca1a7e-e2ae-4794-9fb8-a2cf4f1428ac · outbound

This paper cites VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VideoAgent: A Memory-augmented Multimodal Agent for Video Understanding

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.422447Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.422447Z digest=sha256:c12df9eed2addee930db4ff0b5f1633fb9eeb4a81c1f5eebf4ba0add914a1bc3

Observation c0c07cc0-f991-4305-80a8-37b2a2c996c5 · outbound

This paper cites VideoAgent: Long-form Video Understanding with Large Language Model as Agent.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VideoAgent: Long-form Video Understanding with Large Language Model as Agent

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.425309Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.425309Z digest=sha256:85aa3ac2dcde42dc888295c73b6808d1c5acb8201031a1a4a220e4ce2b0263e6

Observation 28f45b09-2bf9-437a-841d-7b2e7489b3af · outbound

This paper cites VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VideoTree: Adaptive Tree-based Video Representation for LLM Reasoning on Long Videos

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.428376Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.428376Z digest=sha256:efaa5a1cd2f0ed4151e23662f5682953028b0d8a26b89056d8cbcaf2347630d8

Observation 307dc06e-81df-428d-b77f-398c091a9c6c · outbound

This paper cites Retrieval-augmented generation for knowledge-intensive nlp tasks.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Retrieval-augmented generation for knowledge-intensive nlp tasks

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.431459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.431459Z digest=sha256:76344e6067a5a83909000049bc96c667372fa01ba6dce537cdda21f1c833b40d

Observation b84ea8fc-79e5-4af5-89b9-ee20d9c47a53 · outbound

This paper cites Active retrieval augmented generation.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Active retrieval augmented generation

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.434530Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.434530Z digest=sha256:ff95fcf8616c4f2cfce552ae1c286e63dbe970ac744c1c4f435babd8c7043388

Observation 8afc8e91-0f1c-47c2-8a98-f9cd1ff1daed · outbound

This paper cites Sentence-level prompts benefit composed image retrieval.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Sentence-level prompts benefit composed image retrieval

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.437517Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.437517Z digest=sha256:b6a4fd05cbb10a4b06127e6a10554fa87406c219f63ab18fea92922175ce3cdd

Observation 7e026fbc-b4f4-4fb3-806c-4d987713efd9 · outbound

This paper cites VQA4CIR: Boosting Composed Image Retrieval with Visual Question Answering.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VQA4CIR: Boosting Composed Image Retrieval with Visual Question Answering

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.440763Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.440763Z digest=sha256:ca222cfe11858469747f3d3914948da2bed52025c5a6bbd9b68d0eb424cf544f

Observation fd6652c5-5c9d-4bb3-9842-a4c75f95a24c · outbound

This paper cites Searching for Best Practices in Retrieval-Augmented Generation.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Searching for Best Practices in Retrieval-Augmented Generation

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.443996Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.443996Z digest=sha256:4fa88c7050433c36729810644aa74ee45c8c4a7dafb423d36907f0e8c3946392

Observation 211330bc-be32-4ab6-8105-170281393d0f · outbound

This paper cites Retrieval-Augmented Generation for Natural Language Processing: A Survey.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Retrieval-Augmented Generation for Natural Language Processing: A Survey

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.447621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.447621Z digest=sha256:9e8ae1c921e241853c2f6585a37161b85d931aff17a9c36bd41fac0b7e81c29f

Observation 3de5b9a1-3313-48b8-a56f-8d2929889cf1 · outbound

This paper cites A Survey on Retrieval-Augmented Text Generation for Large Language Models.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks A Survey on Retrieval-Augmented Text Generation for Large Language Models

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.450836Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.450836Z digest=sha256:b59bc1f5f30fabf4d9e6a444ed1d7f23750b5d79f343aae0120ddfeb9e0f0bb9

Observation cea28d7a-102d-4cb3-856a-122697f58eb1 · outbound

This paper cites Retrieval-Augmented Generation for AI-Generated Content: A Survey.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Retrieval-Augmented Generation for AI-Generated Content: A Survey

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.453946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.453946Z digest=sha256:08006c4ba97a8dd1eb91889a8e32e8fcf50e0a8a5162baa125b0b7f50b992d3e

Observation bf8931c0-9ddf-442f-9271-4b1c7f1279ff · outbound

This paper cites Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Self-RAG: Learning to Retrieve, Generate, and Critique through Self-Reflection

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.457193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.457193Z digest=sha256:7841e03f2ad8432612a6cd595fde12117a9653b61ee082f2c5c8c73a0cce6e99

Observation 82e24110-d08c-4dbe-9695-4b2b63acb26f · outbound

This paper cites Realm: Retrieval- augmented language model pre-training.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Realm: Retrieval- augmented language model pre-training

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.460122Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.460122Z digest=sha256:9f3c74b8394aabb601186e43b00f094e8e844a0ca3883a8b4b0ebab5b8f03ae7

Observation a1e3442f-a6d3-46a9-8103-6820d27ca7e6 · outbound

This paper cites SAIL: Search-Augmented Instruction Learning.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks SAIL: Search-Augmented Instruction Learning

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.463160Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.463160Z digest=sha256:51dafc0e5ebfa40ab2e168f6f65d3069fa448301fd827a40c236d396610e6049

Observation f6995908-3bb9-49a2-8ee2-fe0fc8b4e038 · outbound

This paper cites Dense passage retrieval for open-domain question answering.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Dense passage retrieval for open-domain question answering

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.465906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.465906Z digest=sha256:8d650c47c1344abcb8119dcbc3db28e034b4a6ec33d8e828015015b271a438ca

Observation 9d1c0ee8-f18d-4add-88fc-4f40f10f5aea · outbound

This paper cites Document haystacks: Vision-language reasoning over piles of 1000+ documents.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Document haystacks: Vision-language reasoning over piles of 1000+ documents

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.469041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.469041Z digest=sha256:f84a67f2b2a940603cc167d8918197a9766f8e8d44f39fcb8daa40dfcc299ec4

Observation ecef7520-2cff-45d5-9b6d-d3d4fb5e731a · outbound

This paper cites an unresolved cited work.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Unresolved cited work

Reference 81

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.472251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.472251Z digest=sha256:e8b02c8e489874f32a10126736901dfa51b07047b3cdee262fcc983d42364350

Observation 03f03589-0a7e-44de-a1b4-023e511b63e0 · outbound

This paper cites Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Visual Haystacks: A Vision-Centric Needle-In-A-Haystack Benchmark

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.475228Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.475228Z digest=sha256:18950882b9cdc4d5186688393927fa11b1054a0c1ab28a294b3c1f4ef58fa7ca

Observation b601c34f-c801-408f-b65c-b77ebb3da369 · outbound

This paper cites VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VisRAG: Vision-based Retrieval-augmented Generation on Multi-modality Documents

Reference 83

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.478155Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.478155Z digest=sha256:1fffc55f7f381d88902dd963a2035d45dc19a4d26e223a2fff221041309c26cd

Observation 326e72fd-02cb-4c9e-924a-f48db5577fca · outbound

This paper cites Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Visual-RAG: Benchmarking Text-to-Image Retrieval Augmented Generation for Visual Knowledge Intensive Queries

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.481450Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.481450Z digest=sha256:075ae5890297a5e9a9d85226b786dec450758fe6b742935f9a22ed6a7c2d80d5

Observation ffd5fd2a-22ec-48b0-bdf5-303b0fafa39b · outbound

This paper cites Vdocrag: Retrieval-augmented generation over visually-rich documents.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Vdocrag: Retrieval-augmented generation over visually-rich documents

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.484538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.484538Z digest=sha256:05790d850e5f2a56b8e9b9956ef9c3922199a28429bb6a1d72ad7ca6e22aec17

Observation cc435739-a5a2-4dbf-bfbf-127ac913f76d · outbound

This paper cites Retrieval Augmented Visual Question Answering with Outside Knowledge.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Retrieval Augmented Visual Question Answering with Outside Knowledge

Reference 86

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.487275Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.487275Z digest=sha256:6a8f5147934d002ba7d23db4646552a0b277a41b4ae496270b9026479c0c6750

Observation 13bf5ae5-ae74-47dc-b20a-7252476bba83 · outbound

This paper cites Fact-Aware Multimodal Retrieval Augmentation for Accurate Medical Radiology Report Generation.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Fact-Aware Multimodal Retrieval Augmentation for Accurate Medical Radiology Report Generation

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.490369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.490369Z digest=sha256:90c64e576b5e92c3063888238306ee4c5c4affbb239778a26f74ff47e334f060

Observation 3a3c1727-6ef2-48e4-88fb-1affe423f913 · outbound

This paper cites Rule: Reliable multimodal rag for factuality in medical vision language models.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Rule: Reliable multimodal rag for factuality in medical vision language models

Reference 88

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.493573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.493573Z digest=sha256:d20e4b848e2945cb0b3fc618addd67761feeddc3b8d1c0a4874899ff69c40af5

Observation fe7e62fc-4e71-4d67-8fa1-ece2db426d5e · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 89

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.496853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.496853Z digest=sha256:46d10a7740e55cea654f190c2b0a4d855d3b83d5994ed8cacb9f659b4db0aa3a

Observation 27713dc0-bfae-43f4-8ab4-52514fc9136e · outbound

This paper cites Gpt-4o: Enhanced multimodal language model.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Gpt-4o: Enhanced multimodal language model

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.499783Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.499783Z digest=sha256:df87cfa7c8f838166afd81a1d539ce1f602e8a230eae21b5cbe88f1236e26f2c

Observation 740f4cb2-fa07-4a92-9683-d5d6d5972ec0 · outbound

This paper cites Minigpt-4: Enhancing vision-language understanding with advanced large language models.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Minigpt-4: Enhancing vision-language understanding with advanced large language models

Reference 91

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.502745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.502745Z digest=sha256:a82ec5c23182b539c6e2e01a626cecb8ba735987224815a03d24ac5524cdd763

Observation 8129eeaa-5041-402e-90f5-2c55f99f7835 · outbound

This paper cites Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Unified-IO: A Unified Model for Vision, Language, and Multi-Modal Tasks

Reference 92

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.505625Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.505625Z digest=sha256:66223a60c8de2443749b788276845ce3593bd371a0d82b40ebaf03ebc7eb609a

Observation e9e009aa-e610-432d-a75e-7b9b36171587 · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 93

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.509123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.509123Z digest=sha256:2c1d43fef36cd9db50caaff995fc3843830a8164c62647ae81adc3297a1b0d82

Observation f397db9d-cebe-4215-ac62-ec113453365f · outbound

This paper cites Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Making the V in VQA matter: Elevating the role of image understanding in Visual Question Answering

Reference 94

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.511845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.511845Z digest=sha256:2431ed3bf904ca128a7dc262379f3e7394bae8e1cccc2fc24e722b929f507678

Observation fd642cb6-bd08-4575-a76d-a41451365bf0 · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 95

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.514613Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.514613Z digest=sha256:db84418ad318136d67c75db8c5cfcc56090354db5085b0a28454f81d560a5943

Observation 10e2c647-8096-410e-bede-17e3d463131c · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 96

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.517861Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.517861Z digest=sha256:e6bebf305eb8172d4904e067e437d82c8038a89416812d083f449e44bdab9de2

Observation 8712754c-c291-4b51-9ca2-36203477e74a · outbound

This paper cites VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks VideoRAG: Retrieval-Augmented Generation with Extreme Long-Context Videos

Reference 97

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.521074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.521074Z digest=sha256:0f48445495cbd416d1fd808f0ef02e1a6697438e4c4f70ba331473fbcb508a08

Observation 94a4611f-a6a2-4342-a19f-f5cbfc809bc6 · outbound

This paper cites V-desirr: Very fast deep embedded single image reflection removal.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks V-desirr: Very fast deep embedded single image reflection removal

Reference 98

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.523879Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.523879Z digest=sha256:04378b3a0af4c64d0a41be58e550e0f9f70b0044070c1ee4f84c22e0b2f16a89

Observation cb64ee82-bb51-4f1b-a730-c2ccfc1ca690 · outbound

This paper cites Measured albedo in the wild: Filling the gap in intrinsics evaluation.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Measured albedo in the wild: Filling the gap in intrinsics evaluation

Reference 99

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.526839Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.526839Z digest=sha256:d394062c0ee90b96cc5c824e85525279a9e02a88cdb5ddf36fa8a660f9275f8c

Observation ccf0e7a5-89b1-422e-b981-79f3699f5364 · outbound

This paper cites Adverb: Visually guided audio dereverberation.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Adverb: Visually guided audio dereverberation

Reference 100

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.530209Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.530209Z digest=sha256:13aeac4040b04df1d5f1b75ea658893dab7f93d00479c70d8d72830185fbca5a

Observation 24a613e3-e1f0-43fe-89fb-6bf65e1a63a2 · outbound

This paper cites Melfusion: Synthesizing music from image and language cues using diffusion models.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Melfusion: Synthesizing music from image and language cues using diffusion models

Reference 101

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.533720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.533720Z digest=sha256:48401a5a7b093f5cdcd2b9b7ca8d7ec3e2d56ad13cb3b1eb5ca26491e5ae04fd

Observation ece97e52-8192-434d-8c9a-7279ccfaf2c2 · outbound

This paper cites Foleygen: Visually-guided audio generation.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Foleygen: Visually-guided audio generation

Reference 102

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.536648Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.536648Z digest=sha256:34b6274f2cd488050249ee10fc3ad2c6e8a8c3fa35fb7f4fc1b1d37417ba8b92

Observation e61e0399-ea2b-4839-8675-591d47f90ef7 · outbound

This paper cites Codi-2: In-context interleaved and interactive any-to-any generation.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Codi-2: In-context interleaved and interactive any-to-any generation

Reference 103

Resolution
unresolved
no resolver link, observed 2026-08-07T05:49:53.539318Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:49:53.539318Z digest=sha256:9798794c1ed31d786b60d426cd83860267d7a320b8b9178ca682df178aa89648

Observation 95e292b2-e55e-489f-8221-178eeb0a9785 · outbound

This paper cites Listen to the pixels.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Listen to the pixels

Reference 104

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:54.393061Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:49:53.542548Z digest=sha256:8144e97400dfcba65ba1d8df9de0296376f3342c33e03bbb7918da9ca973ed80

Observation 8df9726a-60c8-4102-83a5-372640adcbd6 · outbound

This paper cites Audvisum: Self- supervised deep reinforcement learning for diverse audio-visual summary generation.

MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks Audvisum: Self- supervised deep reinforcement learning for diverse audio-visual summary generation

Reference 105

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:49:54.385050Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-08-07T05:49:53.545730Z digest=sha256:e62e18dbf18d00e79aca629dc543c6c0e0ae17cb7e57973f39b6718df62e44ca

Pith citing papers

Observation b7e657f6-9660-4ad7-a8ed-513e02acdef8 · inbound

EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception cites this paper.

EgoAdapt: Adaptive Multisensory Distillation and Policy Learning for Efficient Egocentric Perception MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T22:41:05.001058Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T22:41:05.001058Z digest=sha256:b8ba62f348532ce6ce4a33edafe3e130b71d6a7944352ecbfca0aacaf6793922

Observation dbe2db90-2cb4-4672-b286-ff4a14a70b56 · inbound

EgoSound: Benchmarking Sound Understanding in Egocentric Videos cites this paper.

EgoSound: Benchmarking Sound Understanding in Egocentric Videos MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-15T21:46:42.485110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-05-15T21:44:47.636912Z digest=sha256:0c751af4656cc95d681a90ffac688705fc982924329b482c494862976c20c46e

Observation ea59bb54-f6e6-48c1-b475-63155e155898 · inbound

Through the PRISM: Principle-Aware, Interpretable, and Multi-Scale Evaluation of Visual Designs cites this paper.

Through the PRISM: Principle-Aware, Interpretable, and Multi-Scale Evaluation of Visual Designs MAGNET: A Multi-agent Framework for Finding Audio-Visual Needles by Reasoning over Multi-Video Haystacks

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-06-28T19:22:34.455260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-28T19:19:09.199731Z digest=sha256:790ae2d4a29433e3e090561343b22b908d68272ba611a231190cd96c693772f4