Pith. sign in

Paper Citation Record · LEDGER

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment

As of 13 August 2026, this Paper Citation Record lists 68 of 68 outbound references and 2 inbound Pith citation observations for arXiv:2501.18314.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2501.18314 v2

Coverage vector

measured 68 of 68 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-10T00:02:26.319596Z

measured 70 of 70 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:29:05.574108Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-06T05:00:19.327534Z

Reference resolution

68 of 68 outbound references displayed

  • verified exact2
  • verified fuzzy43
  • unresolved23
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fdf31bf0-5965-4ee2-8d37-07a1a9098af5 · outbound

This paper cites Prime Voice AI, 2023.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Prime Voice AI, 2023

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:27.200658Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.074865Z digest=sha256:161ec45503f5b3bef257dd40e297de8b43bb6d8044c9a812a60527d7e90121fc

Observation 7c0b15e2-e613-4299-bc34-ce60acd262d7 · outbound

This paper cites Accessed December 28, 2023 [Online] https://www.pika.art/, 2023.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Accessed December 28, 2023 [Online] https://www.pika.art/, 2023

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:27.187743Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.079755Z digest=sha256:d90a81a4a2a59c63226c2fe5ae75f25793e73cba0a86b60ac6aea5968703579b

Observation 21cbe542-4314-4966-a68f-7f2eb23af791 · outbound

This paper cites Accessed June 17, 2024 [Online] https://runwayml.com/research/introducing-gen-3-alpha, 2024.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Accessed June 17, 2024 [Online] https://runwayml.com/research/introducing-gen-3-alpha, 2024

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:27.175505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.083409Z digest=sha256:e9832a072c5a7f8c8e934185c3d16dea74373646807200cb99ddb7dd0f29c97d

Observation 46e2f2ee-4c14-4881-a339-80f2d8423dca · outbound

This paper cites Accessed June 6, 2024 [Online] https://klingai.kuaishou.com/, 2024.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Accessed June 6, 2024 [Online] https://klingai.kuaishou.com/, 2024

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:27.164098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.086752Z digest=sha256:7bd935969b0bddf5dfb16eba59f84027aa30414bf23c388b71db4caa5db92be2

Observation 92a46a10-49f5-4508-8728-34d5a329ffb4 · outbound

This paper cites Accessed February 16, 2024 [Online] https://openai.com/sora/, 2024.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Accessed February 16, 2024 [Online] https://openai.com/sora/, 2024

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:27.152193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.089994Z digest=sha256:97901b406787dc6cae025ffa0ab2e767f5e325a5f0307d2e140daeab34520604

Observation 67604226-62bc-4995-8c68-4cb83de72b8f · outbound

This paper cites MusicLM: Generating Music From Text.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment MusicLM: Generating Music From Text

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-10T00:02:26.093378Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:02:26.093378Z digest=sha256:5b9d824d59d2bf017445846dc54cd2809545475870016bcecc4dea583224f640

Observation 54079295-2452-42d4-aea7-2fe544a8b568 · outbound

This paper cites G., Schmidmer, C., Berger, J., Obermann, M., Ullmann, R., Pomy, J., and Keyhl, M.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment G., Schmidmer, C., Berger, J., Obermann, M., Ullmann, R., Pomy, J., and Keyhl, M

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:27.140844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.097895Z digest=sha256:148d3495c86a1583676e459615563b8eee3c77c7475c962d1910ebcdb6d927d8

Observation b5aba429-cc3e-46f8-adc6-3f7f1351b3d9 · outbound

This paper cites Attention-guided neural networks for full-reference and no-reference audio-visual quality assessment.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Attention-guided neural networks for full-reference and no-reference audio-visual quality assessment

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:27.130287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.101036Z digest=sha256:410e2c1d22aa828321e9f293181a237c42e7e985b88d280238aea16d2c06326d

Observation 0b5b111c-0ede-4c07-8e30-3bce1046ada1 · outbound

This paper cites Subjective and objective audio-visual quality assessment for user generated content.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Subjective and objective audio-visual quality assessment for user generated content

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:27.118856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.104801Z digest=sha256:be8ff1fe0ad2ef63a68b6c72c791472b67e0bb152f95f3ce84141744c19294c5

Observation f8009c33-1708-4b14-8d49-2048beb884b1 · outbound

This paper cites Semantically consistent Video-to-Audio Generation using Multimodal Language Large Model.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Semantically consistent Video-to-Audio Generation using Multimodal Language Large Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-10T00:02:26.108555Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:02:26.108555Z digest=sha256:c46f2597c223ad191118fd820c53013d8fb99598a693682a94252c5dca94aa67

Observation 9f910086-50e1-41d5-bf02-9b05428547b6 · outbound

This paper cites Vggsound: A large-scale audio-visual dataset.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Vggsound: A large-scale audio-visual dataset

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:27.108116Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.112847Z digest=sha256:b0ab7499ded7bc4745ccbcc305af02e26b68dce3fe713e57a96094c4017380fe

Observation f0af9eec-e1f8-4f70-a573-8304309ffba5 · outbound

This paper cites Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Vast: A vision-audio-subtitle-text omni-modality foundation model and dataset

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:27.096437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.116824Z digest=sha256:76f3a69e3e866ec19e78cd3e25beffe54866bae4afa60cec9e1024fce6c9b57b

Observation 9da84270-4d0e-43ae-a854-9f56b0425268 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-10T00:02:26.120459Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:02:26.120459Z digest=sha256:7cc14e4d5153d9f1bc2ac763def5d741c7f615252e42a1a18e0de6b3ac81772f

Observation 8c990ca2-3e07-4f28-81de-fabe653a10d8 · outbound

This paper cites Simple and controllable music generation.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Simple and controllable music generation

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:27.084185Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.124527Z digest=sha256:207e44e74848019439e671fa2f350cd633f5692a1bd73c21f0999dae53dd9c17

Observation b185952f-2392-43f5-a3e2-bff21ac70847 · outbound

This paper cites PAM: Prompting Audio-Language Models for Audio Quality Assessment.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment PAM: Prompting Audio-Language Models for Audio Quality Assessment

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-10T00:02:26.128138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:02:26.128138Z digest=sha256:77c7da05176d86eecc6d578bad9845143f709af1cef8466b9ad581743a4ae7f3

Observation 5ea53f22-e938-4b26-8d79-e8b12116018a · outbound

This paper cites Toward universal text-to-music retrieval.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Toward universal text-to-music retrieval

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:27.071912Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.132102Z digest=sha256:92c13eb5e147826a7daf88299cd17e61d7bcc12af80e3375013d3d20d4cadc4f

Observation 9aaf1ed9-a3e1-4064-9eb8-702427f8af1f · outbound

This paper cites Clap learning audio concepts from natural language supervision.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Clap learning audio concepts from natural language supervision

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:27.059316Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.135897Z digest=sha256:07691ec06018b82d6408f8d3c286b45be3f81aeb825a015cfd120264b9877bb7

Observation 7b3f034a-05a7-460c-97fd-1be3d9d0069f · outbound

This paper cites VITA: Towards Open-Source Interactive Omni Multimodal LLM.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment VITA: Towards Open-Source Interactive Omni Multimodal LLM

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-10T00:02:26.139518Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:02:26.139518Z digest=sha256:6e75cd6614ddfd243ec448bdaadd0ccfd17048587e46bc8ea8a3beb11fec0d0a

Observation f2e25c47-60b7-42a2-a8ee-70597ef18393 · outbound

This paper cites Gotta Hear Them All: Towards Sound Source Aware Audio Generation.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Gotta Hear Them All: Towards Sound Source Aware Audio Generation

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-10T00:02:26.143618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:02:26.143618Z digest=sha256:38b2e088b30d676a24ff9926fe13a1d280bf2038698e835073a0bce465a99133

Observation b1a8fd6b-f6cf-41f4-a5d7-5944ee8d1e82 · outbound

This paper cites AnimateDiff : Animate your personalized text-to-image diffusion models without specific tuning.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment AnimateDiff : Animate your personalized text-to-image diffusion models without specific tuning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:27.046444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.147525Z digest=sha256:f9202e3504d4da2559a4f282c6c53d453b7093c5c323ba1a688baa6a213f0a93

Observation 95381a4c-0dce-46b3-b59e-285693eb614e · outbound

This paper cites Onellm: One framework to align all modalities with language.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Onellm: One framework to align all modalities with language

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:27.035546Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.151192Z digest=sha256:8d9f5bffec39e8d0b1752c6da0cc06f37d64fb2e84a1e24f0c603b9dace097c8

Observation 9cc67da7-92c8-46f2-a49a-51bac96a791f · outbound

This paper cites an unresolved cited work.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-10T00:02:27.023540Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.154855Z digest=sha256:8d62a26d5e7ff949dde2d5f95bf77bcf8339d041a00de447a88dd0fc46e8799a

Observation cefd2335-c222-4665-a4ee-3d47f23f13af · outbound

This paper cites CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-10T00:02:26.158614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:02:26.158614Z digest=sha256:76eec632a878016090e2eb83c9647d50c67e0d0651c44a1ea83e499ffe7b4a12

Observation 53accb63-7a54-44d1-bfd5-4d6a6c717ec7 · outbound

This paper cites Video-to-Audio Generation with Fine-grained Temporal Semantics.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Video-to-Audio Generation with Fine-grained Temporal Semantics

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-10T00:02:26.162701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:02:26.162701Z digest=sha256:3ddc4a109fdd5209570cf91729a395eee7b6650d3699b97630af48d6d4a1ee1c

Observation 5d45ad44-df55-4936-996f-d2a5faf2cf52 · outbound

This paper cites Vbench: Comprehensive benchmark suite for video generative models.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Vbench: Comprehensive benchmark suite for video generative models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:27.009115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.166500Z digest=sha256:cdab5f6bff1acec0ffeaf79873d2130dda9764ea2bfa39318fe5889ef8bc5577

Observation 073488ba-5728-47ff-9e5d-05e5cf7d8ba3 · outbound

This paper cites GPT-4o System Card.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment GPT-4o System Card

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-10T00:02:26.171351Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:02:26.171351Z digest=sha256:2322df70b0366a57d8f2ef29daafc5e02e9fbc449f10c75433be2850b8c2547e

Observation e6f7b572-5f85-4443-939e-3414a4ad8fa8 · outbound

This paper cites Taming Visually Guided Sound Generation.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Taming Visually Guided Sound Generation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-10T00:02:26.175169Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:02:26.175169Z digest=sha256:df515fb8813efd38b05d8c5ec7d4d3f62590b3863e448d13501a9dbe9f0b7c27

Observation ad621faf-ae9f-4405-9a81-1df5f46d6ee6 · outbound

This paper cites Read, watch and scream! sound generation from text and video.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Read, watch and scream! sound generation from text and video

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:26.996922Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.178962Z digest=sha256:a42b099de3c9563415d0463b8cdac431e2fde2b148ae85cf781102b7e951917f

Observation 76f07674-9edd-45ae-86cc-f7209d1b0187 · outbound

This paper cites T2VBench : Benchmarking temporal dynamics for text-to-video generation.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment T2VBench : Benchmarking temporal dynamics for text-to-video generation

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:26.985567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.182673Z digest=sha256:0a8b0c986eadf5f04af275c422e490e8352abdf285e63a69abcec772927e7603

Observation 405c95da-158c-48a6-90e4-7c1e87db9807 · outbound

This paper cites D., Kim, B., Lee, H., and Kim, G.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment D., Kim, B., Lee, H., and Kim, G

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:26.973264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.185848Z digest=sha256:4af7f810c85dd4b67b2c1334a96346c86ac5b427a57c9fa7bdcd4be11a7cb13e

Observation bc894573-ea6a-4ad9-8454-bd76e50531c5 · outbound

This paper cites Subjective-aligned dataset and metric for text-to-video quality assessment.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Subjective-aligned dataset and metric for text-to-video quality assessment

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:26.961500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.189063Z digest=sha256:93a20860bc80cacdafaba4b02398cc2747a1afa528812d0772d7f8021e21a6d0

Observation 0423f0dc-9b3c-48f9-9046-375a6fed13ca · outbound

This paper cites AudioGen: Textually Guided Audio Generation.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment AudioGen: Textually Guided Audio Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-10T00:02:26.192304Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:02:26.192304Z digest=sha256:ef5e5cda5ba2953678fff229612768ea58e5a2f7b6709b675051ed22173b242d

Observation 71c8e2b8-dd21-450b-910e-5d819958bc9a · outbound

This paper cites Groundinggpt: Language enhanced multi-modal grounding model.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Groundinggpt: Language enhanced multi-modal grounding model

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:26.950750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.195824Z digest=sha256:c4460de0c8994f45a18e35f63d907c02e810dd5e7c9e1dbe388a6ea1614250df

Observation 25eba396-9234-4e1d-b685-16c12c5f394d · outbound

This paper cites VALOR : Vision-audio-language omni-perception pretraining model and dataset.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment VALOR : Vision-audio-language omni-perception pretraining model and dataset

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:26.938998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.199202Z digest=sha256:635d66ddf86322fdb0bcf4c3b57db59bcf5bd3b9db568f8ef07a94e53e02dfb7

Observation d14bee3c-402e-427f-820d-a7e57fb52728 · outbound

This paper cites Fetv: A benchmark for fine-grained evaluation of open-domain text-to-video generation.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Fetv: A benchmark for fine-grained evaluation of open-domain text-to-video generation

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:26.926543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.202480Z digest=sha256:f1837271d7c00ad1395f7690f289f5e93e230a0f263651a8473dcdbcded51e09

Observation 4f97cb71-53a0-442b-ae76-0286d41da35d · outbound

This paper cites Mosnet: Deep learning based objective assessment for voice conversion.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Mosnet: Deep learning based objective assessment for voice conversion

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:26.908571Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.205464Z digest=sha256:d4de7dd585e6ce15e633e6ca0361d3139fc6968b8ca3566c3eeee6d2add68238

Observation e32b041a-0311-420d-8a11-f4ef76a4ad8f · outbound

This paper cites Unified-IO 2: Scaling autoregressive multimodal models with vision language audio and action.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Unified-IO 2: Scaling autoregressive multimodal models with vision language audio and action

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:26.890302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.208314Z digest=sha256:791606c0023b847d1732f1ef746082fa2673851444573e5a70ba3a51b9e9eb1c

Observation 9de5644b-981c-4461-9eb9-988615a2d865 · outbound

This paper cites Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Diff-foley: Synchronized video-to-audio synthesis with latent diffusion models

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:26.864803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.211651Z digest=sha256:4f97ed9111970ed7c1b38d2939e876184c613a9063f35e3d443a478f15bcf252

Observation 781d5118-91c8-4bd0-859c-7de8313a11cf · outbound

This paper cites Speech Quality Assessment through MOS using Non-Matching References.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Speech Quality Assessment through MOS using Non-Matching References

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-10T00:02:26.215147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:02:26.215147Z digest=sha256:bc652ae8bbfd47dbca0f47c61c5f6f6ce1a3c079001ff5148cc99374a1fbfcfa

Observation 87d68138-c513-4a26-8ed8-011dae96b6dc · outbound

This paper cites C., and Bovik, A.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment C., and Bovik, A

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:26.849603Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.218837Z digest=sha256:53640aaf9c2695eef44263f4a561ba369bc46838155619062036f9a80cb7a0d5

Observation f69d7519-19da-4049-a174-c2f0afd16265 · outbound

This paper cites Nisqa: A deep cnn-self-attention model for multidimensional speech quality prediction with crowdsourced datasets.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Nisqa: A deep cnn-self-attention model for multidimensional speech quality prediction with crowdsourced datasets

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:26.831081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.222182Z digest=sha256:61ed0a5acd20bfb1dd0382f9161e9607c0d94e776b4a1474dd33af273be58a53

Observation a318e777-902e-41cd-a38b-63a683fa931b · outbound

This paper cites Audio-visual instance discrimination with cross-modal agreement.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Audio-visual instance discrimination with cross-modal agreement

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:26.815819Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.226070Z digest=sha256:6f8412b199abfb1799a76fcafd50fa200909fca4605b247d5e9685bf4ee70781

Observation 0173fc3d-43ff-461d-8c7a-f56bc89cdaf7 · outbound

This paper cites STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment STA-V2A: Video-to-Audio Generation with Semantic and Temporal Alignment

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-10T00:02:26.229598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:02:26.229598Z digest=sha256:2250d5d9a6e17e4c388f99d78df2831307d511363fecb791736c729cc2841a53

Observation 76b4a802-6fbc-4208-b57d-19e5ffa2df38 · outbound

This paper cites W., Beerends, J.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment W., Beerends, J

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:26.798275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.233649Z digest=sha256:a939dcadd02e604c17f141292197b2725003c8ef1d0a7411bf25131ebb6ab8a4

Observation b427e816-91cf-41df-910b-1708a806493d · outbound

This paper cites and Adi, Y.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment and Adi, Y

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:26.782611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.237491Z digest=sha256:56aad9d16b32d63e05d2dbd75ea6f037b6e69e9504a2a64bde8bad6f911fe022

Observation dcfaa58c-c1fb-4998-ab91-54afefd11800 · outbound

This paper cites PandaGPT: One Model To Instruction-Follow Them All.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment PandaGPT: One Model To Instruction-Follow Them All

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-10T00:02:26.241116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:02:26.241116Z digest=sha256:cfb7bc082d7c4b634bc19215e3e139469f7f964ea5059c6e3b1df1ee4d6f0565

Observation 7d449d2c-7c70-461c-89cd-588cb68bb726 · outbound

This paper cites The influence of text-guidance on visual attention.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment The influence of text-guidance on visual attention

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:26.769698Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.244867Z digest=sha256:1085b7980bb30b3b4e010a78fcdb59cd43b1c820e95ddcd5edd741c2e92a046e

Observation 120815b1-bc21-40e2-bfd9-36e100e1ba76 · outbound

This paper cites How is visual attention influenced by text guidance? database and model.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment How is visual attention influenced by text guidance? database and model

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:26.754506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.248517Z digest=sha256:d41aaf79fd6467619e6958185f8145ff1533a0f00fbce6e22288c8d0a8226b4b

Observation 8e2c7403-36c5-456d-8351-a3f6fc1535da · outbound

This paper cites Explore the Hallucination on Low-level Perception for MLLMs.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Explore the Hallucination on Low-level Perception for MLLMs

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-08-10T00:02:26.450950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.252376Z digest=sha256:b41459d845ca4b6663c17000024a9c69b93d29a369acab613b4a52ed61797a54

Observation acb85e0e-9cfe-4d1f-8f7e-428eaa8e3022 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-10T00:02:26.256454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:02:26.256454Z digest=sha256:cb2e2afd2ac7bb3f3d3228d2160a98e3d4f6a7572a61202ab5604c7ec2d205d8

Observation aa8832c5-e30d-471c-a17d-7504612c1fe3 · outbound

This paper cites Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-10T00:02:26.260452Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:02:26.260452Z digest=sha256:ef60c0ac6d1049d88e36d83264bdb5bd7e085d6a75f3cf7b8a4aecc38630a98c

Observation af93185d-f391-436d-bdd6-ecd0f6fea59f · outbound

This paper cites V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment V2a-mapper: A lightweight solution for vision-to-audio generation by connecting foundation models

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:26.741748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.264346Z digest=sha256:298806369c26a8135ce2071a8e34e62fdbd6949211fa352497a518988946c5d9

Observation a06682c8-17b6-480a-b6c1-0b2897661010 · outbound

This paper cites Quality Assessment for AI Generated Images with Instruction Tuning.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Quality Assessment for AI Generated Images with Instruction Tuning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-10T00:02:26.267826Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:02:26.267826Z digest=sha256:967165708b7514528b5b0b79342d6d972a4a8e45223c8c9a229aecf0419a1b59

Observation cebbe6fc-9da6-49a0-a4a0-2a4688c00092 · outbound

This paper cites AIGV-Assessor: Benchmarking and Evaluating the Perceptual Quality of Text-to-Video Generation with LMM.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment AIGV-Assessor: Benchmarking and Evaluating the Perceptual Quality of Text-to-Video Generation with LMM

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-10T00:02:26.271791Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:02:26.271791Z digest=sha256:6900ac8ec1088fcb11d20d2d6396d80242c7f3ff46001c80eb4cb7f3925d70a9

Observation 0a76a232-dce4-41ed-abc9-687eea36cec2 · outbound

This paper cites Tiva: Time-aligned video-to-audio generation.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Tiva: Time-aligned video-to-audio generation

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:26.729335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.275625Z digest=sha256:794de484c92cd59bb6f711ab86b84522aea07f71d69343eca13aa16bf3ced42e

Observation 2fb4dc11-4839-43de-bd19-a2a57c1a7e7a · outbound

This paper cites A recipe for scaling up text-to-video generation with text-free videos.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment A recipe for scaling up text-to-video generation with text-free videos

Reference 56

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:26.714931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.279119Z digest=sha256:7c6b354076084449ba2776278e1cbaff7e978b8ec6ca7c2f61d65268dd83df0f

Observation 018796f9-3fd2-43e6-b295-73f652f6c019 · outbound

This paper cites Modaverse: Efficiently transforming modalities with llms.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Modaverse: Efficiently transforming modalities with llms

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:26.703112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.282662Z digest=sha256:2639a1ee0594c15db09a99fe7afbeeeff0f0d858a8632528c244d26fddab66c6

Observation 4e73897b-2810-45a3-94e7-3160415148e0 · outbound

This paper cites LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment LAVIE: High-Quality Video Generation with Cascaded Latent Diffusion Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-10T00:02:26.286357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:02:26.286357Z digest=sha256:c1ac3ed22c4a2a6ce181209a5f6ace3571c966986f6b6eac48d14ee4d58d4679

Observation 3a242a52-48b2-4cec-9a7b-8c4e1b088951 · outbound

This paper cites Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Frieren: Efficient Video-to-Audio Generation Network with Rectified Flow Matching

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-10T00:02:26.290373Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:02:26.290373Z digest=sha256:d22a6ed26b0ac37c4f292e4046512993013d75837dc3347dedbc2a34e22de930

Observation 11e0a367-81ea-43cb-9175-79584784ff06 · outbound

This paper cites Q-Align : Teaching LMMs for visual scoring via discrete text-defined levels.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Q-Align : Teaching LMMs for visual scoring via discrete text-defined levels

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:26.689470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.293796Z digest=sha256:c27e526755cedde61381933271ddcc9dd3d4b61ac6887fd80ed41fcee8728381

Observation f2864ae5-a3c3-4d5d-806b-0b7be6a7b944 · outbound

This paper cites NExT-GPT : Any-to-any multimodal llm.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment NExT-GPT : Any-to-any multimodal llm

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:26.675299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.297158Z digest=sha256:a9a433cab858c73641467d4ac45e54c0da115c51e6deb653ff59e3e7dd6ad081

Observation 50d2b6f4-94a9-42b8-a00c-386de598363e · outbound

This paper cites Sonicvisionlm: Playing sound with vision language models.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Sonicvisionlm: Playing sound with vision language models

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:26.661344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.300422Z digest=sha256:977dfc40dce69ef2141b3a94fc9f3818aa9ee3caa94de68ae402fe83e3d43228

Observation 910fe5ed-f1b0-48c2-8584-ca20a38da188 · outbound

This paper cites Efficient Video to Audio Mapper with Visual Scene Detection.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Efficient Video to Audio Mapper with Visual Scene Detection

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-08-10T00:02:26.371834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.303436Z digest=sha256:50adc0bd0bb2472f4412ed7a9859345088e29bc900ddb308781c281efb1d9719

Observation 795ccb3c-3b8b-435e-840b-12bb25dca64c · outbound

This paper cites Telepresence video quality assessment.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Telepresence video quality assessment

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:26.644619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.306497Z digest=sha256:16bb7b94ab8c30208c65cb714f08ae4ecb4b2c4c37b14605b9378d18c42e5f78

Observation 22a6b40c-e1a5-4ea8-8e30-a5d03663809d · outbound

This paper cites E., Fu, S.-W., Fuh, C.-S., Tsao, Y., and Wang, H.-M.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment E., Fu, S.-W., Fuh, C.-S., Tsao, Y., and Wang, H.-M

Reference 65

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:26.632511Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.309233Z digest=sha256:adff6eae8dedd01500f389e46fed16d0ea20e9d575668434c5aecca5b04c26cc

Observation 24cb6520-edae-4e6e-ab85-24c89b42df46 · outbound

This paper cites FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment FoleyCrafter: Bring Silent Videos to Life with Lifelike and Synchronized Sounds

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-10T00:02:26.312439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:02:26.312439Z digest=sha256:f07fc7c3c29edecd893b6d63494a8091d2cc37e66b3273f6dd822f22c67c73cd

Observation 0e247d60-34a4-48fc-93ba-a0493d534c6d · outbound

This paper cites Lmm-pcqa: Assisting point cloud quality assessment with lmm.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment Lmm-pcqa: Assisting point cloud quality assessment with lmm

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-08-10T00:02:26.620940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-10T00:02:26.316126Z digest=sha256:3c486b4992a52c51ceb3e198db3c2929255168787c4606b582fc4e26cd650056

Observation 07557746-ac0d-4d5e-b0b6-243fc6bfc735 · outbound

This paper cites write newline.

AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment write newline

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-10T00:02:26.319596Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T00:02:26.319596Z digest=sha256:033cb48af5e6169fee8b1b2f118a5ef408f5b54d4f7fa84f3ac01ca8be4c77be

Pith citing papers

Observation 13e43b0e-11d7-4144-a940-7bd6e6e57ace · inbound

Efficient Face Image Quality Assessment via Self-training and Knowledge Distillation cites this paper.

Efficient Face Image Quality Assessment via Self-training and Knowledge Distillation AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:05.574108Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:05.574108Z digest=sha256:08220c982538e83e2cafa124b117ca16aa4fb7778d80d757302dc8b04dac00a1

Observation cfc4fff1-41c7-4d93-badb-65f6addd245b · inbound

Engagement Prediction of Short Videos with Large Multimodal Models cites this paper.

Engagement Prediction of Short Videos with Large Multimodal Models AGAV-Rater: Adapting Large Multimodal Model for AI-Generated Audio-Visual Quality Assessment

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-08-06T05:00:19.331823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-06T05:00:14.789121Z digest=sha256:c8838b735827575b14dd412a3f7384a72f99e521b9c9fb12044fc2a9b2868f8d