Pith. sign in

Paper Citation Record · LEDGER

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models

As of 13 August 2026, this Paper Citation Record lists 47 of 47 outbound references and 0 inbound Pith citation observations for arXiv:2411.16201.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.16201 v1

Coverage vector

measured 47 of 47 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T13:31:10.814698Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

47 of 47 outbound references displayed

  • verified exact0
  • verified fuzzy6
  • unresolved41
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eca11ac4-b118-4b84-b791-5d36912e82cd · outbound

This paper cites Flamingo: a visual language model for few-shot learning.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Flamingo: a visual language model for few-shot learning

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.521538Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.521538Z digest=sha256:f3fb2fad276d9ce7b81b4feaf92850e2d16a1e523726bb7a9481e79eee20b1d9

Observation 22bd658a-6679-4d56-b378-404ad7ada325 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.527255Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.527255Z digest=sha256:833ca7fcc5c854e3f73cc9cb2389fb76b760e3e76b089729811523874a8918ee

Observation ace32fdb-527b-4fd2-946b-bb00af8afd99 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.533202Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.533202Z digest=sha256:2a2232529f0033ecfa55af249be937d156ac09f88a175323e5ed8a97a8cc2b4e

Observation bd396c2f-d606-4388-97d0-7a2753996647 · outbound

This paper cites LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.539991Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.539991Z digest=sha256:2b8aa35ff1995baa04965fc2dd2d913cfcff8ee8d87d0b9e10d6e93cc1316262

Observation dcb0f39c-00e6-47a1-8c81-6bcd4cf07bb0 · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.545848Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.545848Z digest=sha256:c3efcaf92af6f80463e2644cb290ea68224c4472eae4468c7ed48cd591f9c73c

Observation f2222748-57a1-49c7-8b5f-2c75b0e93df1 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.552918Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.552918Z digest=sha256:a85c09a50878ca20e1339b2c1b24d247b5deed336146ca96ec8fa1a053bd872c

Observation 853a2a25-2629-437f-8fe0-cda44967dfa9 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.559017Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.559017Z digest=sha256:a3cad5c756dd82c5c7dc1547db82492a7d60c34152b9d414ad0f634eadcf90c1

Observation 08710526-0657-4472-8222-192f6e320dff · outbound

This paper cites LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models LanguageBind: Extending Video-Language Pretraining to N-modality by Language-based Semantic Alignment

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.564173Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.564173Z digest=sha256:a6047b6d74ec06a1f442dd253c300a9730e93bdfd7cfda10c187fd91b42f6e8a

Observation f72d6367-31f0-4582-9c47-499e9a7b9109 · outbound

This paper cites Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Direct Preference Optimization of Video Large Multimodal Models from Language Model Reward

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.570432Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.570432Z digest=sha256:5b3098cff369dd90282fe955fb0030fb73e8e28480616f3b44448f1304f8d001

Observation 6b977978-c600-4e54-9f9e-2b06636631bb · outbound

This paper cites PaLM-E: An Embodied Multimodal Language Model.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models PaLM-E: An Embodied Multimodal Language Model

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.576440Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.576440Z digest=sha256:8b91d468ac52bc67cf42c73f2e05c283952b33b00363a20727d150f4075049de

Observation 50ad9708-b22c-4188-9ed3-a8cfdd831e80 · outbound

This paper cites Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.582066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.582066Z digest=sha256:acda6e236751569d12786e2840a68691631c468cac0d78ef2e5e8f3e120113ec

Observation d0eebc09-dac3-4705-9906-85b88d1675ab · outbound

This paper cites Evaluating Object Hallucination in Large Vision-Language Models.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Evaluating Object Hallucination in Large Vision-Language Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.587661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.587661Z digest=sha256:8bfdf544a2ccf8a0189adb3f30f9671f867ed31b5fecf66f462f2bb512d261ae

Observation 2c70d2ac-9b1d-4da8-b970-aff52c43e887 · outbound

This paper cites Aligning Large Multimodal Models with Factually Augmented RLHF.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Aligning Large Multimodal Models with Factually Augmented RLHF

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.594338Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.594338Z digest=sha256:80574042944c69953844185ede188c2728859457df77e4d7aa4369049f998d1c

Observation e4e687df-5993-40e2-9eaa-6a61212db21c · outbound

This paper cites GPT-4o System Card.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models GPT-4o System Card

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.600039Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.600039Z digest=sha256:b51211b869d47df55c0f5d115c80eece32c9d245d29bade4d9e78b2832df9f49

Observation 0118922d-762f-4fdc-bcff-11827726ed82 · outbound

This paper cites Self-Rewarding Language Models.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Self-Rewarding Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.605734Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.605734Z digest=sha256:d2a3683f96a605d4c4206a066e05d1bfe5a1a7d4493e212ed458b099866c07a6

Observation 6897a2d7-2a33-4c9b-8678-c23035d08cca · outbound

This paper cites Direct preference optimization: Your language model is secretly a reward model.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Direct preference optimization: Your language model is secretly a reward model

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.611788Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.611788Z digest=sha256:52885f6825618c19cb070a5e19548f08bcec7203018d9f9eb6ed1aff438c514c

Observation 32855b2e-1994-4440-8ade-2291b1ab1ee9 · outbound

This paper cites Statistical Rejection Sampling Improves Preference Optimization.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Statistical Rejection Sampling Improves Preference Optimization

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.616988Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.616988Z digest=sha256:6d93c7f6010c0f83bb1668e1ee08d4a3100c4367154e7219961f43c622bbdfe0

Observation 8c071df3-41bb-493c-a256-f81fcce82abc · outbound

This paper cites SimPO: Simple Preference Optimization with a Reference-Free Reward.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models SimPO: Simple Preference Optimization with a Reference-Free Reward

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.621710Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.621710Z digest=sha256:ae04953711c6fe3e85effbc632ee9d6e2171cc4be4aa08a24072cf6dd4cde302

Observation fdf776f6-ceca-4d22-9551-1afc8ad641ff · outbound

This paper cites Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Rlhf-v: Towards trustworthy mllms via behavior alignment from fine-grained correctional human feedback

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.627777Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.627777Z digest=sha256:ac185aaa0c55f3ad5acf91d12f59f0bdf04d55698e7c9d0fcc78bc2514cab48f

Observation 9332e79f-d7ce-46ff-82ee-da5bfda82e81 · outbound

This paper cites Silkie: Preference Distillation for Large Visual Language Models.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Silkie: Preference Distillation for Large Visual Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.635516Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.635516Z digest=sha256:348a6e8f77a0bfac94bfb8202be90a235a1421c0be910dc3ac201984b816a1d8

Observation d7e5fbeb-dd43-4e40-a0fe-5cad899f7f54 · outbound

This paper cites Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Tuning Large Multimodal Models for Videos using Reinforcement Learning from AI Feedback

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.642858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.642858Z digest=sha256:91e559e7459012cb82d694e52b4f96c04e92f1522553760044bd093760f1ef26

Observation 88b3a2d7-0385-44ec-8608-466cfb54084c · outbound

This paper cites Rank analysis of incomplete block designs: I.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Rank analysis of incomplete block designs: I

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.650203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.650203Z digest=sha256:a95ca6dd71cb606694d5644a2bd36e4e2dce6d64030e3314115c6a8226765dea

Observation 9a7dfcbb-7853-45c8-8df7-051d97f7878c · outbound

This paper cites Model Extrapolation Expedites Alignment.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Model Extrapolation Expedites Alignment

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.657000Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.657000Z digest=sha256:be79f2bc2dc10dde4339058b8d1df41f09b918ef3307eaff672370d552d79966

Observation f3a7c618-105f-479a-b8e0-f5dd33172327 · outbound

This paper cites Activitynet: A large-scale video benchmark for human activity understanding.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Activitynet: A large-scale video benchmark for human activity understanding

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.663745Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.663745Z digest=sha256:7e48e79dd0acd005a08f162a5ee4f25e323d8bb2379a537392d998f050a89ef5

Observation 191dc859-7ff7-4a0e-998f-720dd03601ae · outbound

This paper cites Frozen in time: A joint video and image encoder for end-to-end retrieval.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Frozen in time: A joint video and image encoder for end-to-end retrieval

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.670735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.670735Z digest=sha256:2e6461872c4ed343fd74040db7ddb09dcb0cb415fd02cef378d9afcbbcc4f64a

Observation 7b80de64-9711-4c68-befb-f88f3430b5dc · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Llama-vid: An image is worth 2 tokens in large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:31:11.745619Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:31:10.676275Z digest=sha256:98e51ba9b0ef2a356c18b52e33c5c7bb70b815b0a41e66442c94b9d6d2a00812

Observation da689573-6cbe-4ae8-b9ce-9f1e8505ddf7 · outbound

This paper cites Llava-next: A strong zero-shot video understanding model, April 2024.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Llava-next: A strong zero-shot video understanding model, April 2024

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.681872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.681872Z digest=sha256:3f987e900241e0f6cfce17250943d5f4a0604f0a8e80389732b20a015ce1989e

Observation ea46d369-dad9-4098-a159-5b1315716da8 · outbound

This paper cites VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.687598Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.687598Z digest=sha256:e917c84c3098ec9f25810a59bc348cb85572949f5e5a8728f56334c176a8441b

Observation fe0c5fab-4a56-46fa-ad5d-29576a414594 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Learning transferable visual models from natural language supervision

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.695262Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.695262Z digest=sha256:338345a5b297afae12ea6cc16beab53f9daa071d633a01ee7c864f9851e32214

Observation ae951733-419c-4de0-8bee-6c3d47550f85 · outbound

This paper cites Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Beyond Hallucinations: Enhancing LVLMs through Hallucination-Aware Direct Preference Optimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.702395Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.702395Z digest=sha256:a2765dcf33acc23e44c0fa0610047d8e741ab6dab1b46a4ec21ede07a63bd4bc

Observation 0f6da3eb-9320-488a-b35a-d615b36ee3fd · outbound

This paper cites Understanding Reference Policies in Direct Preference Optimization.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Understanding Reference Policies in Direct Preference Optimization

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.714278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.714278Z digest=sha256:fa72d4589f4c1b0eb09925724ec212e8fbe7ed35f7df948c9b8f00591557dc2d

Observation 92c1f849-6227-4133-b55c-5fdca36ea965 · outbound

This paper cites Averaging Weights Leads to Wider Optima and Better Generalization.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Averaging Weights Leads to Wider Optima and Better Generalization

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.721135Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.721135Z digest=sha256:8abb300944960bfe8f304727234fcd8b673ff5f88dea41f4388f49edf1074133

Observation b6b0144e-f424-4ec9-89cc-34b2917dd750 · outbound

This paper cites Mitigating the alignment tax of rlhf.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Mitigating the alignment tax of rlhf

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:31:11.688096Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:31:10.730400Z digest=sha256:5271b54ad8c1c16e94a40ec44e54b51211962c8223f9d0f362ff88fc9bf1ce61

Observation d9f3d4ec-cfa5-40b9-8788-4841e330076c · outbound

This paper cites Spurious Feature Diversification Improves Out-of-distribution Generalization.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Spurious Feature Diversification Improves Out-of-distribution Generalization

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.736454Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.736454Z digest=sha256:74f4489989c14534e9129e5565e19aac38572362e82d0c22f6ae711ec38bb4f7

Observation 3fed1565-deae-4028-ac33-0bf978d6e7f1 · outbound

This paper cites Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Model soups: averaging weights of multiple fine-tuned models improves accuracy without increasing inference time

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.743616Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.743616Z digest=sha256:a81943afa104880d5d85baf9b6d34fa8f87ec5c53ea67e60fe661ae8c3e0ccf6

Observation 08502e6c-3882-4202-b606-44fdd5c2b591 · outbound

This paper cites Msr-vtt: A large video description dataset for bridging video and language.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Msr-vtt: A large video description dataset for bridging video and language

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.751541Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.751541Z digest=sha256:eda15c3ea82ffacbfa7ef0ff959f2885d7c9302b1a10adc214575e088c7017ec

Observation 4448ab51-afb0-46ab-a2be-246f99138b23 · outbound

This paper cites Collecting highly parallel data for paraphrase evaluation.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Collecting highly parallel data for paraphrase evaluation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:31:11.646667Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:31:10.759683Z digest=sha256:733b005596837be441b7b01dd9fc0d02eb7dc2886578d066811ad394ce8c6765

Observation b7491e61-2905-4ed1-9a7f-8e0f7626bc67 · outbound

This paper cites Tgif: A new dataset and benchmark on animated gif description.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Tgif: A new dataset and benchmark on animated gif description

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.765343Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.765343Z digest=sha256:3d49603302e6093b74f18c50e786d83c51aa32ea11b3a7829e22bcd49d219c0d

Observation f9000a92-d14b-4285-8f02-7219c1b1f778 · outbound

This paper cites something something.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models something something

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.770443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.770443Z digest=sha256:f7526db64b55c6d594141332ef181aba18a77f88f03c9aca199896e3ac16ac4a

Observation 7b5a80dc-c49c-460d-96ba-9fa08f53d769 · outbound

This paper cites Eva: Exploring the limits of masked visual representation learning at scale.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Eva: Exploring the limits of masked visual representation learning at scale

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.776113Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.776113Z digest=sha256:4fd7773b83a50e3e1034f3088c32fd1f6b975f353c600da5fbdaa0049c952311

Observation 4499c109-e71c-4b1a-a225-533eb47d1388 · outbound

This paper cites Vision transformer with quadrangle attention.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Vision transformer with quadrangle attention

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:31:11.591064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:31:10.781161Z digest=sha256:94dfc67363630795ba34b931f555ff9b59f94453a5686c7d8eb2b6c0eacad4c3

Observation dc08f4c8-2f55-45f4-801d-1e6d1db9056e · outbound

This paper cites Judging llm-as-a-judge with mt-bench and chatbot arena.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Judging llm-as-a-judge with mt-bench and chatbot arena

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.786726Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.786726Z digest=sha256:d014362cf1f030bbd684ee390c78909ade8e9cda8db8bbb9e8142fafbd481a0a

Observation ea81f325-c03c-41fb-8884-478455918ba4 · outbound

This paper cites Decoupled Weight Decay Regularization.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Decoupled Weight Decay Regularization

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-12T13:31:10.792205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T13:31:10.792205Z digest=sha256:1ce971e48b6b3299b6f6d14536d78eb7b635d2a02ff8283e1afa9a3b65c68843

Observation 28c2e256-00f5-4b48-bbb5-c1a4675cc820 · outbound

This paper cites an unresolved cited work.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-12T13:31:11.560023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:31:10.797644Z digest=sha256:91a15f7b8d50cf7b54b124cc01bb373c390495e755174dafcab6f91eb151931f

Observation eb050356-f8a8-4c6f-96bb-d281a68530e1 · outbound

This paper cites an unresolved cited work.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Unresolved cited work

Reference 45

Resolution
unresolved
raw_fallback, observed 2026-08-12T13:31:11.536951Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:31:10.803316Z digest=sha256:3d92f09498b598638ed60008b7599772c140e9a8159da6f2468af12ba0b047e0

Observation 43217553-0d2e-4953-9722-a874f9a72be0 · outbound

This paper cites Consider the following criteria for evaluation: -**Relevance**:Evaluate how relevant the model's predicted answer is to the question asked.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Consider the following criteria for evaluation: -**Relevance**:Evaluate how relevant the model's predicted answer is to the question asked

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:31:11.513344Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:31:10.808796Z digest=sha256:ec2acd17391b64ab16ba66929cad6e0834552703423fc32c2a03ba8dfea382b2

Observation ddb17ddb-16c1-414f-9989-3f1d82ec9922 · outbound

This paper cites Temporal Consistency:Does the answer appropriately reflect the temporal progression and events of the video?3.

Video-Text Dataset Construction from Multi-AI Feedback: Promoting Weak-to-Strong Preference Learning for Video Large Language Models Temporal Consistency:Does the answer appropriately reflect the temporal progression and events of the video?3

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T13:31:11.492286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-08-12T13:31:10.814698Z digest=sha256:656aea6e4e0e10e83216a8344f9da0731cd4aa84225e2306c82e31bc0c409c6c

Pith citing papers

No inbound Pith citation observations are available.