Pith. sign in

Paper Citation Record · LEDGER

Team of One: Cracking Complex Video QA with Model Synergy

As of 8 August 2026, this Paper Citation Record lists 9 of 9 outbound references and 0 inbound Pith citation observations for arXiv:2507.13820.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.13820 v1

Coverage vector

measured 9 of 9 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:19:20.572832Z

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

9 of 9 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7942b469-8135-4613-b723-615543de5f56 · outbound

This paper cites write newline.

Team of One: Cracking Complex Video QA with Model Synergy write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:19:19.594085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:19:19.594085Z digest=sha256:10a436758b997de03fd2e44e6dd7c7f6f59f1ce385d75c068ddbae5798556392

Observation a26eab60-698e-4f3f-a062-4c57e6fe0641 · outbound

This paper cites Gemini 2.5 Pro Preview (2025-03-25) , 2025.

Team of One: Cracking Complex Video QA with Model Synergy Gemini 2.5 Pro Preview (2025-03-25) , 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:19:21.159092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:19:19.680262Z digest=sha256:1c6714bfae2a0f18f25c8785468f6935cd8b601d8a8aa19b793bff35ca3f1ab6

Observation 51d3faf4-ad0f-42bd-afbe-a5773eadb2e6 · outbound

This paper cites How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs.

Team of One: Cracking Complex Video QA with Model Synergy How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T16:19:19.770396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:19:19.770396Z digest=sha256:1df81ba07a306b74dfac677387c7aee272d533b706806028a1592513d4a66752

Observation fb465385-0d6d-433f-ad89-d4a14ff65ae1 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Team of One: Cracking Complex Video QA with Model Synergy Llama-vid: An image is worth 2 tokens in large language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:19:20.890382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-06T16:19:19.916856Z digest=sha256:ac684a5b3ed0b49f337b801598879214cbc2322f7a5cadd58a5ca2b2386b42f4

Observation c4b533c1-c4c3-466b-99f3-ac902afaac00 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Team of One: Cracking Complex Video QA with Model Synergy Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T16:19:20.053425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:19:20.053425Z digest=sha256:dc25bfe456175e457539df20e27c969f1245ce4d7808fd800a405f30928ac0a1

Observation 7c45b3d7-ddaa-431b-adf1-0af8a14fca00 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Team of One: Cracking Complex Video QA with Model Synergy Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:19:20.204460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:19:20.204460Z digest=sha256:1b34f274e22b62dafe5ae824188ba1f6248dfc0a8daf921cc13be0f82cebc7d2

Observation 3f310098-559d-497b-aa17-2aa12af0fde3 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Team of One: Cracking Complex Video QA with Model Synergy Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:19:20.329167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:19:20.329167Z digest=sha256:6d073c7b016e88294dd5118c3aeba65bec97ac6ae804a7cf8ec38f92d2bc9484

Observation df85bc52-675d-40b2-9061-7dfa2486f749 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Team of One: Cracking Complex Video QA with Model Synergy Chain-of-thought prompting elicits reasoning in large language models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:19:20.430057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:19:20.430057Z digest=sha256:79a91710fb8bd0a6bb4423552b46376706129953f68d5f58956e3f3de498d781

Observation 1b7c2315-9f7b-48f5-9ee5-19ff8e9f5dee · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Team of One: Cracking Complex Video QA with Model Synergy Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:19:20.572832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:19:20.572832Z digest=sha256:c3ee7a2e8e4e37eeed14545dfb48f20f432df22bb3a90779ee950586979e1d52

Pith citing papers

No inbound Pith citation observations are available.