Pith. sign in

Paper Citation Record · LEDGER

Team of One: Cracking Complex Video QA with Model Synergy

As of 16 August 2026, this Paper Citation Record lists 9 of 9 outbound references and 0 inbound Pith citation observations for arXiv:2507.13820.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.13820 v1

Coverage vector

measured 9 of 9 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T16:19:20.572832Z

measured 9 of 9 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

9 of 9 outbound references displayed

  • verified exact0
  • verified fuzzy2
  • unresolved7
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 7942b469-8135-4613-b723-615543de5f56 · outbound

This paper cites write newline.

Team of One: Cracking Complex Video QA with Model Synergy write newline

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T16:19:19.594085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:19:19.594085Z digest=sha256:10a436758b997de03fd2e44e6dd7c7f6f59f1ce385d75c068ddbae5798556392

Observation a26eab60-698e-4f3f-a062-4c57e6fe0641 · outbound

This paper cites Gemini 2.5 Pro Preview (2025-03-25) , 2025.

Team of One: Cracking Complex Video QA with Model Synergy Gemini 2.5 Pro Preview (2025-03-25) , 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:19:21.159092Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T16:19:19.680262Z digest=sha256:9af56f31f07dfed75c2981d3e13feb0414e9543b91b341b919901e306e9533ae

Observation 51d3faf4-ad0f-42bd-afbe-a5773eadb2e6 · outbound

This paper cites How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs.

Team of One: Cracking Complex Video QA with Model Synergy How Good is my Video LMM? Complex Video Reasoning and Robustness Evaluation Suite for Video-LMMs

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-06T16:19:19.770396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:19:19.770396Z digest=sha256:7264d94f82f9c9e9818529432c9493d9263cb081dd9f2b30d46259411556f223

Observation fb465385-0d6d-433f-ad89-d4a14ff65ae1 · outbound

This paper cites Llama-vid: An image is worth 2 tokens in large language models.

Team of One: Cracking Complex Video QA with Model Synergy Llama-vid: An image is worth 2 tokens in large language models

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T16:19:20.890382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-08-06T16:19:19.916856Z digest=sha256:0fdfca3fcf8422e4e35bfc5d53f63c70a0b49fab4ea2a42f2781f4371e6d20f3

Observation c4b533c1-c4c3-466b-99f3-ac902afaac00 · outbound

This paper cites Video-LLaVA: Learning United Visual Representation by Alignment Before Projection.

Team of One: Cracking Complex Video QA with Model Synergy Video-LLaVA: Learning United Visual Representation by Alignment Before Projection

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T16:19:20.053425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:19:20.053425Z digest=sha256:dc25bfe456175e457539df20e27c969f1245ce4d7808fd800a405f30928ac0a1

Observation 7c45b3d7-ddaa-431b-adf1-0af8a14fca00 · outbound

This paper cites Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models.

Team of One: Cracking Complex Video QA with Model Synergy Video-ChatGPT: Towards Detailed Video Understanding via Large Vision and Language Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T16:19:20.204460Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:19:20.204460Z digest=sha256:1b34f274e22b62dafe5ae824188ba1f6248dfc0a8daf921cc13be0f82cebc7d2

Observation 3f310098-559d-497b-aa17-2aa12af0fde3 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

Team of One: Cracking Complex Video QA with Model Synergy Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T16:19:20.329167Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:19:20.329167Z digest=sha256:f096b6df0d00568ee8c2538f55aed3a91c688b58e2abcb46c9904fdade69731c

Observation df85bc52-675d-40b2-9061-7dfa2486f749 · outbound

This paper cites Chain-of-thought prompting elicits reasoning in large language models.

Team of One: Cracking Complex Video QA with Model Synergy Chain-of-thought prompting elicits reasoning in large language models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-06T16:19:20.430057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:19:20.430057Z digest=sha256:79a91710fb8bd0a6bb4423552b46376706129953f68d5f58956e3f3de498d781

Observation 1b7c2315-9f7b-48f5-9ee5-19ff8e9f5dee · outbound

This paper cites Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding.

Team of One: Cracking Complex Video QA with Model Synergy Video-LLaMA: An Instruction-tuned Audio-Visual Language Model for Video Understanding

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T16:19:20.572832Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T16:19:20.572832Z digest=sha256:583b1923e2602f932ab1c27b9044dd74756cd32784aa8e6c750afd1c16efaee5

Pith citing papers

No inbound Pith citation observations are available.