Pith. sign in

Paper Citation Record · LEDGER

WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 12 inbound Pith citation observations for arXiv:2308.10755.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2308.10755 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 12 of 12 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 12 of 12 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T15:41:15.009671Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T22:17:25.674418Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 11e2c600-e6e8-4053-bd5c-c219331e5efe · inbound

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation cites this paper.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.533004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:c4c80693bfee2b5a813904a7200036149c43d10e8f0c22eb2b45966b4e9eb562

Observation ab9cdb45-d6db-471a-8850-4b4cf2a48598 · inbound

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition cites this paper.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:48.755358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:d38066e076b2f22986c0d2295aabdadd712086584c70e531366838db35ccb720

Observation 81d10b28-61e4-4208-ab89-e0b742d66318 · inbound

InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model cites this paper.

InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:30:27.714888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T05:30:27.667126Z digest=sha256:1e8178c7ebe647d93cb4935a88a2205a665d5e80f1d4b43bc1833e8d5262be1c

Observation 3ea15edf-c5e7-492d-9b82-16c2425f82c5 · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:58.983099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:1b4538dd4eb72f56c7db47bf216f963b929d274c414f5d1dfe74f075d3ec5901

Observation 0c52e197-a7db-49dd-b8d2-3b387d26c3e4 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.734923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:09c4fefbc60deab938f2cfb3a0371e8b5cbffe7bff61df87e64087690df6dac5

Observation e6bfa9cf-62c2-40fc-91be-52d0081facb4 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.347493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:6f12829f2c7a4e07fb8d74417ca54cec3dd909f5e575462bc231da801274c84d

Observation 30c50c22-0043-4515-8728-683cebbc098e · inbound

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding cites this paper.

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:09:23.376048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-11T10:09:21.542356Z digest=sha256:93d6ef063c28aeb00c5ad5df403763cb5672519eda22fecbfeef507984791255

Observation bb25d460-dadb-47c3-a2fc-99d052f46874 · inbound

FuxiMT: Sparsifying Large Language Models for Chinese-Centric Multilingual Machine Translation cites this paper.

FuxiMT: Sparsifying Large Language Models for Chinese-Centric Multilingual Machine Translation WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:15.009671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:41:15.009671Z digest=sha256:e0f25d9c3da9da01ec67b88508015466f193bede6e500958ce242c42cbdf2f8f

Observation 3d17bb68-d896-410f-921f-b62d7d2c94aa · inbound

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models cites this paper.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.722409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.722409Z digest=sha256:5aa9409a46b67e9f866c6743afac335b3f7419835267ed9e25575d515bf1cfef

Observation 9efff52f-2622-409b-b59b-b7b3d2ed25cd · inbound

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale cites this paper.

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 214

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:17:25.676164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-07-02T22:10:59.568675Z digest=sha256:b80d5cc10396a321fdac9ec0b4c8807f90db9aa09075486c5c3177d76e296a3c

Observation cd0150c9-a0aa-4f2e-aeb4-96cedfc54c7f · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:45:47.713578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-06-30T01:16:16.834861Z digest=sha256:8872292225d5a46d1d80962b1fac67b78c02203594ad768c07ea62b005cf4e1f

Observation e9f6c4f1-71e5-415e-9bc6-f65b012bfcca · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.114264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:fa3ac8a23b2cfb6bf7b53885e33e7f05399ee8f87e29c6954da73402bdc0733f