Pith. sign in

Paper Citation Record · LEDGER

WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

As of 15 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 14 inbound Pith citation observations for arXiv:2308.10755.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2308.10755 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 14 of 14 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-15T06:32:42.880941+00:00

measured 14 of 14 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-12T19:39:56.129958Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T22:17:25.674418Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 11e2c600-e6e8-4053-bd5c-c219331e5efe · inbound

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation cites this paper.

InternVid: A Large-scale Video-Text Dataset for Multimodal Understanding and Generation WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-15T06:30:22.533004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-15T06:30:22.431538Z digest=sha256:4cb0758c48bad687903fc3be759f6d15ec21b2338d4124529df795a2a1fc383f

Observation ab9cdb45-d6db-471a-8850-4b4cf2a48598 · inbound

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition cites this paper.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:48.755358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:804a38bccfda7931f2065e456e290d2916b0e4f22733ab888ede7b27de8357e3

Observation 81d10b28-61e4-4208-ab89-e0b742d66318 · inbound

InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model cites this paper.

InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-17T05:30:27.714888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T05:30:27.667126Z digest=sha256:5f02439dd1b3f4b5d2c83f439259a8296a00245d08da31c302b4f51d08a71f25

Observation 3ea15edf-c5e7-492d-9b82-16c2425f82c5 · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:58.983099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:54452d5f5d1eaca890a5e7fd951c5ef99e4b3825db725470459f2b58e27b9df3

Observation 0c52e197-a7db-49dd-b8d2-3b387d26c3e4 · inbound

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output cites this paper.

InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-17T10:46:28.734923Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-17T10:46:28.447347Z digest=sha256:a1eb6711abe595dbe7bfd845e1ec81178318405a6ec1def7b060bdd4e090ee7e

Observation e6bfa9cf-62c2-40fc-91be-52d0081facb4 · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 81

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:23:58.347493Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:14fc95c848c8a7041a6b9f5d94be871294d20dd6b4cb08a692b45233502995bb

Observation 30c50c22-0043-4515-8728-683cebbc098e · inbound

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding cites this paper.

DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:09:23.376048Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-05-11T10:09:21.542356Z digest=sha256:6b20eb79cee53f712c6b80797e25cca56c37ac381b7c3b73e4c1906056ace672

Observation 63e25bf4-9410-445f-8098-cb5f101cb6e8 · inbound

YuLan-Mini: An Open Data-efficient Language Model cites this paper.

YuLan-Mini: An Open Data-efficient Language Model WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-11T05:17:55.610187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T05:17:55.610187Z digest=sha256:5cdcf2ff9928f158b028897568db5a649bfe67da12531db36093af6e47352fad

Observation bb25d460-dadb-47c3-a2fc-99d052f46874 · inbound

FuxiMT: Sparsifying Large Language Models for Chinese-Centric Multilingual Machine Translation cites this paper.

FuxiMT: Sparsifying Large Language Models for Chinese-Centric Multilingual Machine Translation WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T15:41:15.009671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T15:41:15.009671Z digest=sha256:c857c143a4688caf3044cba20c9d4002317d40ac0ebe2ed6aa33204911bedde2

Observation 3d17bb68-d896-410f-921f-b62d7d2c94aa · inbound

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models cites this paper.

Mono-InternVL-1.5: Towards Cheaper and Faster Monolithic Multimodal Large Language Models WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T16:50:04.722409Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T16:50:04.722409Z digest=sha256:0313257a31d7797752aa385bdd857fa72fbdeebf1b5c42005bb33faa2c05bb1d

Observation 9efff52f-2622-409b-b59b-b7b3d2ed25cd · inbound

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale cites this paper.

Ling and Ring 2.6 Technical Report: Efficient and Instant Agentic Intelligence at Trillion-Parameter Scale WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 214

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T22:17:25.676164Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=arxiv_source observed=2026-07-02T22:10:59.568675Z digest=sha256:408fe4f709bfe88b5f19a6cbc045c9706a8f4c6469b90c0f05301aae23fb459a

Observation cd0150c9-a0aa-4f2e-aeb4-96cedfc54c7f · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-07-01T15:45:47.713578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-06-30T01:16:16.834861Z digest=sha256:3e046f93d3a374418710fb85039e553dcecfe3994f4b823d69a9bb1f9a52153b

Observation e9f6c4f1-71e5-415e-9bc6-f65b012bfcca · inbound

DataComp-VLM: Improved Open Datasets for Vision-Language Models cites this paper.

DataComp-VLM: Improved Open Datasets for Vision-Language Models WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-07-02T21:17:24.114264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-15T06:32:42.880941+00:00.

source=pdf_text observed=2026-07-02T21:10:10.548489Z digest=sha256:2fb092e47577eb628c9fcc141d357b70a25e4a548dfbd3c46813cc593b8c66dc

Observation 107d21f6-23b2-4e4e-942b-d1deab169e11 · inbound

Auditing Chinese Web-scale Corpora via Sampled BPE Token Statistics cites this paper.

Auditing Chinese Web-scale Corpora via Sampled BPE Token Statistics WanJuan: A Comprehensive Multimodal Dataset for Advancing English and Chinese Large Models

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-12T19:39:56.129958Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-12T19:39:56.129958Z digest=sha256:03c084c3ceaf82c9b8a8e157524e20ed739cd72179b49aa222904abb3ca8717a