Pith. sign in

Paper Citation Record · LEDGER

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models

As of 10 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 0 inbound Pith citation observations for arXiv:2506.07463.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07463 v1

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:41:14.636513Z

measured 40 of 40 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

40 of 40 outbound references displayed

  • verified exact0
  • verified fuzzy20
  • unresolved20
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 1bb2cb67-650b-4bc9-bbeb-2e6be8b0b00d · outbound

This paper cites Rethinking reflection in pre-training, 2025.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Rethinking reflection in pre-training, 2025

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:15.061372Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:41:14.508198Z digest=sha256:b5e3940d648d6829887cfe022dc63eda30d76e26f57308334448f1a6e78db0d8

Observation ea004250-d416-4924-ade8-c31ed815c17a · outbound

This paper cites COIG-CQIA: Quality is All You Need for Chinese Instruction Fine-tuning.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models COIG-CQIA: Quality is All You Need for Chinese Instruction Fine-tuning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.511955Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.511955Z digest=sha256:b313f17cb4631ee02e147255d7cd895c31768fdb48c62c99793d5d58831d3a3a

Observation 761eba16-ca57-421d-bf7b-81fca97c88fa · outbound

This paper cites Careful selection of knowledge to solve open book question answering.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Careful selection of knowledge to solve open book question answering

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:15.052572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:41:14.515368Z digest=sha256:bb17c9c4ea07268d7864cd065d34ec724b3db259cae3307973fe078774301c5f

Observation 5ffeffb1-d8a8-400a-bdcc-4bd59c1a7ac4 · outbound

This paper cites CCI-Data [Data set].

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models CCI-Data [Data set]

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:15.042870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:41:14.518445Z digest=sha256:b9d73b8514f74ae2b1fec51e4cbe48a7090707f9c163964641095f593f4f71ab

Observation 0db9bfae-f766-41b9-963b-b303ddd72d8b · outbound

This paper cites CCI2-Data [Data set].

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models CCI2-Data [Data set]

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:15.032842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:41:14.521642Z digest=sha256:842403ad5d433a5f410717b025a286d0cbea178edb3efeebc182c3b6b4360f8a

Observation 821bc742-0f46-4d42-ab6c-093a5f3eb79b · outbound

This paper cites WuDaoCorporaText [Data set].

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models WuDaoCorporaText [Data set]

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:15.022976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:41:14.524471Z digest=sha256:08a671f8f9f74b57a0f47b1e74b28ef22505305b539eb9295f0529367706b12d

Observation a77e3deb-fb75-4756-be2c-09d6d7bc5a95 · outbound

This paper cites Piqa: Reasoning about physical commonsense in natural language, 2019.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Piqa: Reasoning about physical commonsense in natural language, 2019

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.527577Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.527577Z digest=sha256:0f2fe4a2101956cdee5f009674a2b2133b9d27f0917b22501e991d7da21a2eda

Observation c76ddd3f-3f49-4c0c-bd78-aa2c8867756a · outbound

This paper cites an unresolved cited work.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:41:15.007881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:41:14.530537Z digest=sha256:42047ac847d3fd73f2968250cc6f96fab8d7db28806a2e1862207238a5081326

Observation eb896623-f4b0-4fca-92b1-f8139b77609c · outbound

This paper cites Data- juicer: A one-stop data processing system for large language models.Companion of the 2024 International Conference on Management of Data, 2023.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Data- juicer: A one-stop data processing system for large language models.Companion of the 2024 International Conference on Management of Data, 2023

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:14.999494Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:41:14.533269Z digest=sha256:93be95e6df5dcb564cdd3aa12013cf701fa16289e9f35b30b38d4de39d3bfe43

Observation 6016aaf2-4ce7-44a1-a002-f5a9d56d1b65 · outbound

This paper cites Chinesewebtext: Large-scale high-quality chinese web text extracted with effective evaluation model, 2023.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Chinesewebtext: Large-scale high-quality chinese web text extracted with effective evaluation model, 2023

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:14.990030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:41:14.536439Z digest=sha256:7168bfd4a66297a6096d6a14608e723b4b91be4d5ec6ea1d189a55bde189c5aa

Observation 1b0e76a6-13b4-4204-baec-7bb14c2f216c · outbound

This paper cites Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Think you have solved question answering? try arc, the ai2 reasoning challenge, 2018

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.539286Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.539286Z digest=sha256:e7fcedfe21522309cd847795027dc7571ff70213145b2b198e1cb544d49cbd52

Observation 2e138668-d626-46a3-9a64-3ba295c84ae2 · outbound

This paper cites Unsupervised Cross-lingual Representation Learning at Scale.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Unsupervised Cross-lingual Representation Learning at Scale

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.542023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.542023Z digest=sha256:e29bc10a9971a747ad028050330fc4b6ecaa37f6bd51bced3dfd848597708451

Observation 0d71b92f-79bb-419f-ae51-fcffb048db63 · outbound

This paper cites Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:14.975953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:41:14.545952Z digest=sha256:cd6027348038d14cfa0e3fa22c88c0abee09677a362d0458cff55c83bb7d4024

Observation 61725105-411f-4f92-a3ee-38d037407545 · outbound

This paper cites Lighteval: A lightweight framework for llm evaluation, 2023.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Lighteval: A lightweight framework for llm evaluation, 2023

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:14.965967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:41:14.549699Z digest=sha256:86a3042f7fae39f0e0abda67878daf6cafb26165cf559c499e953053e131d235

Observation 491272bb-96a7-46ad-990d-94452d443c50 · outbound

This paper cites Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Cognitive Behaviors that Enable Self-Improving Reasoners, or, Four Habits of Highly Effective STaRs

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.552614Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.552614Z digest=sha256:33a6e9d1eacd1eaa76103220f4d511ee5986fded5ef2313db475b6887ddbe99c

Observation eedb7d1d-8808-4cc4-afce-714e6be0fb58 · outbound

This paper cites The Pile: An 800GB Dataset of Diverse Text for Language Modeling.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models The Pile: An 800GB Dataset of Diverse Text for Language Modeling

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.555798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.555798Z digest=sha256:2a11a659c4eaa36567932b5052a949e220567b4d6b527ceeb02a2fa5338698c8

Observation 568e34bd-b5b9-4742-a8e1-d580cb3dea04 · outbound

This paper cites Wanjuan: A comprehensive multimodal dataset for advancing english and chinese large models, 2023.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Wanjuan: A comprehensive multimodal dataset for advancing english and chinese large models, 2023

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:14.956085Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:41:14.559100Z digest=sha256:60a7464a666aae75d173c27c49528f5560558f1c88e72f34f5d7881b565b8c52

Observation 124c3a16-07df-4d3c-9311-ab3121e7cefe · outbound

This paper cites Measuring massive multitask language understanding, 2021.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Measuring massive multitask language understanding, 2021

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.562041Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.562041Z digest=sha256:40031c276c32b97b031564fb10c90ad9aa16752695d3aa3cb4745bc5b0e339bc

Observation 21ec9a21-6c4c-414c-a150-2e4b90653a8e · outbound

This paper cites C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models C-eval: A multi-level multi-discipline chinese evaluation suite for foundation models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.565274Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.565274Z digest=sha256:98960672e106c22e3fdd2e17a79bb97f83778df1c935074a32f4081f209c39eb

Observation 8a95f47b-695d-4749-9c5a-66e148574b86 · outbound

This paper cites Bag of Tricks for Efficient Text Classification.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Bag of Tricks for Efficient Text Classification

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.568240Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.568240Z digest=sha256:115aa7348e674df67cf6df2aa7ff939dda4d2a2f8756d530587a7ac6d4a5845e

Observation 3d10ba18-7116-4154-add2-7ca26760cf8e · outbound

This paper cites Deduplicating training data makes language models better, 2022.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Deduplicating training data makes language models better, 2022

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.571434Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.571434Z digest=sha256:b7ddf2cc3b506557f2cc8aff03750e95d2971b6d86a154f03990729582a9a707

Observation adeb30bd-1aff-4068-a773-356b85376e54 · outbound

This paper cites Levesque, Ernest Davis, and Leora Morgenstern.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Levesque, Ernest Davis, and Leora Morgenstern

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:14.930480Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:41:14.574169Z digest=sha256:f32122f3c84d0e8eb2f90ec40c703d5224caf6cd0c9e42fba25d3dc815f866ac

Observation f4787769-6224-46fb-bbf7-6ecd8ae82858 · outbound

This paper cites Cmmlu: Measuring massive multitask language understanding in chinese, 2024.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Cmmlu: Measuring massive multitask language understanding in chinese, 2024

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:14.921385Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:41:14.577122Z digest=sha256:35dc2d0378ca46043a6791a195bce8f91773299af74c971668d1bec7ba06de7e

Observation eecd1c91-24f1-4d04-af7c-e816aa38a308 · outbound

This paper cites Datacomp-lm: In search of the next generation of training sets for language models, 2024.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Datacomp-lm: In search of the next generation of training sets for language models, 2024

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:14.912583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:41:14.579822Z digest=sha256:bd025afb084a2d83e9fd862873d88ddac6bd797f45a463e995cc6a27237f1945

Observation 5a40faad-a015-4d50-b153-f5e2fc448e83 · outbound

This paper cites Openhermes 2.5-zh: A partial chinese translation of openhermes-2.5, 2024.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Openhermes 2.5-zh: A partial chinese translation of openhermes-2.5, 2024

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:14.903381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:41:14.582561Z digest=sha256:6cb23ad604361e42f359775c56fbba846d21b83a32561fd5f07d508946aab2fb

Observation f80f6c1e-32d9-412a-82e4-9f296db90657 · outbound

This paper cites Fineweb2: A sparkling update with 1000s of languages, December 2024.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Fineweb2: A sparkling update with 1000s of languages, December 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:14.894325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:41:14.585479Z digest=sha256:b732f00d4b719187a94bdabae323f3008871ad6f74e780c99775fb3b36f0d175

Observation 170dd499-10b3-44f6-a796-c6a1bb4bd353 · outbound

This paper cites The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models The RefinedWeb Dataset for Falcon LLM: Outperforming Curated Corpora with Web Data, and Web Data Only

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.588446Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.588446Z digest=sha256:5d7beb205f4831bd17c80c1f3c99462534e34b6c19856bfd856063e87cbebac4

Observation ec09a55b-c1db-4934-9655-83e2df2b2ec1 · outbound

This paper cites Deduplicate Text Datasets.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Deduplicate Text Datasets

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:14.884898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:41:14.592064Z digest=sha256:1932b6833b9061d3edf2ced38b1faacfc84247dddba5fa20f10d758ea46beb18

Observation 0bee756a-db61-495c-b69a-e91352b8ca2b · outbound

This paper cites Socialiqa: Com- monsense reasoning about social interactions, 2019.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Socialiqa: Com- monsense reasoning about social interactions, 2019

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.596279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.596279Z digest=sha256:9007cee1cfad03130a02f890dec7377a92ee7a0a13dd321d9d0141f85ec37400

Observation edbe1d46-58dd-4497-939c-f4d2edd9d3f6 · outbound

This paper cites Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Dolma: an Open Corpus of Three Trillion Tokens for Language Model Pretraining Research

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.600683Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.600683Z digest=sha256:2989b7a5ee1a5d3ac798ee84492c597bcd72efced2e376eb09c1b7fc80a23c5b

Observation ae53137f-0530-4d27-a82b-da4d9d9ba88d · outbound

This paper cites Nemotron-cc: Transforming common crawl into a refined long-horizon pretraining dataset, 2024.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Nemotron-cc: Transforming common crawl into a refined long-horizon pretraining dataset, 2024

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:14.788616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:41:14.604175Z digest=sha256:cd3bd2f44ca2432179011617170b067a106ee678b7e69c90dc407b47da984667

Observation d4c02cd1-5746-45b0-9ac3-8b0cc7ddabc6 · outbound

This paper cites The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models The FineWeb Datasets: Decanting the Web for the Finest Text Data at Scale

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.610652Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.610652Z digest=sha256:110e933a5ebccc964905f5baafcf27286570d4b2c58604d4b7e55ecd3bf3b5b5

Observation 5ff8f4b7-0c07-4cd2-83bb-0dc2b5b4905b · outbound

This paper cites Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Nemotron-CC: Transforming Common Crawl into a Refined Long-Horizon Pretraining Dataset

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.613803Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.613803Z digest=sha256:61d776d7666cb087910a022d7264521833168340babf1ee2096b4d6404a57573

Observation 62abf318-fd8f-4daf-bd0b-cb4e3abf1b12 · outbound

This paper cites Qwen2.5: A party of foundation models, September 2024.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Qwen2.5: A party of foundation models, September 2024

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.616833Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.616833Z digest=sha256:30017cc215be5c0a2c3c222d07a3fa3ac6e0a9323e1f9f770488d5222f1aa18f

Observation bfb42c5d-8ef1-4ffb-a744-bcb6f04f7312 · outbound

This paper cites Cci3.0-hq: a large-scale chinese dataset of high quality designed for pre-training large language models, 2024.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Cci3.0-hq: a large-scale chinese dataset of high quality designed for pre-training large language models, 2024

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:14.773012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:41:14.619787Z digest=sha256:10f87ef1e163d7dacea930586c3e95836e86dc0e6938db879b3893972d34cae3

Observation ee8c4ba0-2457-4119-97b4-aa53e8736227 · outbound

This paper cites Qwen2 Technical Report.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Qwen2 Technical Report

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.622934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.622934Z digest=sha256:a43c819641a3c287092deb209f615a1baef78d1bf8e689cfbd26eedba33d7078

Observation fec69a07-0e43-46e1-91ea-2e38f3306176 · outbound

This paper cites Opencsg chinese corpus: A series of high-quality chinese datasets for llm training, 2025.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Opencsg chinese corpus: A series of high-quality chinese datasets for llm training, 2025

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:14.763329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:41:14.626461Z digest=sha256:94ca67c4d90bf1dc3a9bae724d7bfabaf62fa4e0409b0a0bda371436ca9f4902

Observation 094a470d-2194-4559-bdd4-ab9441d562f3 · outbound

This paper cites Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Does Reinforcement Learning Really Incentivize Reasoning Capacity in LLMs Beyond the Base Model?

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:41:14.629752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:41:14.629752Z digest=sha256:a9d2e6c56d3170f52a08ccdcc69bc9cd757084cabd5024f323ec8382c690d652

Observation 730a1797-7f91-41a9-a855-48e534836ef8 · outbound

This paper cites Games" and.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Games" and

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:41:14.753261Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:41:14.633110Z digest=sha256:5f89d2f33aed70f225c91cc5d0a3cb0ae801ceddac1c3af50107f561fff9d06c

Observation 02d8a56d-45ad-4bdd-ac3c-98ac11d90f2d · outbound

This paper cites an unresolved cited work.

CCI4.0: A Bilingual Pretraining Dataset for Enhancing Reasoning in Large Language Models Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:41:14.743827Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-08-07T05:41:14.636513Z digest=sha256:0f5730b4cad4cdc4dad74cce13cd47e831754badcce12a15ef8264c03c076f77

Pith citing papers

No inbound Pith citation observations are available.