Pith. sign in

Paper Citation Record · LEDGER

AlignBench: Benchmarking Chinese Alignment of Large Language Models

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 23 inbound Pith citation observations for arXiv:2311.18743.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2311.18743 v4

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 23 of 23 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 23 of 23 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T22:17:26.844573Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

1
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation ca841481-233e-462b-bdbd-1c2da779d8ed · inbound

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism cites this paper.

DeepSeek LLM: Scaling Open-Source Language Models with Longtermism AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 147

Resolution
verified exact
arxiv_id, observed 2026-05-11T06:08:05.831970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-11T06:08:05.550346Z digest=sha256:aa3927d2212237520b4592d5dc8d1e5108f7dbbbca9e835b1f8c00010fcbf802

Observation 512de56b-abb6-4b6d-bed1-c997bf629b57 · inbound

InternLM2 Technical Report cites this paper.

InternLM2 Technical Report AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 175

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T11:44:38.323496Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-15T11:44:38.066501Z digest=sha256:9842af6d630cbc2f535e7010893386879a375e7bfeecd681939a14f4279d5173

Observation b7a622bc-38a5-4a9e-81e6-929238bb9bcd · inbound

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model cites this paper.

DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 142

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:36:26.438167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-11T05:36:26.207359Z digest=sha256:6e0bbaa009e55c85f68a87acbf0193ee9018d0c207a43e3b7348f2d80d25b88c

Observation 1b350f83-636f-4bf5-be29-05ba3e868105 · inbound

Yi-Lightning Technical Report cites this paper.

Yi-Lightning Technical Report AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-12T04:34:12.079217Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:34:12.079217Z digest=sha256:b92cff23b0a84b8ea52821415da6a0fcb55805877f4bce590ef94bc4e121f65b

Observation 5ae0d3f8-e3d0-47d1-bc0c-5f2e3e7b2f09 · inbound

Building a Family of Data Augmentation Models for Low-cost LLM Fine-tuning on the Cloud cites this paper.

Building a Family of Data Augmentation Models for Low-cost LLM Fine-tuning on the Cloud AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T21:13:58.128006Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T21:13:58.128006Z digest=sha256:725ce39e242539861b3b515d110d4d6cc165da08be2f20616346f70e08203253

Observation 0c5619da-cf35-4326-9b10-1b2861cc74ba · inbound

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods cites this paper.

LLMs-as-Judges: A Comprehensive Survey on LLM-based Evaluation Methods AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 148

Resolution
verified exact
arxiv_id, observed 2026-05-11T23:08:37.299351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-11T23:08:34.312466Z digest=sha256:d1336540cb90695d65e1021d122a6acc42ee0b70b85ded4e51972eaf6bd0365b

Observation 32fec98b-78cc-4f39-9dd4-3f5ff8efc405 · inbound

Does RLHF Scale? Exploring the Impacts From Data, Model, and Method cites this paper.

Does RLHF Scale? Exploring the Impacts From Data, Model, and Method AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-11T20:10:16.113831Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T20:10:16.113831Z digest=sha256:df3851b3dc4c6f12230b683a4f6fbe7b1142b41f638f4914d9f362a96beeaae9

Observation 66f26733-5eae-44f1-b1d6-7627e89aeeab · inbound

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios cites this paper.

RuleArena: A Benchmark for Rule-Guided Reasoning with LLMs in Real-World Scenarios AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:31.603431Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T17:25:31.603431Z digest=sha256:9017beb6c032b03b5fbc10cb6d9c86d1123e73ce259a54b37f4746068cf841b0

Observation 65a52709-18eb-4c80-acfd-56ec685d144a · inbound

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models cites this paper.

SPaR: Self-Play with Tree-Search Refinement to Improve Instruction-Following in Large Language Models AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-11T14:51:18.619871Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T14:51:18.619871Z digest=sha256:218d290ff7cb71e1b3f71421d21ef6d6f251c89efd3e11f949961e6d471d5cb8

Observation 2effc0c7-cc26-48f1-b2bd-3739b840c7f0 · inbound

Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons cites this paper.

Training an LLM-as-a-Judge Model: Pipeline, Insights, and Practical Lessons AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-09T10:26:06.931661Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-09T10:26:06.931661Z digest=sha256:6a745d5b1c744a38ac0d906adbb79d9fb384513b05fe5175d73a9c3f98758ad6

Observation 0672a3dd-9996-4a33-88b4-ddb0bd77d6ce · inbound

MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining cites this paper.

MiMo: Unlocking the Reasoning Potential of Language Model -- From Pretraining to Posttraining AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-15T22:17:26.844573Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T22:17:26.844573Z digest=sha256:e0aa613d6accbe29eb8d33f1b3d93ece43150bea49a973c2ce08696bd7252063

Observation 1c40b935-71f4-4362-ba07-cd6b9392cee5 · inbound

An Empirical Study of Many-to-Many Summarization with Large Language Models cites this paper.

An Empirical Study of Many-to-Many Summarization with Large Language Models AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-15T20:26:19.784181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T20:26:19.784181Z digest=sha256:6c468b2389b6103220516f53b97bb5cef7d3f280088c0bf0abb047f5e291fbfd

Observation c96a5d48-19fd-4db4-8fa2-81e2d1274caf · inbound

Characterizing Bias: Benchmarking Large Language Models in Simplified versus Traditional Chinese cites this paper.

Characterizing Bias: Benchmarking Large Language Models in Simplified versus Traditional Chinese AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T13:08:35.123660Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:08:35.123660Z digest=sha256:3a26e4e98329550b7a84a402e2b22e19a06ffc2ec5940028cc5478623e9f964e

Observation df62893e-d5f8-4916-a8f8-3c71b13c46a2 · inbound

Beyond the Surface: Measuring Self-Preference in LLM Judgments cites this paper.

Beyond the Surface: Measuring Self-Preference in LLM Judgments AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T11:26:04.837491Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:26:04.837491Z digest=sha256:ab5b5d088341ee5e034a6411d7fc8911572405f327e44f454d38dfa857f8f5bc

Observation 5982dec6-ecb1-4708-abd5-cc5daca3049b · inbound

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations cites this paper.

A Survey of Automatic Evaluation Methods on Text, Visual and Speech Generations AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 226

Resolution
unresolved
no resolver link, observed 2026-08-07T10:17:47.030027Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:17:47.030027Z digest=sha256:1c71f709cccb575f826f7b6907e2aa13514531839974a51508f7483b248395fb

Observation ed44843f-1df3-46f5-af14-4712977d543d · inbound

MinosEval: Distinguishing Factoid and Non-Factoid for Tailored Open-Ended QA Evaluation with LLMs cites this paper.

MinosEval: Distinguishing Factoid and Non-Factoid for Tailored Open-Ended QA Evaluation with LLMs AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-15T19:45:00.374438Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-15T19:45:00.374438Z digest=sha256:20b427318c909b8a8e547d9d8944c84d7311022cb59a121086882e194b0da8e0

Observation 71fbc989-12a5-4017-a47c-e9e65b85ced1 · inbound

Chengyu-Bench: Benchmarking Large Language Models for Chinese Idiom Understanding and Use cites this paper.

Chengyu-Bench: Benchmarking Large Language Models for Chinese Idiom Understanding and Use AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T23:27:19.485869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T23:27:19.485869Z digest=sha256:59437c92b6bb1c599dffb687768f324cba52aeefff8499e89e7b4a606338ed38

Observation e1774273-c23d-4431-9cdd-c9135588f87a · inbound

Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning cites this paper.

Does Math Reasoning Improve General LLM Capabilities? Understanding Transferability of LLM Reasoning AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 227

Resolution
verified exact
arxiv_id, observed 2026-05-19T01:01:09.994861Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-05-19T01:01:09.840919Z digest=sha256:da14da9d66db85cde2b336351a9f6b853af8520bd3aa96c22c2717fbea4a9606

Observation 2a47dbe2-5407-4ed4-b574-e7ec36c316b5 · inbound

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework cites this paper.

Diagnosing Failures in Large Language Models' Answers: Integrating Error Attribution into Evaluation Framework AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T18:23:30.510620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:23:30.510620Z digest=sha256:9d9db50e511b26ee8b5fae2fe01a02889caf54e0ab36411df5453130e7d607af

Observation 2285295d-a65f-46bd-b053-39e90300167b · inbound

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards cites this paper.

CompassJudger-2: Towards Generalist Judge Model via Verifiable Rewards AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-06T18:12:26.444272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T18:12:26.444272Z digest=sha256:7c9aacacb5b2d40e185aa059420422a09da66a96d24e31ef05bf531b193785b1

Observation 6463cf01-e5ad-47ae-b3f7-88993610ba89 · inbound

Technical Report of TeleChat2, TeleChat2.5 and T1 cites this paper.

Technical Report of TeleChat2, TeleChat2.5 and T1 AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-06T14:43:22.372244Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T14:43:22.372244Z digest=sha256:4a35d31edb909860985756e6e522930f45d1ae3cc8c00f407051bc8b12d0a411

Observation 09e90af4-6a09-44df-80c5-dfea3a39aa04 · inbound

AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications cites this paper.

AgentScope 1.0: A Developer-Centric Framework for Building Agentic Applications AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-05T17:26:28.871708Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T17:26:28.871708Z digest=sha256:411a07475046973c40bdd8271b6cfe9616f474c8924746d306a72dacaf1c6466

Observation c534328b-7f8f-4641-9ce8-0ae001490020 · inbound

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges cites this paper.

Chain-of-Models: Cross-Model Auditing for Bias-Robust LLM Judges AlignBench: Benchmarking Chinese Alignment of Large Language Models

Reference 140

Resolution
unresolved
no resolver link, observed 2026-08-03T00:55:29.295048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-03T00:55:29.295048Z digest=sha256:a24a8767a56ca10b40bbc43057bd24b58ad1028f6efddef3de7a6b6ee47f5e72