Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T15:57:59.321417Z
Paper Citation Record · LEDGER
As of 21 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 1 inbound Pith citation observation for arXiv:2411.13775.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T15:57:59.321417Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-07T12:11:13.400147Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-08-07T12:11:14.121736Z
53 of 53 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 9d0efe89-a40c-431e-ac48-ae31a1f1696b · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Is ChatGPT A Good Translator? Yes With GPT-4 As The Engine
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbe487d5-a951-413b-8d0f-309c9bad3bdc · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Document-level machine translation with large language models,
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 63042b39-a464-43cb-b053-733fbb546681 · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels From LLM to NMT: Advancing Low-Resource Machine Translation with Claude
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a624b256-c8a9-4ddf-911b-268ae80c9198 · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Towards making the most of llm for translation quality estimation,
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 94d8dc8e-2b00-467b-be11-163a05b6a8d9 · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Adapting Large Language Models for Document-Level Machine Translation
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 040dfe16-cf74-4f4d-ab1e-ea2f7f7491ce · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels How Good Are GPT Models at Machine Translation? A Comprehensive Evaluation
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2eed2e66-825f-4584-b336-3d59179f807f · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Towards making the most of chatgpt for machine translation,
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4521e6fd-a7a4-4801-a42e-e6e01c5d95b5 · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels What is the best way for ChatGPT to translate poetry?
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 40392461-fb42-4707-a20c-a9e2adb02bbc · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Revisiting cross- lingual summarization: A corpus-based study and a new benchmark with improved annotation
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 36f26620-a2c8-4401-97bc-bd21a8460b2e · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels The price of debiasing automatic metrics in natural language evalaution,
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation a529aaaa-5a33-48ed-aa01-d5f2195f0b05 · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Results of the WMT19 metrics shared task: Segment-level and strong MT systems pose big challenges,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e541fbc6-452f-4ca9-91c9-d3589c0e7d54 · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels BLEURT: Learning robust metrics for text generation,
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e0fc2071-05aa-4c97-8976-220791929180 · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Experts, errors, and context: A large- scale study of human evaluation for machine translation,
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 052a4e68-7256-4c08-a965-1578c85b7550 · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels CLUE: A Chinese Language Understanding Evaluation Benchmark
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f49f40d-d55d-4dbf-afc5-3cf55ccf8a29 · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Benchmarking llms via uncertainty quantification,
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 47059008-98e6-42f5-b91d-5d7d1585a89c · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Measuring massive multitask language understanding,
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3e33ce1-657a-4438-9c61-dd64e4ca600f · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Revisiting out-of-distribution robust- ness in nlp: Benchmark, analysis, and llms evaluations,
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ceef260e-c2a2-4512-9b37-a834514b8025 · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Is chatgpt a good translator? yes with gpt-4 as the engine,
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d4979cb6-76c4-4fed-a6aa-8192173d70ea · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Benchmarking Large Language Models on CFLUE -- A Chinese Financial Language Understanding Evaluation Dataset
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11b321c1-7735-4960-a90c-2d2cd927fb7d · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Evaluating large language models for radiology natural language processing,
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ccdd6389-30e5-4080-8e1b-4334600d34f8 · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels News Summarization and Evaluation in the Era of GPT-3
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 99907c1f-f5f4-40a4-88da-88d13fbc5820 · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Using gpt-4 to provide tiered, formative code feedback,
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 6675944a-8299-45c5-a0c1-173f2a8ab9be · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels A comparison of human and gpt-4 use of probabilistic phrases in a coordination game,
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 09016438-5a3e-4ff5-994e-e6e3a58b9bf9 · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Chatgpt and gpt-4 for professional translators: Exploring the potential of large language models in translation,
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f9a77c46-8bc0-4462-828e-8a16b304319f · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Does GPT-4 surpass human performance in linguistic pragmatics?
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 0b1fa082-81c4-4d5a-b0c8-ce7b3fdf81fc · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Comparative analysis of gpt-4vision, gpt-4 and open source llms in clinical diag- nostic accuracy: A benchmark against human expertise,
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation b7e5562b-f896-46f4-a1e3-6f78f68b3fc9 · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Continuous measurement scales in human evaluation of machine translation,
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 4612f68a-b83f-43df-96f6-7e48143c47cc · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Findings of the 2021 conference on machine translation (wmt21),
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 55aec91a-71f0-4c61-afd5-9525fa23d7d3 · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Findings of the 2022 conference on machine translation (wmt22),
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 3866e6d7-c00c-423e-9089-7539a6e75933 · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Findings of the 2023 conference on machine translation (wmt23): Llms are here but not quite there yet,
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 815e74eb-38f0-45d2-8199-15bda42cf4cd · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Assessing inter-annotator agreement for translation error annotation,
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 796d7dd5-0dc0-47d1-ad58-eb0afc309510 · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Quantitative fine-grained human evaluation of machine translation systems: a case study on english to croatian,
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 25f529b7-f594-427c-a936-a7512d97cad7 · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels COMET: A neural framework for MT evaluation,
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation c8e05a55-ca3c-4249-88c7-e45e900b1e0e · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Results of WMT22 metrics shared task: Stop using BLEU – neural metrics are better and more robust,
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation ca1b0cb6-503b-402e-a1cb-f2a5db47431c · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Results of WMT23 metrics shared task: Metrics might be guilty but references are not innocent,
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 2b1ab9cf-1439-4de9-b7e5-6037483a0aa7 · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Achieving Human Parity on Automatic Chinese to English News Translation
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b39280d5-49d4-4aa1-873c-0b83a13ff5f4 · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels aubli, “What’s the difference between professional human and machine translation? a blind multi-language study on domain-specific MT,
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 82db6a20-27ee-4c56-bab4-4bebadfa6ee3 · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Attaining the unattainable? reassessing claims of human parity in neural machine translation,
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 7e3fd4bb-7fc1-4298-a0ec-e95af82c0bea · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels The suboptimal wmt test sets and its impact on human 12 parity,
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 76b702ca-ae6b-47da-9c1d-6f2725f1e02b · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels On" human parity
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 34d3d709-df9e-465d-9e95-1f426ae8fd33 · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Assessing human-parity in machine translation on the segment level,
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 39592dcf-237c-44b2-b543-72f49a0aed08 · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels SeamlessM4T: Massively Multilingual & Multimodal Machine Translation
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6b340e10-fbc3-401c-ac2e-0901e82f103c · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Calibrate before use: Improving few-shot performance of language models,
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73d907f9-783f-4a66-bd08-7cf2a060edad · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing,
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation f883a8c5-8272-489e-800f-53b8809e77db · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels A paradigm shift in machine translation: Boosting translation performance of large language models,
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e9b5c8e7-679d-4625-b7e4-300baf8296ee · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Comet: A neural framework for mt evaluation,
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation fbe27486-29a6-4d6b-9499-dbdc126b2d23 · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels doccano: Text annotation tool for human,
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation e367a4c2-752a-4c2b-bd0e-c341c8d8d00f · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels A coefficient of agreement for nominal scales,
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 53ce2c5b-636f-40f0-92f0-98f26bd75074 · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Validity in content analysis,
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 05f34536-60c7-496e-90dd-73795c35a7ea · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Not all countries celebrate thanksgiving: On the cultural dominance in large language models,
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.
Observation 23c12a05-1e9e-4938-8d03-15ae2f193816 · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4edbe24d-32fc-4c82-9288-930eef5be2e2 · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels TEaR: Improving LLM-based Machine Translation with Systematic Self-Refinement
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83a043c7-7aea-4856-a638-746d09a6a0a8 · outbound
Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels (Perhaps) Beyond Human Translation: Harnessing Multi-Agent Collaboration for Translating Ultra-Long Literary Texts
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 44a04cea-a814-4844-97a7-edf5517ed247 · inbound
Dual Debiasing for Noisy In-Context Learning for Text Generation Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.