Pith. sign in

Paper Citation Record · LEDGER

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels

As of 21 August 2026, this Paper Citation Record lists 53 of 53 outbound references and 1 inbound Pith citation observation for arXiv:2411.13775.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2411.13775 v1

Coverage vector

measured 53 of 53 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-12T15:57:59.321417Z

measured 54 of 54 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-20T06:33:59.587034+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:11:13.400147Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-07T12:11:14.121736Z

Reference resolution

53 of 53 outbound references displayed

  • verified exact2
  • verified fuzzy34
  • unresolved17
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 9d0efe89-a40c-431e-ac48-ae31a1f1696b · outbound

This paper cites Is ChatGPT A Good Translator? Yes With GPT-4 As The Engine.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Is ChatGPT A Good Translator? Yes With GPT-4 As The Engine

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-12T15:57:59.020774Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:57:59.020774Z digest=sha256:cfda23feb06c0464532b198ac820e7669818149ab15b29df6d6bb695b875ad80

Observation bbe487d5-a951-413b-8d0f-309c9bad3bdc · outbound

This paper cites Document-level machine translation with large language models,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Document-level machine translation with large language models,

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:00.469453Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.027455Z digest=sha256:e402a8d4de647268f195a962288169f7ac51c4fa33da6de3c19b26fe2d6e19af

Observation 63042b39-a464-43cb-b053-733fbb546681 · outbound

This paper cites From LLM to NMT: Advancing Low-Resource Machine Translation with Claude.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels From LLM to NMT: Advancing Low-Resource Machine Translation with Claude

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-12T15:57:59.033093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:57:59.033093Z digest=sha256:10aa505af07900acac4cdafc4b533a77b78f2ea4f6143c47320b9a7ac877bc19

Observation a624b256-c8a9-4ddf-911b-268ae80c9198 · outbound

This paper cites Towards making the most of llm for translation quality estimation,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Towards making the most of llm for translation quality estimation,

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:00.453133Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.042943Z digest=sha256:aa565e92c5061f733277c3c3bd8f9acd4da2d54ab36f8398bd2cf4aa6714c3bb

Observation 94d8dc8e-2b00-467b-be11-163a05b6a8d9 · outbound

This paper cites Adapting Large Language Models for Document-Level Machine Translation.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Adapting Large Language Models for Document-Level Machine Translation

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-12T15:57:59.049681Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:57:59.049681Z digest=sha256:67f0f78fc6dfd7081421f1063c511e4d0ed292477d5283ddc173c35ed476a736

Observation 040dfe16-cf74-4f4d-ab1e-ea2f7f7491ce · outbound

This paper cites How Good Are GPT Models at Machine Translation? A Comprehensive Evaluation.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels How Good Are GPT Models at Machine Translation? A Comprehensive Evaluation

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-12T15:57:59.057405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:57:59.057405Z digest=sha256:80b71d5cdf0abec796b1b7882ccb56b016386b9f147eb0d43c60af5691aae60d

Observation 2eed2e66-825f-4584-b336-3d59179f807f · outbound

This paper cites Towards making the most of chatgpt for machine translation,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Towards making the most of chatgpt for machine translation,

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:00.436179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.064112Z digest=sha256:7ba7d58f431cae9825070bf6cf4f278770ed19fa615830356b9aafb31fabb595

Observation 4521e6fd-a7a4-4801-a42e-e6e01c5d95b5 · outbound

This paper cites What is the best way for ChatGPT to translate poetry?.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels What is the best way for ChatGPT to translate poetry?

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:00.419264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.075239Z digest=sha256:4fcde840893648af955aced8b2f1f379753a9ce8b9b9b5a006c976a89fed8ec5

Observation 40392461-fb42-4707-a20c-a9e2adb02bbc · outbound

This paper cites Revisiting cross- lingual summarization: A corpus-based study and a new benchmark with improved annotation.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Revisiting cross- lingual summarization: A corpus-based study and a new benchmark with improved annotation

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:00.404078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.080774Z digest=sha256:8a767bca2ff2b4a0cab3ce7ec9827c9d439699b45918c3efb3b154db6c82b3ab

Observation 36f26620-a2c8-4401-97bc-bd21a8460b2e · outbound

This paper cites The price of debiasing automatic metrics in natural language evalaution,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels The price of debiasing automatic metrics in natural language evalaution,

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:00.387011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.086657Z digest=sha256:07a6f68c0d144b3e0a101e24794752480fdacae6ccd628752a1f43156ee41ea3

Observation a529aaaa-5a33-48ed-aa01-d5f2195f0b05 · outbound

This paper cites Results of the WMT19 metrics shared task: Segment-level and strong MT systems pose big challenges,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Results of the WMT19 metrics shared task: Segment-level and strong MT systems pose big challenges,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:00.367418Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.091795Z digest=sha256:35c8f59ff63293b637d249023124490055d2bf554243ea7d452f9640d3db69db

Observation e541fbc6-452f-4ca9-91c9-d3589c0e7d54 · outbound

This paper cites BLEURT: Learning robust metrics for text generation,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels BLEURT: Learning robust metrics for text generation,

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:00.327588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.102593Z digest=sha256:f121e28b4c764aac0363ae119a6f89d5e6cc8926b364f25b935290aa81bc6709

Observation e0fc2071-05aa-4c97-8976-220791929180 · outbound

This paper cites Experts, errors, and context: A large- scale study of human evaluation for machine translation,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Experts, errors, and context: A large- scale study of human evaluation for machine translation,

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:00.309794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.107173Z digest=sha256:8c3b113a9b583d55dff92ff019cd614718e56b5183fa43319516a73a2bc585f8

Observation 052a4e68-7256-4c08-a965-1578c85b7550 · outbound

This paper cites CLUE: A Chinese Language Understanding Evaluation Benchmark.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels CLUE: A Chinese Language Understanding Evaluation Benchmark

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-12T15:57:59.112412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:57:59.112412Z digest=sha256:3b314e12cf4052c30811aa6ef1a695746e75cfd6673aea0166bc0b59e803afe3

Observation 7f49f40d-d55d-4dbf-afc5-3cf55ccf8a29 · outbound

This paper cites Benchmarking llms via uncertainty quantification,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Benchmarking llms via uncertainty quantification,

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:00.162940Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.117860Z digest=sha256:79a80af45b84cfe9bd98da8cd39800f6a08ac274b94f74745e51ddb9da1be74f

Observation 47059008-98e6-42f5-b91d-5d7d1585a89c · outbound

This paper cites Measuring massive multitask language understanding,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Measuring massive multitask language understanding,

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-12T15:57:59.125050Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:57:59.125050Z digest=sha256:2d4646a4404e66d31b227f29274b0c44a6bf2f99a61d8590ba4389be1dfcbfa0

Observation b3e33ce1-657a-4438-9c61-dd64e4ca600f · outbound

This paper cites Revisiting out-of-distribution robust- ness in nlp: Benchmark, analysis, and llms evaluations,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Revisiting out-of-distribution robust- ness in nlp: Benchmark, analysis, and llms evaluations,

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:00.135736Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.131265Z digest=sha256:68bd2a3e9c82978ec3d4cd950b40265f9b539004daedef79146ce55d0ee1d080

Observation ceef260e-c2a2-4512-9b37-a834514b8025 · outbound

This paper cites Is chatgpt a good translator? yes with gpt-4 as the engine,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Is chatgpt a good translator? yes with gpt-4 as the engine,

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-12T15:57:59.139603Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:57:59.139603Z digest=sha256:17eedfb04b082498999fd343f9d367e61b003fb616a78fdf0589377e801bf112

Observation d4979cb6-76c4-4fed-a6aa-8192173d70ea · outbound

This paper cites Benchmarking Large Language Models on CFLUE -- A Chinese Financial Language Understanding Evaluation Dataset.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Benchmarking Large Language Models on CFLUE -- A Chinese Financial Language Understanding Evaluation Dataset

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-12T15:57:59.145357Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:57:59.145357Z digest=sha256:979176e64491af63ed22a9b404831708e6acd14c3750edaec8262d8b3d297bc2

Observation 11b321c1-7735-4960-a90c-2d2cd927fb7d · outbound

This paper cites Evaluating large language models for radiology natural language processing,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Evaluating large language models for radiology natural language processing,

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-12T15:57:59.150231Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:57:59.150231Z digest=sha256:37b839d6a88b7f0dda9b63001cb98e683e1aa86b18415fbad68aab4b9da233ef

Observation ccdd6389-30e5-4080-8e1b-4334600d34f8 · outbound

This paper cites News Summarization and Evaluation in the Era of GPT-3.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels News Summarization and Evaluation in the Era of GPT-3

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-12T15:57:59.155212Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:57:59.155212Z digest=sha256:ee7bfd57b4565efd15fa63522290ca61b18627e2d70024a6c0ca9acf6d8aa666

Observation 99907c1f-f5f4-40a4-88da-88d13fbc5820 · outbound

This paper cites Using gpt-4 to provide tiered, formative code feedback,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Using gpt-4 to provide tiered, formative code feedback,

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:00.107347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.160248Z digest=sha256:649db50d949f0c36ccc1c5d9f7c636b30b00b81c8c326b41e124e2f3ef227127

Observation 6675944a-8299-45c5-a0c1-173f2a8ab9be · outbound

This paper cites A comparison of human and gpt-4 use of probabilistic phrases in a coordination game,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels A comparison of human and gpt-4 use of probabilistic phrases in a coordination game,

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:00.091070Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.165898Z digest=sha256:0a8a619bf35ec830baaa73a4c4d0c4b8df0802fa57f9e6ee6c9c521b3a49a6cf

Observation 09016438-5a3e-4ff5-994e-e6e3a58b9bf9 · outbound

This paper cites Chatgpt and gpt-4 for professional translators: Exploring the potential of large language models in translation,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Chatgpt and gpt-4 for professional translators: Exploring the potential of large language models in translation,

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:00.071481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.172079Z digest=sha256:04d4c44cc7ced39f5116db260bc3bc870c176120bb1e0935885d922874a5dff6

Observation f9a77c46-8bc0-4462-828e-8a16b304319f · outbound

This paper cites Does GPT-4 surpass human performance in linguistic pragmatics?.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Does GPT-4 surpass human performance in linguistic pragmatics?

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-08-12T15:57:59.511027Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.177099Z digest=sha256:d794b9934701f78df73a335cca53daaf58721981e8e732fee5e02a79aa2c2cf5

Observation 0b1fa082-81c4-4d5a-b0c8-ce7b3fdf81fc · outbound

This paper cites Comparative analysis of gpt-4vision, gpt-4 and open source llms in clinical diag- nostic accuracy: A benchmark against human expertise,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Comparative analysis of gpt-4vision, gpt-4 and open source llms in clinical diag- nostic accuracy: A benchmark against human expertise,

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:00.047439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.182876Z digest=sha256:39dedca158e892ed1200b1bcc62ea0650e3463b25bf111691a252a0f39af8a6c

Observation b7e5562b-f896-46f4-a1e3-6f78f68b3fc9 · outbound

This paper cites Continuous measurement scales in human evaluation of machine translation,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Continuous measurement scales in human evaluation of machine translation,

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:00.023366Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.187765Z digest=sha256:705bb26e94b855d87d11523829865dfb64f2bc3b8ab3173f65c01fb431b259e7

Observation 4612f68a-b83f-43df-96f6-7e48143c47cc · outbound

This paper cites Findings of the 2021 conference on machine translation (wmt21),.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Findings of the 2021 conference on machine translation (wmt21),

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:57:59.998572Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.193129Z digest=sha256:b971050b339b25a3b7500c831d47a362fbd5cae95e6decbc69eac848008ba103

Observation 55aec91a-71f0-4c61-afd5-9525fa23d7d3 · outbound

This paper cites Findings of the 2022 conference on machine translation (wmt22),.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Findings of the 2022 conference on machine translation (wmt22),

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:57:59.981751Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.197791Z digest=sha256:dfc7e6d7059b496e29ea97bb4271e937123d153fac1ed4a52712b9613ba55efe

Observation 3866e6d7-c00c-423e-9089-7539a6e75933 · outbound

This paper cites Findings of the 2023 conference on machine translation (wmt23): Llms are here but not quite there yet,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Findings of the 2023 conference on machine translation (wmt23): Llms are here but not quite there yet,

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:57:59.963860Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.203542Z digest=sha256:bf361251daeed295bf0321df8647f1fde35381715294f35bb56f1c45621da2c2

Observation 815e74eb-38f0-45d2-8199-15bda42cf4cd · outbound

This paper cites Assessing inter-annotator agreement for translation error annotation,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Assessing inter-annotator agreement for translation error annotation,

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:57:59.947859Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.208256Z digest=sha256:6469b24f8f3af076a82d97bead41cf68bbf600340b0f04c66dabcbb1e17655a5

Observation 796d7dd5-0dc0-47d1-ad58-eb0afc309510 · outbound

This paper cites Quantitative fine-grained human evaluation of machine translation systems: a case study on english to croatian,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Quantitative fine-grained human evaluation of machine translation systems: a case study on english to croatian,

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:57:59.932115Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.212977Z digest=sha256:6c5b83521dab03a2ba934af4fe25bbd0fb70407fd399ecb2f9201a9ba664b189

Observation 25f529b7-f594-427c-a936-a7512d97cad7 · outbound

This paper cites COMET: A neural framework for MT evaluation,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels COMET: A neural framework for MT evaluation,

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:58:00.345427Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.217974Z digest=sha256:828b01ec803624d71991802897a4c582fdc009253bd87fae4b498eb7a4c24b92

Observation c8e05a55-ca3c-4249-88c7-e45e900b1e0e · outbound

This paper cites Results of WMT22 metrics shared task: Stop using BLEU – neural metrics are better and more robust,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Results of WMT22 metrics shared task: Stop using BLEU – neural metrics are better and more robust,

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:57:59.915721Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.222913Z digest=sha256:f7a7d59666c626c46f41ed18678c3fd455e9308d6ff661376d842a7d7199402a

Observation ca1b0cb6-503b-402e-a1cb-f2a5db47431c · outbound

This paper cites Results of WMT23 metrics shared task: Metrics might be guilty but references are not innocent,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Results of WMT23 metrics shared task: Metrics might be guilty but references are not innocent,

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:57:59.898699Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.228477Z digest=sha256:37172b80ee5a0f715f408e6516feb6063d15a1aaf22daa1b809b8e57eb03a954

Observation 2b1ab9cf-1439-4de9-b7e5-6037483a0aa7 · outbound

This paper cites Achieving Human Parity on Automatic Chinese to English News Translation.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Achieving Human Parity on Automatic Chinese to English News Translation

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-12T15:57:59.233007Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:57:59.233007Z digest=sha256:4a6eb710e4cdb1856c90eb5347aabfb62f473fea20474a3c99bd8ba051c77f7b

Observation b39280d5-49d4-4aa1-873c-0b83a13ff5f4 · outbound

This paper cites aubli, “What’s the difference between professional human and machine translation? a blind multi-language study on domain-specific MT,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels aubli, “What’s the difference between professional human and machine translation? a blind multi-language study on domain-specific MT,

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:57:59.883055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.238211Z digest=sha256:9d1244b42b65cec75d070f3947fda7419ac48930329eff9599d3e20a8487ec27

Observation 82db6a20-27ee-4c56-bab4-4bebadfa6ee3 · outbound

This paper cites Attaining the unattainable? reassessing claims of human parity in neural machine translation,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Attaining the unattainable? reassessing claims of human parity in neural machine translation,

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:57:59.866785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.243424Z digest=sha256:8fda901a2fcf64dbd5e069ca0bf76821b127041d8620ffb856e4e0c8a06b5ceb

Observation 7e3fd4bb-7fc1-4298-a0ec-e95af82c0bea · outbound

This paper cites The suboptimal wmt test sets and its impact on human 12 parity,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels The suboptimal wmt test sets and its impact on human 12 parity,

Reference 41

Resolution
verified exact
raw_fallback, observed 2026-08-12T15:57:59.472395Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.248907Z digest=sha256:e34da5a358ecd821e9cbd7ac6d4ae1fbb2c8618e521d99f7de7afe7b91d1d9c3

Observation 76b702ca-ae6b-47da-9c1d-6f2725f1e02b · outbound

This paper cites On" human parity.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels On" human parity

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:57:59.851342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.253359Z digest=sha256:26ca2ba62fb2eccd2072e614c6d35f2db1a4bf35b8a3fc866aafa160c34a3089

Observation 34d3d709-df9e-465d-9e95-1f426ae8fd33 · outbound

This paper cites Assessing human-parity in machine translation on the segment level,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Assessing human-parity in machine translation on the segment level,

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:57:59.835335Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.257760Z digest=sha256:3de4ba1f7f11d03879aa9b26568ba64b9a9d0b81e298820786e7ac956e4e25d0

Observation 39592dcf-237c-44b2-b543-72f49a0aed08 · outbound

This paper cites SeamlessM4T: Massively Multilingual & Multimodal Machine Translation.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels SeamlessM4T: Massively Multilingual & Multimodal Machine Translation

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-12T15:57:59.262609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:57:59.262609Z digest=sha256:a4de511025be8b09ce2687aa9ce3dc8dc00f707e7faa34fde65d1c059012f117

Observation 6b340e10-fbc3-401c-ac2e-0901e82f103c · outbound

This paper cites Calibrate before use: Improving few-shot performance of language models,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Calibrate before use: Improving few-shot performance of language models,

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-12T15:57:59.267659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:57:59.267659Z digest=sha256:befaaefac6921ade4a7a098ad47fdf83e75b4573c33da7686513aa3a302e4a00

Observation 73d907f9-783f-4a66-bd08-7cf2a060edad · outbound

This paper cites Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Pre-train, prompt, and predict: A systematic survey of prompting methods in natural language processing,

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:57:59.808512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.272793Z digest=sha256:22a32dae8c223c39c7a74de2ecfd7f2a8015a6d2509a8a7f12bc939ebf4c1212

Observation f883a8c5-8272-489e-800f-53b8809e77db · outbound

This paper cites A paradigm shift in machine translation: Boosting translation performance of large language models,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels A paradigm shift in machine translation: Boosting translation performance of large language models,

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:57:59.792365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.279195Z digest=sha256:c04168172a8078986e2ecd5302f01de48abf1eac883f6a63c8e6a9cbb498bf88

Observation e9b5c8e7-679d-4625-b7e4-300baf8296ee · outbound

This paper cites Comet: A neural framework for mt evaluation,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Comet: A neural framework for mt evaluation,

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:57:59.775722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.284784Z digest=sha256:70602080040003fd717e9a6846c227867dfd22cb08439098f2b930ae82ac7b07

Observation fbe27486-29a6-4d6b-9499-dbdc126b2d23 · outbound

This paper cites doccano: Text annotation tool for human,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels doccano: Text annotation tool for human,

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:57:59.759733Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.291396Z digest=sha256:28e2cda1724e9798603d50d87233d6acef020d32f3bc71e77c5c953596ec6002

Observation e367a4c2-752a-4c2b-bd0e-c341c8d8d00f · outbound

This paper cites A coefficient of agreement for nominal scales,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels A coefficient of agreement for nominal scales,

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-12T15:57:59.296481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:57:59.296481Z digest=sha256:42a9ae5d285e9185917f752ee6508f9f2f9e4df966d4f7af97a6bd7ea9f22278

Observation 53ce2c5b-636f-40f0-92f0-98f26bd75074 · outbound

This paper cites Validity in content analysis,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Validity in content analysis,

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:57:59.733467Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.301131Z digest=sha256:d039bb20cfe49ba8db190cd8fe786459d4260b8f90cb712838f427a721aaa33a

Observation 05f34536-60c7-496e-90dd-73795c35a7ea · outbound

This paper cites Not all countries celebrate thanksgiving: On the cultural dominance in large language models,.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Not all countries celebrate thanksgiving: On the cultural dominance in large language models,

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-12T15:57:59.716804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=pdf_text observed=2026-08-12T15:57:59.306255Z digest=sha256:b16f8735080d06c6215436ef877af82febeac9b356b32c6e725f84cc5a13784b

Observation 23c12a05-1e9e-4938-8d03-15ae2f193816 · outbound

This paper cites Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels Siren's Song in the AI Ocean: A Survey on Hallucination in Large Language Models

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-12T15:57:59.311066Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:57:59.311066Z digest=sha256:f7b7deb4a20d0b62aad17ae101f230e443f464cf9112501d79c32c07b66e5691

Observation 4edbe24d-32fc-4c82-9288-930eef5be2e2 · outbound

This paper cites TEaR: Improving LLM-based Machine Translation with Systematic Self-Refinement.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels TEaR: Improving LLM-based Machine Translation with Systematic Self-Refinement

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-12T15:57:59.316216Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:57:59.316216Z digest=sha256:c8b9b68cccef596b6c4a65b6925203ec7e28a4eb48eccb719b70ad9038dccd6a

Observation 83a043c7-7aea-4856-a638-746d09a6a0a8 · outbound

This paper cites (Perhaps) Beyond Human Translation: Harnessing Multi-Agent Collaboration for Translating Ultra-Long Literary Texts.

Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels (Perhaps) Beyond Human Translation: Harnessing Multi-Agent Collaboration for Translating Ultra-Long Literary Texts

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-12T15:57:59.321417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T15:57:59.321417Z digest=sha256:60dff5e1e20896ff1ed0ab73dd757f9e6231f78f62b21d66693af398bc8d8729

Pith citing papers

Observation 44a04cea-a814-4844-97a7-edf5517ed247 · inbound

Dual Debiasing for Noisy In-Context Learning for Text Generation cites this paper.

Dual Debiasing for Noisy In-Context Learning for Text Generation Benchmarking GPT-4 against Human Translators: A Comprehensive Evaluation Across Languages, Domains, and Expertise Levels

Reference 25

Resolution
verified exact
local_arxiv, observed 2026-08-07T12:11:14.152040Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-20T06:33:59.587034+00:00.

source=arxiv_source observed=2026-08-07T12:11:13.400147Z digest=sha256:95db8513456f905d403f6c0e22695c9fd835a00bf29e1c0b022009bc610ad494