Pith. sign in

Paper Citation Record · LEDGER

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation

As of 8 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 1 inbound Pith citation observation for arXiv:2506.05070.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05070 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:31:58.854906Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T13:58:53.430492Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T14:00:28.739440Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact3
  • verified fuzzy2
  • unresolved59
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4aef4f8a-4efb-448e-9302-1fb752988ee2 · outbound

This paper cites URL: " 'urlintro :=.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation URL: " 'urlintro :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.341810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.341810Z digest=sha256:8c4b19e4f8d2e5ede6f59227c886ef97de6158070ce1f725f04450859969ec4e

Observation 43a10ab0-bd84-49f3-ab43-6a4e4d51e9ae · outbound

This paper cites write newline.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.348130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.348130Z digest=sha256:2971f8476091ba4edf2d5fdd093896b368622e8c9c17293a8e1e2be62c01ca27

Observation d3c297ee-c723-4649-ad59-260d500d7ba2 · outbound

This paper cites GPT-4 Technical Report.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.352960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.352960Z digest=sha256:529deaffee993200f058616147c61c7218747a36cbb18fa93b1b2ea328fdf8d1

Observation 814d1cc0-7757-44a5-8247-f66e06bdb408 · outbound

This paper cites Scalable Ensembling For Mitigating Reward Overoptimisation.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Scalable Ensembling For Mitigating Reward Overoptimisation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.359316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.359316Z digest=sha256:af61a7f5216cfafdd95c3104720d244a2d9f493378264badd7342f0a43105b32

Observation c495aa36-a0a6-4e35-897b-b39e32dacadb · outbound

This paper cites Tower: An Open Multilingual Large Language Model for Translation-Related Tasks.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Tower: An Open Multilingual Large Language Model for Translation-Related Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.366539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.366539Z digest=sha256:a47b63754fe8be509e5795ac567474744dc96fa01f0ab6fa3bfc29686acfd88e

Observation e808fdcb-b6c2-4d25-bb7d-d73c297dc441 · outbound

This paper cites Concrete Problems in AI Safety.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Concrete Problems in AI Safety

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.377706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.377706Z digest=sha256:530f6ce5e03208bac8f0a64d5934d826e3656b6d5aada9c0c76bd2b6ac02051e

Observation fe80d579-e51f-41e1-96b1-0655ba1804a0 · outbound

This paper cites Qwen Technical Report.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Qwen Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.382769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.382769Z digest=sha256:176c1ad9260b30cd4d025487cdf83bef2c15d4d1b99544773dabd60684fb8a30

Observation dc428be4-0dc3-4508-900e-89bd237f8020 · outbound

This paper cites Emergent Complexity via Multi-Agent Competition.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Emergent Complexity via Multi-Agent Competition

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.387429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.387429Z digest=sha256:238991c4c189c4c1169d983a73d334b795a29731149b5b1cbc1ec67ac0a7f700

Observation cb90f93b-66ae-418b-9270-a838e25382b9 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:32:00.716563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.396487Z digest=sha256:55e385eaddb4a190009fc1357a75f23ff6b7ffef6ad2b429d289de9d487c8cee

Observation 33556244-a5fb-4791-ae68-32141301413a · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:32:00.699664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.414533Z digest=sha256:c81910ccd1bdf31271b6b0625f2faa419a153699957f58adbc90d70d7f29bcc4

Observation 0845b035-c6e9-4ef5-b840-5557464b2c83 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:32:00.678724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.425470Z digest=sha256:46cef6a2dc0f20e79062b9cf0440dda0d0faec38b204d51b432d074b31371880

Observation 56e3d274-f9a2-4493-9e26-90213e9bc965 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.432600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.432600Z digest=sha256:c7f853566758b2254aad5c1d6f1e37fd605dea19ddfc8625a815b3ef021070bf

Observation 31b62ac5-511a-400e-8239-fa51a5c12ace · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:32:00.660012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.438800Z digest=sha256:0155ab3c1239fd8d2c776ea646a67325216181367ca6fe50102f948335ad1e51

Observation 53cc20fb-ca4b-4433-9ce8-119417ded970 · outbound

This paper cites Christiano, Jan Leike, Tom B.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Christiano, Jan Leike, Tom B

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:32:00.645154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.448565Z digest=sha256:66759268af9f2bd989713149903fa35ed155e670d11e46731c81d186e2de5a80

Observation b807f088-fd29-4a3c-88ff-08690762dcc7 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:32:00.628423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.456237Z digest=sha256:b0725b5a93227aacc234d5e078ad80e18caa436429ce7b307dbd7309e6068259

Observation f6b9c1cb-0d09-459e-b7fe-5221d89fd310 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.473720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.473720Z digest=sha256:63ebd1dbb20a99cede630043bbc2e95d9f7b9c5bddc910d567242b293d59f943

Observation d2cdc9af-7ef8-41d7-850b-0726031690e7 · outbound

This paper cites Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.481403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.481403Z digest=sha256:335b37d615569ca33b12175d1b6467471ee2902109e59a6ebb253aa4fcd25d8d

Observation 806d647f-eff4-49d1-b43c-f88187e4ae0c · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:32:00.603929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.492488Z digest=sha256:74156e8be072b1453967415ccf46f43e409fc07bb8329c317b5ee7b4f9483d03

Observation 37f8064f-43b1-4588-80f6-b4f8e37ecb1f · outbound

This paper cites Sharkey, Jacob Pfau, and David Krueger.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Sharkey, Jacob Pfau, and David Krueger

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:32:00.573914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.503016Z digest=sha256:6ff455062c31589adcda7fef4d46f0415a803a5cef89a86123768534b5c4d553

Observation e21614d8-da94-49b4-a126-d7340da0b202 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.511170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.511170Z digest=sha256:77e3528b4fb4ee3acf7a4d8d052d9c498604e5eb39ecf51e074e5e69e8fc313f

Observation 7f8677c2-5b30-41ac-a233-fa9c0bdc1e58 · outbound

This paper cites Reward Tampering Problems and Solutions in Reinforcement Learning: A Causal Influence Diagram Perspective.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Reward Tampering Problems and Solutions in Reinforcement Learning: A Causal Influence Diagram Perspective

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.520812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.520812Z digest=sha256:e5b9f1a53eca1eb4d4e6c2955c97ea26891270ed32b0666aa5949a865a987fc8

Observation 71103d20-cb28-43c6-8e79-6262c0ce05c4 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:32:00.552490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.527197Z digest=sha256:1c664b46bc10b6cb8278e2a9563a76a6949e4f23953e4a6b8d33340c6a52da65

Observation 44331071-acc5-481f-9936-d2f72986b32e · outbound

This paper cites MT-R1-Zero: Advancing LLM-based Machine Translation via R1-Zero-like Reinforcement Learning.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation MT-R1-Zero: Advancing LLM-based Machine Translation via R1-Zero-like Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.538085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.538085Z digest=sha256:c869ac6b2e154d4d118b6ffe780c478b23dd8bede66908f2c1f69c7502f31531

Observation 334c0700-d8a3-47ae-8f5d-f2d957034ff9 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 24

Resolution
verified exact
doi, observed 2026-08-07T10:31:59.022555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.547197Z digest=sha256:58972631cc29eedc667dd2c6911aa65e12c8b3c49f18dccfda436a28fc161b4c

Observation 6a3ba601-d6c1-4ce7-aef7-cc5dab0bda08 · outbound

This paper cites Adversarial Policies: Attacking Deep Reinforcement Learning.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Adversarial Policies: Attacking Deep Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.552488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.552488Z digest=sha256:ec1ddc3b7eff81e2b7603c977024487a9cb0bf5bc4edeedf63f79ea233fa43c9

Observation cbef659f-1911-4505-b718-a40d9cc3ad60 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:32:00.524308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.558415Z digest=sha256:5a97f6c58d84af6774092b1a458f230ade9c46888988969b6b43c2a89ce6a35b

Observation 32586a40-4472-4f8b-9681-debe3d640808 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 27

Resolution
verified exact
doi, observed 2026-08-07T10:31:59.005534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.563586Z digest=sha256:2f4dd10b5a33cd2704922f9cfdcd9de8f5317374cb94716019f7da46c3d27b35

Observation 8c8e82d1-5f1e-4856-a566-c97b63036b82 · outbound

This paper cites The Llama 3 Herd of Models.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation The Llama 3 Herd of Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.569830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.569830Z digest=sha256:c6c768a93afdb8e730eb067ead5a3a081cae95f1e03c836a05788c4a04db346b

Observation 3523feee-8ce4-4f6e-85b7-074511019571 · outbound

This paper cites R1-T1: Fully Incentivizing Translation Capability in LLMs via Reasoning Learning.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation R1-T1: Fully Incentivizing Translation Capability in LLMs via Reasoning Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.579714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.579714Z digest=sha256:7aaf8e24a320f86e070a43dabeaf1bf9606fe0cbbbbd57e542a1767f891e45c4

Observation a0995178-3ee5-4158-934c-53e7407fa2ca · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.585982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.585982Z digest=sha256:0fdefce7359631fd08307886f5b03417b39fe6eea4f9f7b8123beb737add7bfe

Observation bfef10d8-d7b7-4db8-8025-09de6bd6f92d · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.592766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.592766Z digest=sha256:c64b8d7b431b800c9e12ccede4c5d70183dec4daf379270b96d7d7e1c6f309be

Observation 649b3153-4ee3-47b6-99a6-77e1b68a46c2 · outbound

This paper cites FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.602250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.602250Z digest=sha256:cc1da7db2a33f77f797b0146179067aef4e062034ecf2ff87c206d46b44e3c23

Observation d9760982-7134-4d01-88d6-9f33f48e48e6 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:32:00.492138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.607897Z digest=sha256:4f7b4c57bca21089e8ae0fa7301781f775c267c329161f7b857725def515e14e

Observation 1498e946-3515-4cf7-a44f-80a8c60823e6 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.616272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.616272Z digest=sha256:b9ff465fbad15993c4f2cd44c7dd7631b0b8040655677105b85455809d3d356f

Observation 3f2dc168-26ff-42c7-9480-3cc7297b2e88 · outbound

This paper cites Data Selection Curriculum for Neural Machine Translation.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Data Selection Curriculum for Neural Machine Translation

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:31:59.584675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.624070Z digest=sha256:37aca772ca2d263c720200a662098d0f76dd2a86f7d8752756eab6365dd481f2

Observation c76aba47-09a2-417d-9cbd-2353896b00d0 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.631752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.631752Z digest=sha256:90b75bca33ae7690e065347b4b3238d460e4537695348c1cf11a69ff028eddfd

Observation de183fdf-da0d-405a-9522-aa91a473317c · outbound

This paper cites OpenAI o1 System Card.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation OpenAI o1 System Card

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.636142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.636142Z digest=sha256:ab94b18b1cb944b8587c2d7bbd3a98ee662331d78a35da66b33c4e587632b9fb

Observation 81dbe7df-1593-4f49-8999-e9be208fac69 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.642529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.642529Z digest=sha256:10f75e3909ddb69a484113c127cacb5ed7ab23e02e76a3df2bf35a70284c252c

Observation ba0f3cae-cacc-4523-ae81-b2088667c66a · outbound

This paper cites The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.649510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.649510Z digest=sha256:b53de462c9cde4da199d58fe0b68d3f14625c3c6faa989a0e48d4363a1cf255c

Observation 6ff6da39-141a-4f59-892b-45aec30628e4 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.658846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.658846Z digest=sha256:6296e7e7712eb43c3ed6d637644fbbf792b3e21289d687f8e0d57d39e0ae2ab7

Observation 72c4d4ec-207c-4e44-af5a-22c86d951f6e · outbound

This paper cites A Deep Reinforced Model for Abstractive Summarization.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation A Deep Reinforced Model for Abstractive Summarization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.667226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.667226Z digest=sha256:453b74b63db952f5e5af07e50d53e08b45cf7b5de403b4e5d2b33d275002f43e

Observation 634703e7-9912-421d-99fa-4981135fea13 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.674801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.674801Z digest=sha256:e1e8aa7db2714e9208bd3efb69b18c3e69c1c5c3f5f5641664c34d2e8d2c5eee

Observation 7546240a-fe56-49ee-9a50-cb308bc9dcce · outbound

This paper cites Sequence Level Training with Recurrent Neural Networks.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Sequence Level Training with Recurrent Neural Networks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.686521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.686521Z digest=sha256:2f06a05b66380f12694012467835a6c35cdab4a9b6139fda5b96b99169eb4760

Observation de7414d2-0c46-47b5-a8cd-bdd307e1ccbd · outbound

This paper cites CometKiwi: IST-Unbabel 2022 Submission for the Quality Estimation Shared Task.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation CometKiwi: IST-Unbabel 2022 Submission for the Quality Estimation Shared Task

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.695659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.695659Z digest=sha256:935f80bb5b42d796e6acbd4b4dcf4983174e43e090ebb2b2a6a471ce546e5b5b

Observation f6b07334-fa51-48c8-93af-475f9dd68877 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:32:00.428159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.708564Z digest=sha256:d27f502fab0c3654915faf1a019ccb0745b0660377a74086fef70876ca14fa9e

Observation 3fbde2c5-7427-45d4-b5ff-54aeb36988ba · outbound

This paper cites Proximal Policy Optimization Algorithms.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Proximal Policy Optimization Algorithms

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.718462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.718462Z digest=sha256:29c2301e329473e1671ed68921fd24e22d7a19228d17c80f0673ab6b69474a72

Observation 7146a338-e57b-45a4-b569-ff02429e57e1 · outbound

This paper cites BLEURT: Learning Robust Metrics for Text Generation.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation BLEURT: Learning Robust Metrics for Text Generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.726456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.726456Z digest=sha256:0a6637bf90076720cb86dca73d8818ec17b0d076aa0c5ab76a9926234fc8da27

Observation db15dab3-617d-45c3-9bc5-5c2e30ac7a59 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.732799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.732799Z digest=sha256:25f9b4862bd49868e9067432fb17bd8b4b07e12c8197cc7ef964c8fe927b545f

Observation 9025b613-d158-4572-a2cf-f3a101bd963c · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation HybridFlow: A Flexible and Efficient RLHF Framework

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.745477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.745477Z digest=sha256:9fe2d562af70b4df208d1a859c01c381fea2d79d7336db77135ef160bc7fffce

Observation 9947951e-4a71-4b89-b52c-05bb67fd4854 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.756060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.756060Z digest=sha256:abf97827b090227a068551fb05436413d190bfcf0b93bfcbe199c37752974991

Observation 2f3e331b-e8ef-4e76-b81d-f74935f5872e · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:32:00.404365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.765145Z digest=sha256:d434aa6b8c7d6d127b2d0c665da95d322ec53c8a66d0de36d31ca706dc941dac

Observation 8a9f80ef-f7d5-49e6-980f-6ae046fda1ad · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.771678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.771678Z digest=sha256:0e62730ca2c417ec9f8f04e70000ad31f8e1c180a8655b5ed1e26fc8c921412f

Observation 33e343d7-4346-4f97-b0a4-f8318d52c449 · outbound

This paper cites Remedy: Learning Machine Translation Evaluation from Human Preferences with Reward Modeling.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Remedy: Learning Machine Translation Evaluation from Human Preferences with Reward Modeling

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.777646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.777646Z digest=sha256:92649f324f4c71af19f84a8e50b9bed7e9414631052af51faf6acac4f293cb16

Observation 63f20127-0258-40f2-ae54-3a42bcec1421 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.786233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.786233Z digest=sha256:8cd97cad89e1dfe37b5b62cbbc21094b70c593737af98c534e2d8c6b5be77585

Observation b2132ac7-0138-43b4-9e8b-6ed0322c2619 · outbound

This paper cites Avoiding Tampering Incentives in Deep RL via Decoupled Approval.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Avoiding Tampering Incentives in Deep RL via Decoupled Approval

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.792853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.792853Z digest=sha256:52bced12522081e84f3004342b460bc316330f382414244b1c15be1b0ea9898f

Observation bc1bbd2b-cabe-4d28-9688-5e48e084313d · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.798141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.798141Z digest=sha256:a391b78d34d9a15bc4dbb29de101c63a5fe6ede5e058a7cbf1a5b512c0154f9a

Observation b456d0f2-4452-4108-b873-bff66967c961 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:32:00.364386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.806096Z digest=sha256:61b4878485d391460671cb1a0aebe9ddbba960f8dd0e6ca206f2539a82f4c0b8

Observation 7db21fb0-8274-483a-855a-770646a5f354 · outbound

This paper cites Large Language Models are Better Reasoners with Self-Verification.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Large Language Models are Better Reasoners with Self-Verification

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.811397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.811397Z digest=sha256:a1d441e0d2850e703d5c06ae8177a49565f2b53f9854180f2b8f574bae4add21

Observation 4f4a20ab-a61a-4ed7-a2f4-e9030f44c884 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.822218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.822218Z digest=sha256:e85eb5debab2383ccfe196ca3f96b37e9a7fa9a9f72419de6807b124594ac00e

Observation 3dba0183-55ae-4acf-8082-250d94d7ab31 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.827925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.827925Z digest=sha256:1332bfe3d3ff45fcadb1f46e5923b802bea3c0de400470f5d854ff6200a3fa57

Observation 8b9d0a21-403c-4f8c-b44b-a0a6dc4c42d8 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.837869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.837869Z digest=sha256:543de9e389b4c6432edbb5b9b13de5b8cef90349d8441016127f222ecd98d83a

Observation bdadc1b5-c4b4-4594-b592-672929e0f370 · outbound

This paper cites A Paradigm Shift in Machine Translation: Boosting Translation Performance of Large Language Models.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation A Paradigm Shift in Machine Translation: Boosting Translation Performance of Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.843811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.843811Z digest=sha256:5ee10523c250380bf883c68f03daa01afe2e2d0f6bc7153f64e6e579885961d8

Observation 74720cda-6b79-43b4-9a2f-e39f780edf78 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:32:00.318317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.850713Z digest=sha256:7ebcd2c356a74b9227dbfb3656ca126102f957cb8eedf8b7b6f8211ea981477c

Observation 39ece709-7d17-4acd-a95f-5c97499f6594 · outbound

This paper cites Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.854906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.854906Z digest=sha256:bf7f84197fc08acadaa02c16c14ee737c3f729d143444b6292f26275dab21c0d

Pith citing papers

Observation 50cc0c47-2c40-480a-bfaa-f44cd5bc1755 · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation

Reference 109

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.742286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:b43326a0d8615bfb8e07a0a04f15621e0712988f78807bd12a180b6f34d1243d