Pith. sign in

Paper Citation Record · LEDGER

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation

As of 13 August 2026, this Paper Citation Record lists 64 of 64 outbound references and 1 inbound Pith citation observation for arXiv:2506.05070.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.05070 v2

Coverage vector

measured 64 of 64 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T10:31:58.854906Z

measured 65 of 65 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-10T13:58:53.430492Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T14:00:28.739440Z

Reference resolution

64 of 64 outbound references displayed

  • verified exact3
  • verified fuzzy2
  • unresolved59
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 4aef4f8a-4efb-448e-9302-1fb752988ee2 · outbound

This paper cites URL: " 'urlintro :=.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation URL: " 'urlintro :=

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.341810Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.341810Z digest=sha256:8ebaf305484f736c1f050011b01deb260eab49aff709c5ad508fad9d6f596a50

Observation 43a10ab0-bd84-49f3-ab43-6a4e4d51e9ae · outbound

This paper cites write newline.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation write newline

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.348130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.348130Z digest=sha256:98d5f9db238db5be015f3bbc52ad57f0414f22eb9ac0111dfa686c121658d21a

Observation d3c297ee-c723-4649-ad59-260d500d7ba2 · outbound

This paper cites GPT-4 Technical Report.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation GPT-4 Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.352960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.352960Z digest=sha256:360f3dc22f66a71a9a72e6b72dfc61381ae36efe1dd439f7648885b1fa6d8428

Observation 814d1cc0-7757-44a5-8247-f66e06bdb408 · outbound

This paper cites Scalable Ensembling For Mitigating Reward Overoptimisation.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Scalable Ensembling For Mitigating Reward Overoptimisation

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.359316Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.359316Z digest=sha256:e518544b1c2118c2a87f6b8893b4058a42a26207732b3dd01433785d61949d78

Observation c495aa36-a0a6-4e35-897b-b39e32dacadb · outbound

This paper cites Tower: An Open Multilingual Large Language Model for Translation-Related Tasks.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Tower: An Open Multilingual Large Language Model for Translation-Related Tasks

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.366539Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.366539Z digest=sha256:d8f2f959f1b908319795d4026a9fcf1b45ac3fd5b959f3ce76fb482925fe9b9e

Observation e808fdcb-b6c2-4d25-bb7d-d73c297dc441 · outbound

This paper cites Concrete Problems in AI Safety.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Concrete Problems in AI Safety

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.377706Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.377706Z digest=sha256:0ff10d7227d743ff748b340519c5f529cce3c5d83e38b5982464ee340eed45e2

Observation fe80d579-e51f-41e1-96b1-0655ba1804a0 · outbound

This paper cites Qwen Technical Report.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Qwen Technical Report

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.382769Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.382769Z digest=sha256:474cc38b120aced1db28ef88ae623c1c1a117ba46ec0835eeefe8e208ac1725a

Observation dc428be4-0dc3-4508-900e-89bd237f8020 · outbound

This paper cites Emergent Complexity via Multi-Agent Competition.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Emergent Complexity via Multi-Agent Competition

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.387429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.387429Z digest=sha256:392a7590d49127bc5d439e354ee3c5e20cb9d6fc82a04bcd42458663ba8e2e48

Observation cb90f93b-66ae-418b-9270-a838e25382b9 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:32:00.716563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.396487Z digest=sha256:2a87e98030a3ca97d9ed6b60cdddf0b26979fb8a5f7d6390d9038098355024eb

Observation 33556244-a5fb-4791-ae68-32141301413a · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 10

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:32:00.699664Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.414533Z digest=sha256:29576674d166a614f88cb5badc20ab39b6c9137169722a0cbda88186b7e4a400

Observation 0845b035-c6e9-4ef5-b840-5557464b2c83 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 11

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:32:00.678724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.425470Z digest=sha256:cc2a7b9c46c2b08590d7eb8814eb9c82bd6ea81250f2e4eee035d573d974d7a1

Observation 56e3d274-f9a2-4493-9e26-90213e9bc965 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.432600Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.432600Z digest=sha256:274ca1b599a265bacf665f999800f90fa092f9efee4826c77d60e7238480d3b6

Observation 31b62ac5-511a-400e-8239-fa51a5c12ace · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:32:00.660012Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.438800Z digest=sha256:a49ece17d5b7512b6f1b86f48f66b8575d16f2c7b876b19d7cacc318fa14a44b

Observation 53cc20fb-ca4b-4433-9ce8-119417ded970 · outbound

This paper cites Christiano, Jan Leike, Tom B.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Christiano, Jan Leike, Tom B

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:32:00.645154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.448565Z digest=sha256:425bf8566a2f4d6e18e8b6cbc2b8d4621a2e920281fe793273250810b03ee811

Observation b807f088-fd29-4a3c-88ff-08690762dcc7 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:32:00.628423Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.456237Z digest=sha256:27497b325b236bb5fe6817b945963f26dd36d1aa541bb8a483b97d354936a3a8

Observation f6b9c1cb-0d09-459e-b7fe-5221d89fd310 · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.473720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.473720Z digest=sha256:c825a140c00cf49c8e6abc4c94c7f2abb488e7c779155ce1f7dd3135579cd444

Observation d2cdc9af-7ef8-41d7-850b-0726031690e7 · outbound

This paper cites Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Sycophancy to Subterfuge: Investigating Reward-Tampering in Large Language Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.481403Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.481403Z digest=sha256:74990d52ebf79fff1412dcfd33c94b04ddea03cffacd63250714f07bcd738ce2

Observation 806d647f-eff4-49d1-b43c-f88187e4ae0c · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 18

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:32:00.603929Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.492488Z digest=sha256:4a61890cf423764771f19bf2e6df502bd5a652cf0e2222cdd0671499ee7b223e

Observation 37f8064f-43b1-4588-80f6-b4f8e37ecb1f · outbound

This paper cites Sharkey, Jacob Pfau, and David Krueger.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Sharkey, Jacob Pfau, and David Krueger

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T10:32:00.573914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.503016Z digest=sha256:d0869d235a0431414f4c6fa2791b56b94c17fb5735ce38da252f4af7bf5595cc

Observation e21614d8-da94-49b4-a126-d7340da0b202 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.511170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.511170Z digest=sha256:8ff42085f50f9f15ec5d50e4c7065ffb96eab546d45648eab140e35d19733c2b

Observation 7f8677c2-5b30-41ac-a233-fa9c0bdc1e58 · outbound

This paper cites Reward Tampering Problems and Solutions in Reinforcement Learning: A Causal Influence Diagram Perspective.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Reward Tampering Problems and Solutions in Reinforcement Learning: A Causal Influence Diagram Perspective

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.520812Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.520812Z digest=sha256:7f9039891e8722fbf131aa36f11929d901e00ea6c453d688ab3be8ebcb6ef4d5

Observation 71103d20-cb28-43c6-8e79-6262c0ce05c4 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:32:00.552490Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.527197Z digest=sha256:5ebb9c363cc463346b94e1c4c1cdded920435a883a9717ac3cddefe1481cd88a

Observation 44331071-acc5-481f-9936-d2f72986b32e · outbound

This paper cites MT-R1-Zero: Advancing LLM-based Machine Translation via R1-Zero-like Reinforcement Learning.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation MT-R1-Zero: Advancing LLM-based Machine Translation via R1-Zero-like Reinforcement Learning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.538085Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.538085Z digest=sha256:7603220ae04a7c832145558a49d975efbda48374e68347fdb39af70470ceb0fe

Observation 334c0700-d8a3-47ae-8f5d-f2d957034ff9 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 24

Resolution
verified exact
doi, observed 2026-08-07T10:31:59.022555Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.547197Z digest=sha256:6d1d5a29d0d948fb2bd4fa2ff403fdbe317c89bd9a02e0f900a0f4e6829366ac

Observation 6a3ba601-d6c1-4ce7-aef7-cc5dab0bda08 · outbound

This paper cites Adversarial Policies: Attacking Deep Reinforcement Learning.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Adversarial Policies: Attacking Deep Reinforcement Learning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.552488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.552488Z digest=sha256:4859aead7ac7d4611cc6ca744b3649161cae2fbdf97ed0b40b700ca3affbefc1

Observation cbef659f-1911-4505-b718-a40d9cc3ad60 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:32:00.524308Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.558415Z digest=sha256:8f3f27a232c32f1bdf23a8e512914fa7d0ff78307e5453c5515b1c2293b6fa09

Observation 32586a40-4472-4f8b-9681-debe3d640808 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 27

Resolution
verified exact
doi, observed 2026-08-07T10:31:59.005534Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.563586Z digest=sha256:476373f3113e9bfbd71bd0e48bb0560dd467048ee9c197d53fc0869dadff44a1

Observation 8c8e82d1-5f1e-4856-a566-c97b63036b82 · outbound

This paper cites The Llama 3 Herd of Models.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation The Llama 3 Herd of Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.569830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.569830Z digest=sha256:0b05d4971c4a5d0cb27d43ea69458002b309928a7f87d81fddeffa6921729fdf

Observation 3523feee-8ce4-4f6e-85b7-074511019571 · outbound

This paper cites R1-T1: Fully Incentivizing Translation Capability in LLMs via Reasoning Learning.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation R1-T1: Fully Incentivizing Translation Capability in LLMs via Reasoning Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.579714Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.579714Z digest=sha256:3340eb17a34f08878a27fe1a03906ea25c668ed3cc4d32d34cada600ed353541

Observation a0995178-3ee5-4158-934c-53e7407fa2ca · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.585982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.585982Z digest=sha256:aefe8647e79b393bdfc13d7237b99246a84b6918c135e2749ad189203d4088d4

Observation bfef10d8-d7b7-4db8-8025-09de6bd6f92d · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.592766Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.592766Z digest=sha256:1a1f7880338baa87deaa71b5a747b5bff6379de6b9dc560561de887ffb96c2c0

Observation 649b3153-4ee3-47b6-99a6-77e1b68a46c2 · outbound

This paper cites FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation FLEUR: An Explainable Reference-Free Evaluation Metric for Image Captioning Using a Large Multimodal Model

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.602250Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.602250Z digest=sha256:926839e50df25d52b2e6cb805dc9c28bdc284fb4ee0d858030e1693cf48ac685

Observation d9760982-7134-4d01-88d6-9f33f48e48e6 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:32:00.492138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.607897Z digest=sha256:1e87695090e462b80b6e24adeff7d26c03b9e548f1e730e520fbfec418b3e319

Observation 1498e946-3515-4cf7-a44f-80a8c60823e6 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.616272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.616272Z digest=sha256:8cdd2877f419909df53a279cbb358a5bc5602177361a2c50db1f52e93e143623

Observation 3f2dc168-26ff-42c7-9480-3cc7297b2e88 · outbound

This paper cites Data Selection Curriculum for Neural Machine Translation.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Data Selection Curriculum for Neural Machine Translation

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-08-07T10:31:59.584675Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.624070Z digest=sha256:8f81fbac1e1c448f78dc16d9d5fd39281f16e3e33543cdcfbe4876c4f8fd639b

Observation c76aba47-09a2-417d-9cbd-2353896b00d0 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.631752Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.631752Z digest=sha256:e06fa651edd0c81b0d60e72008373148fda3f9df45cda26bd43c60718e1478c1

Observation de183fdf-da0d-405a-9522-aa91a473317c · outbound

This paper cites OpenAI o1 System Card.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation OpenAI o1 System Card

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.636142Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.636142Z digest=sha256:05ff3f2644492bda5363c7b959a5db7dc64b21577c8bd444592027348acb42ee

Observation 81dbe7df-1593-4f49-8999-e9be208fac69 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.642529Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.642529Z digest=sha256:5a2c33b5ec454130587eaf190b731cae40661a1ef329351e9d36384ea629a89f

Observation ba0f3cae-cacc-4523-ae81-b2088667c66a · outbound

This paper cites The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation The Effects of Reward Misspecification: Mapping and Mitigating Misaligned Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.649510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.649510Z digest=sha256:723b879644a176bbea56bebc53655b36f7a46967fe19d443a877762eb1e31afe

Observation 6ff6da39-141a-4f59-892b-45aec30628e4 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.658846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.658846Z digest=sha256:b546facc04e2d10000a53b8a7aa8ae915a21fa7dd390f621025d0d77fcf239d5

Observation 72c4d4ec-207c-4e44-af5a-22c86d951f6e · outbound

This paper cites A Deep Reinforced Model for Abstractive Summarization.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation A Deep Reinforced Model for Abstractive Summarization

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.667226Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.667226Z digest=sha256:16ca2b52f772e14737d5cf764e78e86dfbf421665a8c0515b1b7b54bb35208a0

Observation 634703e7-9912-421d-99fa-4981135fea13 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.674801Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.674801Z digest=sha256:d4832d852cb0ea4b94bb0825c6a07593adb5a901bab2957dc8df931231f93fa7

Observation 7546240a-fe56-49ee-9a50-cb308bc9dcce · outbound

This paper cites Sequence Level Training with Recurrent Neural Networks.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Sequence Level Training with Recurrent Neural Networks

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.686521Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.686521Z digest=sha256:f7d3febc70b3bd14d6626a496024d063f621573b4148e2abc0f4696669382c2d

Observation de7414d2-0c46-47b5-a8cd-bdd307e1ccbd · outbound

This paper cites CometKiwi: IST-Unbabel 2022 Submission for the Quality Estimation Shared Task.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation CometKiwi: IST-Unbabel 2022 Submission for the Quality Estimation Shared Task

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.695659Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.695659Z digest=sha256:9a32ad2390c579cf0cf9a54de382ebbe5e0eb0b523c2a8e747365048e2164ed0

Observation f6b07334-fa51-48c8-93af-475f9dd68877 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:32:00.428159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.708564Z digest=sha256:75d8bcb657c9c777cf45f5357d33d8981309bd2ac0797a39df30a6ee49c72e6f

Observation 3fbde2c5-7427-45d4-b5ff-54aeb36988ba · outbound

This paper cites Proximal Policy Optimization Algorithms.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Proximal Policy Optimization Algorithms

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.718462Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.718462Z digest=sha256:355b52533477017bfa8eedee0cabe51171508d48705e0014a6c84f81fd014af8

Observation 7146a338-e57b-45a4-b569-ff02429e57e1 · outbound

This paper cites BLEURT: Learning Robust Metrics for Text Generation.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation BLEURT: Learning Robust Metrics for Text Generation

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.726456Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.726456Z digest=sha256:181840a2247484a62c7532b0c83cd6056a4f6700d66aa3918ee99e16a605bbeb

Observation db15dab3-617d-45c3-9bc5-5c2e30ac7a59 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.732799Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.732799Z digest=sha256:c5af278019df7d0cb7844c98215a62c690066b849ddbe05b45df273cb8391d92

Observation 9025b613-d158-4572-a2cf-f3a101bd963c · outbound

This paper cites HybridFlow: A Flexible and Efficient RLHF Framework.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation HybridFlow: A Flexible and Efficient RLHF Framework

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.745477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.745477Z digest=sha256:ce0cab9bb6271bb1ac5eb667a59050af6cffc5fd5015938666e373ed651085f5

Observation 9947951e-4a71-4b89-b52c-05bb67fd4854 · outbound

This paper cites Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.756060Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.756060Z digest=sha256:9ffcb35cf10c1d78bf7ea0eeff50ef682a02a10cdbd0230eb1e5051d203f4c43

Observation 2f3e331b-e8ef-4e76-b81d-f74935f5872e · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:32:00.404365Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.765145Z digest=sha256:8e527935dd0670bb4a4f11d72633e39cb80f87aa97a514207d4333d8d05bd296

Observation 8a9f80ef-f7d5-49e6-980f-6ae046fda1ad · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.771678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.771678Z digest=sha256:606a7e9593e0af88ce6edc4dd052a52bfb175e667602deb76c34e15ac873add3

Observation 33e343d7-4346-4f97-b0a4-f8318d52c449 · outbound

This paper cites Remedy: Learning Machine Translation Evaluation from Human Preferences with Reward Modeling.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Remedy: Learning Machine Translation Evaluation from Human Preferences with Reward Modeling

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.777646Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.777646Z digest=sha256:cabf6e97c0d2f9cdc79418e52cc1d0cc88a9877967ee0d15d26ec9fb8d616a43

Observation 63f20127-0258-40f2-ae54-3a42bcec1421 · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.786233Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.786233Z digest=sha256:bbd88026403772b517783cc87c9f8c45a85d7f2b52c40b41536201fbd6e49184

Observation b2132ac7-0138-43b4-9e8b-6ed0322c2619 · outbound

This paper cites Avoiding Tampering Incentives in Deep RL via Decoupled Approval.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Avoiding Tampering Incentives in Deep RL via Decoupled Approval

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.792853Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.792853Z digest=sha256:8450f1796316ca1363b891a9f86fb3c24c134fab069d6edc503b010ff7b1569b

Observation bc1bbd2b-cabe-4d28-9688-5e48e084313d · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.798141Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.798141Z digest=sha256:0c97bdb72214e4a26d37b7f4d12fdd9ea85792d0e79033fa2e0fded7a61e3f0e

Observation b456d0f2-4452-4108-b873-bff66967c961 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 58

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:32:00.364386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.806096Z digest=sha256:e7f278ab20ef2205a101342e56654a1b0d3ae75050f7158db477c20ad14e9ec3

Observation 7db21fb0-8274-483a-855a-770646a5f354 · outbound

This paper cites Large Language Models are Better Reasoners with Self-Verification.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Large Language Models are Better Reasoners with Self-Verification

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.811397Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.811397Z digest=sha256:f66519ed01d449b6cdf992ce014fe74cbd81bd004a42b8e783516111d2ce52b2

Observation 4f4a20ab-a61a-4ed7-a2f4-e9030f44c884 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.822218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.822218Z digest=sha256:eace5f1bb46c117149d821d6adc748071f49b738e02407afea0fe820d7d88a10

Observation 3dba0183-55ae-4acf-8082-250d94d7ab31 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.827925Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.827925Z digest=sha256:c5290cd9b7be030bfd0f8ebfdacbb2b73b464e2595316b35d419d117e2665e2e

Observation 8b9d0a21-403c-4f8c-b44b-a0a6dc4c42d8 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.837869Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.837869Z digest=sha256:1f30c23910d6bb8160ddb46314762c29cef7d61d31301493bdfa7a0d3d432599

Observation bdadc1b5-c4b4-4594-b592-672929e0f370 · outbound

This paper cites A Paradigm Shift in Machine Translation: Boosting Translation Performance of Large Language Models.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation A Paradigm Shift in Machine Translation: Boosting Translation Performance of Large Language Models

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.843811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.843811Z digest=sha256:6a9d12f3024cc9335f6397ec4dee999542e3c9772641c561a288e78ca0eedd80

Observation 74720cda-6b79-43b4-9a2f-e39f780edf78 · outbound

This paper cites an unresolved cited work.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Unresolved cited work

Reference 64

Resolution
unresolved
raw_fallback, observed 2026-08-07T10:32:00.318317Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T10:31:58.850713Z digest=sha256:62ce886b5d33a0996c25124e6b45bb8c5a4ca27c09b8070356395ee9b20cc588

Observation 39ece709-7d17-4acd-a95f-5c97499f6594 · outbound

This paper cites Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation.

RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation Contrastive Preference Optimization: Pushing the Boundaries of LLM Performance in Machine Translation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-07T10:31:58.854906Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:31:58.854906Z digest=sha256:2aad60f383a5a406415abe49efa9088f8058a7dd24b19c80f58169aa431c168d

Pith citing papers

Observation 50cc0c47-2c40-480a-bfaa-f44cd5bc1755 · inbound

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges cites this paper.

Reward Hacking in the Era of Large Models: Mechanisms, Emergent Misalignment, Challenges RIVAL: Reinforcement Learning with Iterative and Adversarial Optimization for Machine Translation

Reference 109

Resolution
verified exact
arxiv_id, observed 2026-05-10T14:00:28.742286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T13:58:53.430492Z digest=sha256:54990bef74210c321be05e9b7c57fa6306364abea37fc205d45d19fb6f476c73