Pith. sign in

Paper Citation Record · LEDGER

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models

As of 18 August 2026, this Paper Citation Record lists 45 of 45 outbound references and 2 inbound Pith citation observations for arXiv:2506.07165.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07165 v1

Coverage vector

measured 45 of 45 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:48:04.530071Z

measured 47 of 47 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-16T00:52:06.745056Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T14:02:39.792894Z

Reference resolution

45 of 45 outbound references displayed

  • verified exact1
  • verified fuzzy14
  • unresolved29
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 509ef113-699f-40ce-8791-51ccddbc9a56 · outbound

This paper cites an unresolved cited work.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:48:04.380172Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:48:04.380172Z digest=sha256:9dc4d8b21b268f15bef66c1b17a99dcd20e2fb8fa9868a5f405ab46db2ad219e

Observation 6a775cc4-d1d3-4e49-8a9b-cb25c3ccb4a9 · outbound

This paper cites - (2) Acknowledges both but slight deviations.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models - (2) Acknowledges both but slight deviations

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:48:04.384800Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:48:04.384800Z digest=sha256:5d8d0968d207d0935e0e246bececa09864fce6b88c64fab31782ab81acdf2ccc

Observation 1b1a3283-9453-4c75-b10b-7acce1c8d99c · outbound

This paper cites Unified Preference Optimization: Language Model Alignment Beyond the Preference Frontier.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unified Preference Optimization: Language Model Alignment Beyond the Preference Frontier

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:48:04.758497Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.333706Z digest=sha256:02cb0733d47c68042c5b591b10ddb4b14dd0d415d6b8ba60b888dd484bc3e255

Observation 71380829-229d-44c8-9d61-8c582f171bdb · outbound

This paper cites Based the instruction following rule and given my answer to an instruction, your role is to provide specific and constructive score for me.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Based the instruction following rule and given my answer to an instruction, your role is to provide specific and constructive score for me

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:48:05.206741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.394469Z digest=sha256:d873da63b453eb2bd190ee3b7a78e13cd51189568411c63f503e47477d336cdc

Observation 86162f79-067f-4f88-b0c0-d2f4e8348557 · outbound

This paper cites an unresolved cited work.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work

Reference 5

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:48:04.909769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.489631Z digest=sha256:1a89b7dc171e120df791d5d776f4ae8c429ea087e43f35e1ce25fa7964da47f8

Observation 4aab7911-1989-4d87-9051-75447098113c · outbound

This paper cites Multi-Objective Alignment of Large Language Models Through Hypervolume Maximization.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Multi-Objective Alignment of Large Language Models Through Hypervolume Maximization

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:48:04.350188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:48:04.350188Z digest=sha256:661c0248e13ec22ec9638d7f9ae90b6346307b9b3f6addcbd44c258e1b9d9ebc

Observation fbff26fe-c0c9-46e4-bedc-349cecf533f1 · outbound

This paper cites an unresolved cited work.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work

Reference 7

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:48:04.804950Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.520670Z digest=sha256:d3940e0b1da9078aa97eef5ac5facb48779290a8d29260243d66828dfa86be9e

Observation 603471d7-87ff-4375-9720-010016c55d82 · outbound

This paper cites an unresolved cited work.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work

Reference 8

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:48:05.258845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.360084Z digest=sha256:f61644e4f570960bd1758cda3a15c6d2dc8a277bc95858b3c5ec3088637a731c

Observation 5fbb7736-bb7b-42f0-86de-4dcedd084167 · outbound

This paper cites Al ter na tively, you can use a nav iga tion app like Google Maps or Waze to get the most ac cu rate and up -to -date di rec tions.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Al ter na tively, you can use a nav iga tion app like Google Maps or Waze to get the most ac cu rate and up -to -date di rec tions

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:48:04.774067Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.530071Z digest=sha256:f912bf21c1c7cf83d8147bb89dc7e3c44aa5427a15e4247a4e7ba12394d59c80

Observation a6c9871d-7ba5-4d27-ba1d-24fa4ab96c53 · outbound

This paper cites RRHF: Rank Responses to Align Language Models with Human Feedback without tears.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models RRHF: Rank Responses to Align Language Models with Human Feedback without tears

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:48:04.370319Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:48:04.370319Z digest=sha256:9cbf937f5fd01531066d6fac7018b4420a04818453048cab83e16b6820839eda

Observation 1f6ca1b9-2929-4fd4-a204-d7da3e8cf2d7 · outbound

This paper cites an unresolved cited work.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:48:04.375331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:48:04.375331Z digest=sha256:8f7fe2f597e22038c21395ae6f95d445f6696cfd6da76c97742bf3b711ad488a

Observation d9e92826-9fb6-46ae-8ebc-a6341bdc5fa9 · outbound

This paper cites an unresolved cited work.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:48:04.389549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:48:04.389549Z digest=sha256:0f176fcf70cec4ed08e18d2d042892d87f2abf1b76a2da420392a68ffac3f216

Observation ed371d59-8070-42de-9b40-2ce71ec1f177 · outbound

This paper cites The response completely missed the essence of what the user wanted.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models The response completely missed the essence of what the user wanted

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:48:05.192139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.399006Z digest=sha256:a3d806d7267f00fc4c76e9b4aa4f1b9a4666f1197ba2641cd5145f872c82665e

Observation 0cd52a9f-1b84-4ad1-a6bd-9cbb8fd30d50 · outbound

This paper cites an unresolved cited work.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:48:05.177921Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.403467Z digest=sha256:243cf99cf8fde83964cc8bf5ecf4c6975a6e4a28b7bdfb099f12eeca8ad660ad

Observation cd53e2c5-9267-4541-9df4-ee050e8602fb · outbound

This paper cites The response did not fully satisfy what the user was looking for.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models The response did not fully satisfy what the user was looking for

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:48:05.163942Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.408360Z digest=sha256:f59f8fc1b5430a9c8b50e0a1a56edbd75ee4b7f1a6611a03a14d1a06b74e8cd5

Observation aa9b1903-cb42-4c0a-814c-254e76fae5c4 · outbound

This paper cites an unresolved cited work.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work

Reference 19

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:48:05.150165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.413250Z digest=sha256:70f47efd777b6696cc126a4ddfec165162e0ae067726e3a46f84b8fa89d90679

Observation cbc3e889-c355-4e81-9ff2-9165a1ad4ad2 · outbound

This paper cites Based the helpfulness rule and given my answer to an instruction, your role is to provide specific and constructive score for me.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Based the helpfulness rule and given my answer to an instruction, your role is to provide specific and constructive score for me

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:48:05.136403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.417719Z digest=sha256:6d121e76fa9d923cb16e810fcc6f16a1cd888c4f370b35018b243a80935c72c9

Observation 32325681-7ec1-4dd6-9bdf-f0ea1f3c9990 · outbound

This paper cites All information provided is wrong, false or hallucinated.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models All information provided is wrong, false or hallucinated

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:48:05.122505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.422395Z digest=sha256:3301d2ab107aa61321ec8eae2cb76f9ea86465d7780ea100f4b11a3151b215e9

Observation 52821ed3-340a-4e97-a005-215c70773b3c · outbound

This paper cites The response may contain multiple instances of hallucinations, false information, misleading information, or irrelevant information.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models The response may contain multiple instances of hallucinations, false information, misleading information, or irrelevant information

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:48:05.108613Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.426726Z digest=sha256:501d810178a659b8281c2f264b18a1ebafe1d8b3249df5e4b471a07ad28fceb4

Observation 58aa1c99-ebb1-497d-9d3e-ebc5d1b74f66 · outbound

This paper cites The response may miss some details, contain misleading information, or minor hallucinations, but is more or less aligned with what the prompt asks for.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models The response may miss some details, contain misleading information, or minor hallucinations, but is more or less aligned with what the prompt asks for

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:48:05.094666Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.431058Z digest=sha256:021e144f12dbc0f599412fbab510061269c35fbf9c59ba8491958a752cce8d38

Observation 7fd95f00-d9a4-45b7-9049-a5a45085108f · outbound

This paper cites It contains no misleading information or hallucinations.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models It contains no misleading information or hallucinations

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:48:05.079708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.435493Z digest=sha256:72a844d4ddf8dbba222ee5f4c34af3ce4c5ea9cf2c9c4ecd2b129d1966ecd9c1

Observation 790f3e81-e2ec-499a-aab5-3cfcc4769e69 · outbound

This paper cites preference.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models preference

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:48:05.065370Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.439958Z digest=sha256:5e4bd2f094697ca038e58f0080054de6931332dea9cf3f963c8220eb012b5ae1

Observation 0011564e-56b0-4aab-abea-a22f639a26d9 · outbound

This paper cites an unresolved cited work.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:48:05.049744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.445478Z digest=sha256:9a7f9139775c693ecec84f6a27bec22e05346f787c7a3333a9f7dbefe7e12860

Observation 3f29bfbe-07b8-4c9f-a11d-1b34ce1daa2a · outbound

This paper cites an unresolved cited work.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work

Reference 27

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:48:05.035054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.449913Z digest=sha256:e86f98c6c0b10a1327cd943c309ef2c7c45986c00412e5379563f4bda6323fc6

Observation dce37129-54b0-45fb-8324-9574dd80a10a · outbound

This paper cites an unresolved cited work.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:48:05.020579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.454405Z digest=sha256:31e6eea14890c65f4493c21d7c155054ef96856071fea5fe701cb43e7561ffd0

Observation 4af259a0-678c-44b6-8a77-ec21b6b0bcc1 · outbound

This paper cites an unresolved cited work.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:48:05.006108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.458745Z digest=sha256:49e34ec99c1c29601ef61da01629db8d5d1b6ffb72f0293f8981e00ec1acd5da

Observation 48b3e66d-2b68-455e-8ba5-0529e032da40 · outbound

This paper cites an unresolved cited work.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work

Reference 30

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:48:04.992129Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.463266Z digest=sha256:2bd3813db1b48d02a3d63b1cabd310357bd75c0b9e1c6087f4deb4e16489a2c2

Observation 40e32249-f472-4723-8fff-fd3817ceeb14 · outbound

This paper cites “markdown“‘markdown. This is an example of a code block in Markdown. You can see that it is formatted to look like it’s not part of the regular text flow.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models “markdown“‘markdown. This is an example of a code block in Markdown. You can see that it is formatted to look like it’s not part of the regular text flow

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:48:04.978143Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.467611Z digest=sha256:d16f2a727a90b8db3b4d30ac20f39ed6dd1ef16b1cb4149de6c2c83315ea0213

Observation 0d84453a-c1c0-4a46-8b93-550b18215763 · outbound

This paper cites an unresolved cited work.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:48:04.963674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.472349Z digest=sha256:86b468823bfc11ee12349e753412c3eb1bd904efb83b762c2ec14259e167a8e2

Observation 76909b5b-1482-4ff9-8e44-9f0fcfd9bfd7 · outbound

This paper cites an unresolved cited work.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:48:04.950425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.476696Z digest=sha256:3295f3cb769a6d2decae16f7b603122b3b24bf891723461abf94e693c8c3159f

Observation ebadc2cd-9ee1-40cb-b856-13ae76e0a176 · outbound

This paper cites an unresolved cited work.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:48:04.937069Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.480870Z digest=sha256:ffc5b184653df041679d448c56a93cc338f93ca683686d17019c8886312d44e8

Observation b862183f-ab10-4734-86ba-1069bdc073a7 · outbound

This paper cites an unresolved cited work.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work

Reference 35

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:48:04.923359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.485127Z digest=sha256:20028b7fef16ebd6a7938c8d646a1d3e5302f57d765cec52119e05be2c19f5ba

Observation 41fc8d2a-bdb7-4ad1-a839-ac4003e13cda · outbound

This paper cites an unresolved cited work.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work

Reference 37

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:48:04.895292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.494306Z digest=sha256:96a650fb9453a961911d9980765e71ae651f842ca4009bbd60dbc8f4182732df

Observation 46f191fa-7d8e-42ec-a79a-c6b79afd8bfe · outbound

This paper cites an unresolved cited work.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work

Reference 38

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:48:04.881178Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.498733Z digest=sha256:f3f899c9eb44ef4dd0a8e1d20f204bd74abae6aba0f85a7222d801b4447606a3

Observation 8905fa09-1251-4ee1-ad62-a8170e0d10ff · outbound

This paper cites an unresolved cited work.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work

Reference 39

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T05:48:04.866254Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.502975Z digest=sha256:c805daec636c717aefabfdf5adb996aa0b130f4948f1d600fa31258bc3f241dc

Observation 0666dcdc-6eb4-4a53-9d2a-e916603e2a7f · outbound

This paper cites an unresolved cited work.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:48:04.850567Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.507704Z digest=sha256:51c63ad056c2d4a8c82d22585831c8d12aa44868ae6b3b652d088847de4a3995

Observation 4b3002d7-2969-41ed-bdf5-5b0cce6667fe · outbound

This paper cites an unresolved cited work.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work

Reference 41

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:48:04.834072Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.511997Z digest=sha256:3947d37aeb650d0c9f3c51dc265a0fe227ab913bce73f1ef8e8fb1af1acac6d3

Observation 76cc2232-739b-4d7b-8197-e6c39dab316a · outbound

This paper cites an unresolved cited work.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work

Reference 42

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:48:04.819470Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.516378Z digest=sha256:ebcb957f924c34c14340920614abfd890a915c934b8128a926959e92130180bf

Observation d97ec2f6-8e34-480a-9c64-b69b5dd026f1 · outbound

This paper cites an unresolved cited work.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Unresolved cited work

Reference 44

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:48:04.788887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.525653Z digest=sha256:b748ccc5e6e19529d97e5c376824542ea4e93a695fa05139df8942b48cfb45a0

Observation a13a5ca8-7485-485b-9904-bfad94f57cc4 · outbound

This paper cites Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Rethinking Bradley-Terry Models in Preference-Based Reward Modeling: Foundations, Theory, and Alternatives

Reference 1027

Resolution
unresolved
no resolver link, observed 2026-08-07T05:48:04.364900Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:48:04.364900Z digest=sha256:ebc3e1fc8fc3a839a20dad2a082e6d522d8bb285d83925abd894cf6e39bb9c68

Observation 10c276b0-1f1d-4f37-ba7b-f54921c2139b · outbound

This paper cites KTO: Model Alignment as Prospect Theoretic Optimization.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models KTO: Model Alignment as Prospect Theoretic Optimization

Reference 1983

Resolution
unresolved
no resolver link, observed 2026-08-07T05:48:04.339130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:48:04.339130Z digest=sha256:244859c82ba2d0efc3cb5324f6fa4b0535d864065fce1341b8f9b1227776e57a

Observation 46dcb65d-320c-4102-bf03-cf7d6de321d5 · outbound

This paper cites Ryan Park, Rafael Rafailov, Stefano Ermon, and Chelsea Finn.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Ryan Park, Rafael Rafailov, Stefano Ermon, and Chelsea Finn

Reference 2022

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:48:05.273911Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.355394Z digest=sha256:e776e2d4c58647baf4139ce668ae91b7062c29adf848a5efa386eb84d045ec33

Observation b655c1f6-dca9-4a28-bdb0-b8ca25589367 · outbound

This paper cites Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Personalized Soups: Personalized Large Language Model Alignment via Post-hoc Parameter Merging

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-07T05:48:04.345187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:48:04.345187Z digest=sha256:4b2493351b6e7ecc3dfa1131028aab3dfffe31aa36c586d13fce24f237cf5d1e

Observation 46af202c-d500-4828-b593-0f2cd9e422e8 · outbound

This paper cites Mohammad Gheshlaghi Azar, Zhaohan Daniel Guo, Bi- lal Piot, Rémi Munos, Mark Rowland, Michal Valko, and Daniele Calandriello.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Mohammad Gheshlaghi Azar, Zhaohan Daniel Guo, Bi- lal Piot, Rémi Munos, Mark Rowland, Michal Valko, and Daniele Calandriello

Reference 2024

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:48:05.303657Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.324206Z digest=sha256:d17f4e9f19c959b7dcf12275806f440469b823c1a1e05fa4763d14e910836ec0

Observation 91b171db-f386-42d6-88ab-6aa8242fb34e · outbound

This paper cites Anirudhan Badrinath, Prabhat Agarwal, and Jiajing Xu.

AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models Anirudhan Badrinath, Prabhat Agarwal, and Jiajing Xu

Reference 4455

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:48:05.288823Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-07T05:48:04.329332Z digest=sha256:c0d76d74feb9db37fccbb8832cf0966113114ad24efd9721e2d19aed55755d79

Pith citing papers

Observation 998f2889-8323-4c19-990d-1a109a623c33 · inbound

A Survey on Progress in LLM Alignment from the Perspective of Reward Design cites this paper.

A Survey on Progress in LLM Alignment from the Perspective of Reward Design AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-16T00:52:06.745056Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T00:52:06.745056Z digest=sha256:e3b07568c8ea64e03a2c2adf000601f9e73862101154e407214ebbee3e849b6f

Observation 7f8dbf7d-2cd1-4c5d-8052-3748243927f9 · inbound

Failure Modes of Maximum Entropy RLHF cites this paper.

Failure Modes of Maximum Entropy RLHF AMoPO: Adaptive Multi-objective Preference Optimization without Reward Models and Reference Models

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:02:39.796329Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=arxiv_source observed=2026-05-18T14:02:11.084514Z digest=sha256:eb07a2d31c1365253aaf56f9306359fff56d87793e7381778d8a8a674d609536