Pith. sign in

Paper Citation Record · LEDGER

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models

As of 8 August 2026, this Paper Citation Record lists 32 of 32 outbound references and 0 inbound Pith citation observations for arXiv:2506.08643.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08643 v1

Coverage vector

measured 32 of 32 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:10:47.738852Z

measured 32 of 32 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

32 of 32 outbound references displayed

  • verified exact0
  • verified fuzzy14
  • unresolved18
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 27882341-1330-4df4-a831-7b405c940710 · outbound

This paper cites an unresolved cited work.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:10:48.172701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:10:47.622410Z digest=sha256:2baa9dadcbb79edc5b604a738ebd3e24fbf3d967819ddac33cf84ed9a584ffba

Observation 58cdae0f-28af-4945-9fb0-aee46b596d3b · outbound

This paper cites Deep Reinforcement Learning from Human Preferences.Advances in neural information processing systems, 30, 2017.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Deep Reinforcement Learning from Human Preferences.Advances in neural information processing systems, 30, 2017

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:10:48.161122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:10:47.626913Z digest=sha256:3e21a22a2021a1c3278cfb3de1e8ab403ccae9f33b3bb283a4dd5ff2d38be1e1

Observation bfdf5909-7334-42f5-a450-a7b5d1f9b5c4 · outbound

This paper cites Direct Preference Optimization: Your Language Model is Secretly a Reward Model.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Direct Preference Optimization: Your Language Model is Secretly a Reward Model

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:10:47.630917Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:10:47.630917Z digest=sha256:fd5b8a4d8cd3af51f06b585b7db212fd4c9fd3439b4b9adf2617a1335b2d85f7

Observation dbfc23ff-a6e2-47f1-b1dc-e4eb0066af96 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:10:47.635344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:10:47.635344Z digest=sha256:053d1c7f9d15da963415a0a64d8c5e735a087c7f365de08fcc73aec58d9fc376

Observation 95de08ac-82bd-4bc9-aeae-fa829c9d921b · outbound

This paper cites Self-Exploring Language Models: Active Preference Elicitation for Online Alignment.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Self-Exploring Language Models: Active Preference Elicitation for Online Alignment

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:10:47.640424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:10:47.640424Z digest=sha256:16509fd4a6f75c6c3d5a195c845fbb46c4b1ec2cad1caa31aca1a1e013c4a404

Observation 61805581-0aa0-46b5-bd5d-072e76c637b2 · outbound

This paper cites Rlhf workflow: From reward modeling to online rlhf, 2024.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Rlhf workflow: From reward modeling to online rlhf, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:10:48.147719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:10:47.645289Z digest=sha256:9667ad587983c54ab17bd819b1e17d1a076c9e57dfe65a016039f38bcda0fc96

Observation 62ebd2c6-aafd-4c20-a118-1df51773acc9 · outbound

This paper cites Iterative reasoning preference optimization, 2024.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Iterative reasoning preference optimization, 2024

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:10:48.135347Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:10:47.650814Z digest=sha256:7bdf2c5492ba51c0eb2326e5c1da0974ef1a3c4ea4c194ad1aeff8772f9efac8

Observation 3ea0c8e9-79b3-4378-886f-c3d82cd5f275 · outbound

This paper cites Self-consistencyimproveschainofthoughtreasoninginlanguage models, 2023.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Self-consistencyimproveschainofthoughtreasoninginlanguage models, 2023

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:10:48.123023Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:10:47.654564Z digest=sha256:43beee7015551e338d528b22ff7aede68dd7ce97c5db1849a14c8ffcef927bea

Observation 9dfa1c39-c779-4b98-b9b2-af5917958369 · outbound

This paper cites LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion, 2023.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion, 2023

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:10:48.111650Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:10:47.657915Z digest=sha256:f301a1eecd982cd4327cd1d1ddd113446ce350a05ea732d179ae09c298f7764d

Observation c2b5542a-cc9a-49fd-97c4-98d0d2093555 · outbound

This paper cites Self-refine: Iterative refinement with self-feedback, 2023.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Self-refine: Iterative refinement with self-feedback, 2023

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:10:48.099223Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:10:47.661140Z digest=sha256:98ba2b07b9add68dff4166b7477170e15b3156cdd1d3b384cafeaf5e60ef2032

Observation a094e9af-d346-4177-9362-b7fa7e2527f2 · outbound

This paper cites Mixture-of-Agents Enhances Large Language Model Capabilities.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Mixture-of-Agents Enhances Large Language Model Capabilities

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:10:47.664602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:10:47.664602Z digest=sha256:571be5d5eaaa78c3f022a63a180ff1bba89f7a25fffefe944378613ffcf58031

Observation 8cf393f2-5050-43b1-a125-9ebde517fe12 · outbound

This paper cites Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Math-Shepherd: Verify and Reinforce LLMs Step-by-step without Human Annotations

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:10:47.668467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:10:47.668467Z digest=sha256:5c6955b2c82f6a1f3fdc2e59a2e544c295add500cf4baaa347e5bf31878ebfc1

Observation 987e069c-93af-41e1-af18-175339f600f2 · outbound

This paper cites Let’s verify step by step.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Let’s verify step by step

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:10:47.672230Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:10:47.672230Z digest=sha256:e6b8ff75cb7306f65b40263015b35d012505b6efbe07be17cf9703630046df8c

Observation 0dec1e55-ad11-4e91-9a5c-75e7c0d30d8c · outbound

This paper cites Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Scaling llm test-time compute optimally can be more effective than scaling model parameters, 2024

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:10:47.676429Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:10:47.676429Z digest=sha256:06321f593141d265d2873c206135f0dbc5cb13af6e5f5126d6f101fc23d56bff

Observation ac9d8ef4-c215-4eb4-a48f-6cdc67a58d81 · outbound

This paper cites Proximal Policy Optimization Algorithms, 2017.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Proximal Policy Optimization Algorithms, 2017

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:10:47.680150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:10:47.680150Z digest=sha256:35b2b7e13e2fb68d18ac227601d2f0b2dd945e67c9c5731d3576a79bc9feef8a

Observation beddbc8d-7e06-4ce8-be57-2db901814d09 · outbound

This paper cites OpenAI o1 System Card.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models OpenAI o1 System Card

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:10:47.683724Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:10:47.683724Z digest=sha256:b3114437f3fc3b7b91a13578154999d993831ecd41d84f7d500c801e7fec880e

Observation ad9307ea-2f30-4894-b58f-e8a7cbacd983 · outbound

This paper cites Introducing openai o3 and o4-mini.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Introducing openai o3 and o4-mini

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:10:48.065812Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:10:47.687742Z digest=sha256:a36445598b651d098c8e2ef47c25f84114520998ede00bd2207d40e464eb14b8

Observation 74987d55-c4a2-49a2-a2ef-ccb0736a1bd5 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:10:47.692219Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:10:47.692219Z digest=sha256:04dbedd3f9fdfec06a1aba682d4522ebf605097a04f8d9847e4817edb1bc1c96

Observation 326c78ae-8728-4f82-a649-650eb2bd7896 · outbound

This paper cites Least-to-most prompting enables complex reasoning in large language models, 2023.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Least-to-most prompting enables complex reasoning in large language models, 2023

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:10:47.695845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:10:47.695845Z digest=sha256:04601a2a72a1a38bfee165e28c13898f04d3779b5366ba0bdeae311d728d53df

Observation 6b271064-6b92-40f7-b857-bd7ba9331e1c · outbound

This paper cites LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models LLM-Blender: Ensembling Large Language Models with Pairwise Ranking and Generative Fusion

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-07T05:10:47.699915Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:10:47.699915Z digest=sha256:f0b3b14c36ea5a7ae4bee8aa9bd676560aff0989c2b8479489834d0839e11938

Observation c7a3e8e7-6e2e-4467-9b44-58a8ec10ad32 · outbound

This paper cites University of Michigan Press, 1975.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models University of Michigan Press, 1975

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:10:48.045880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:10:47.703283Z digest=sha256:4a0a7ba7578a8e8dd0a2b10cfbb9e922bd9785ac3c69ac649e4145b750eee546

Observation da481354-9ed3-460d-85e9-a803f3ad467d · outbound

This paper cites an unresolved cited work.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Unresolved cited work

Reference 22

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:10:48.034197Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:10:47.706864Z digest=sha256:66ab8902e0031bb95b4608efcfa1e0e851afb2ce45789d52123168ec98aa186a

Observation 3d61b798-5180-475c-9188-4a74497663fb · outbound

This paper cites Onevolution,search,optimization,geneticalgorithmsandmartialarts: Towards memetic algorithms.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Onevolution,search,optimization,geneticalgorithmsandmartialarts: Towards memetic algorithms

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:10:48.022055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:10:47.710202Z digest=sha256:36cf2ee8e31115239fc7aa950c1b41317a9385a9701b917cc8aba25cc5a1ac45

Observation 4778488d-0245-4a4e-8f9a-8ea03269eb1b · outbound

This paper cites Memetic computation—past, present & future [research frontier].IEEE Computational Intelligence Magazine, 5(2):24–31, 2010.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Memetic computation—past, present & future [research frontier].IEEE Computational Intelligence Magazine, 5(2):24–31, 2010

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:10:48.010168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:10:47.713154Z digest=sha256:1ec9c243a77ca732a8f6df44ce39cd5b0c140aedcfe044cdd3590d818808c530

Observation 26ea52d5-a577-43f9-92ca-c25801b56d7b · outbound

This paper cites Omran, A.P.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Omran, A.P

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:10:47.997441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:10:47.716566Z digest=sha256:a9e7b4c2fe17e3fbf09bb2b8abfe55caa8d00140917b537b25ef88170f3eae71

Observation 5904a2f6-ff94-4461-ac0c-1406a6003c47 · outbound

This paper cites Evolving code with a large language model, 2024.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Evolving code with a large language model, 2024

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:10:47.985176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:10:47.719755Z digest=sha256:0aeb6e4a52e4a9862ffd421dd03a8a32bbcb773709ddc201151c25c3e3b94070

Observation 0c959754-f7ed-4857-b1bf-b5d850e2932a · outbound

This paper cites Evolving deeper llm thinking, 2025.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Evolving deeper llm thinking, 2025

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:10:47.973970Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:10:47.722758Z digest=sha256:acf7ebbb95ddba156e9d2a76cc17a34b4fb5ff0cb1ac5b7bb63d29270b9277f3

Observation 2c1c3d80-f312-4fc0-ba9c-e86e7df256c8 · outbound

This paper cites Llmrefine: Pinpointing and refining large language models via fine-grained actionable feedback, 2024.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Llmrefine: Pinpointing and refining large language models via fine-grained actionable feedback, 2024

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:10:47.961634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T05:10:47.726073Z digest=sha256:6e8f5b1a4e09d7389f6f2511a05b43e3e77e328af4ee47df42ab8f443d4619e8

Observation 6276d0b8-982b-4b02-92cc-5e67f32e6b2e · outbound

This paper cites What makes a reward model a good teacher? an optimization perspective.arXiv preprint arXiv:2503.15477, 2025.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models What makes a reward model a good teacher? an optimization perspective.arXiv preprint arXiv:2503.15477, 2025

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:10:47.729198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:10:47.729198Z digest=sha256:51e69bcda0b0cc43ca22a52ca208f47a6e991bed06fff96f79e20e8ed780eeca

Observation 74b1f979-cc8e-47cb-b73b-fd3f452de62d · outbound

This paper cites Scaling laws for reward model overoptimization.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Scaling laws for reward model overoptimization

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:10:47.732461Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:10:47.732461Z digest=sha256:f8300a3895a4e251478b6392d2c2cb0115ae4a3734a9fd9d226ad3f5fccd500c

Observation 71434e16-9a23-4888-9167-faa6a70701b9 · outbound

This paper cites Countering Reward Over-optimization in LLM with Demonstration-Guided Reinforcement Learning.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Countering Reward Over-optimization in LLM with Demonstration-Guided Reinforcement Learning

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:10:47.735478Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:10:47.735478Z digest=sha256:8d11efd3c900d9e5294628295217a0d29eb8a16d49fe4ee8aba47f59d73ef475

Observation bae79e4c-a56b-4df0-a91f-d32ef34bc8cd · outbound

This paper cites Confronting Reward Model Overoptimization with Constrained RLHF.

MEMETRON: Metaheuristic Mechanisms for Test-time Response Optimization of Large Language Models Confronting Reward Model Overoptimization with Constrained RLHF

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:10:47.738852Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:10:47.738852Z digest=sha256:a1103096cff2800c4d09a01de2638e690213d076820aad57d3e96e4f876de251

Pith citing papers

No inbound Pith citation observations are available.